|
Guanren Zhou
I am a Ph.D. candidate at UC Berkeley, advised by Prof. Khalid M. Mosalam in the STAIR Lab. My research focuses on grounded multimodal reasoning and generation across vision-language and NLP/LLM systems.
I am seeking a Summer 2027 internship in AI/ML research, machine learning engineering, or software engineering for AI systems.
Email /
Scholar /
GitHub /
LinkedIn
|
|
Research
My work asks how models can reason over complex visual and textual inputs while grounding their outputs in identifiable image regions and source passages.
|
|
|
Entity-Guided Expert-Grounded Visual Diagnosis for Structures in the Wild
Guanren Zhou, Xiaolei Chu, Ziqi Wang, Khalid M. Mosalam
Manuscript in preparation
Hierarchical SAM 3 retrieval and DINOv3 specialist classifiers ground a frozen VLM’s fine-grained diagnoses in segmentation masks, enabling traceability without new task-specific annotations or VLM fine-tuning.
|
|
|
GeoScout: Caption-Conditioned Next-Best-View Policies for Generalizable 3D Reconstruction
Xiaolei Chu, Guanren Zhou, Khalid M. Mosalam, Ziqi Wang
Manuscript in preparation
code
GeoScout conditions a PPO next-best-view policy on refined VLM geometry captions, using language as a prior over unobserved shape to reconstruct unseen object instances with fewer views than a matched no-caption policy.
|
|
|
Reinforcement Learning for Intelligent Optimization of Steel Structures: A Review
Yuqing Gao, Yuhang Lu, Meiyu Du, Wei Wang, Guanren Zhou
Smart Construction, 2026(3), 0014
paper
This review contrasts case-specific optimization with reusable RL policies, highlighting cross-instance generalization, constraint satisfaction, and verifiable evaluation.
|
|
|
Social Amplification Dominates Collective Hazard Response
Xiaolei Chu, Guanren Zhou, Marco Broccardo, Didier Sornette, Khalid M. Mosalam, Ziqi Wang
arXiv preprint, 2026
arXiv
Fine-tuned BERTweet estimates state-level stress prevalence from social-media posts to calibrate an interpretable network model, revealing that social influence outweighed direct exposure in over 80% of U.S. states.
|
|
|
Automated Virtual Earthquake Reconnaissance Reporting Using Natural Language Processing
Guanren Zhou, Khalid M. Mosalam
Natural Hazards Review, 2025
project page /
paper
An event-triggered pipeline combines fine-tuned RoBERTa, VLM extraction from event graphics, and citation-constrained LLM synthesis to turn heterogeneous web sources into source-attributed briefings and temporal analyses.
|
|
|
A Large Language Model for Disaster Structural Reconnaissance Summarization
Yuqing Gao, Guanren Zhou, Khalid M. Mosalam
18th World Conference on Earthquake Engineering (WCEE), 2024
paper
A hierarchical LLM pipeline combines CNN-extracted visual attributes with structured metadata to generate per-instance reports and cross-instance summaries from multimodal observations.
|
Target-Conditioned Event Attribution in Multi-Event Documents
Ongoing work
This project formulates proposition-to-event grounding as a four-way target-conditioned classification task that exposes false attribution hidden by document-level relevance; benchmark construction is underway.
|
|