|
Guanren Zhou
I am a Ph.D. candidate at UC Berkeley, advised by Prof. Khalid M. Mosalam in the STAIR Lab. My research focuses on grounded multimodal reasoning and generation across vision-language and NLP/LLM systems.
I am seeking a Summer 2027 internship in AI/ML research, machine learning engineering, or software engineering for AI systems.
Email /
Scholar /
GitHub /
LinkedIn
|
|
Research
My work asks how models can reason over complex visual and textual inputs while grounding their outputs in identifiable image regions and source passages.
|
|
|
Entity-Guided Expert-Grounded Visual Diagnosis for Structures in the Wild
Guanren Zhou, Xiaolei Chu, Ziqi Wang, Khalid M. Mosalam
Manuscript in preparation
We combine hierarchical SAM 3 grounding with context-aware filtering and DINOv3-based specialist models to structure visual inputs for VLM reasoning. Citation-aligned masks and relational cues connect generated claims to the image regions that support them, without new end-to-end diagnosis annotations or VLM fine-tuning.
|
|
|
Caption-Conditioned Next-Best-View Policies for Generalizable 3D Reconstruction
Xiaolei Chu, Guanren Zhou, Khalid M. Mosalam, Ziqi Wang
Manuscript in preparation
code
GeoScout combines a 3D occupancy belief with refined VLM geometry captions to condition a PPO next-best-view policy. The captions provide semantic priors over unobserved shape, helping the policy reconstruct held-out object instances with fewer views than an otherwise matched no-caption policy.
|
|
|
Reinforcement Learning for Intelligent Optimization of Steel Structures: A Review
Yuqing Gao, Yuhang Lu, Meiyu Du, Wei Wang, Guanren Zhou
Smart Construction, 2026(3), 0014
paper
This review contrasts case-specific optimization with reusable RL policies, highlighting cross-instance generalization, constraint satisfaction, and verifiable evaluation.
|
|
|
Social Amplification Dominates Collective Hazard Response
Xiaolei Chu, Guanren Zhou, Marco Broccardo, Didier Sornette, Khalid M. Mosalam, Ziqi Wang
arXiv preprint, 2026
arXiv
Fine-tuned BERTweet estimates state-level stress prevalence from social-media posts to calibrate an interpretable network model, revealing that social influence outweighed direct exposure in over 80% of U.S. states.
|
|
|
Automated Virtual Earthquake Reconnaissance Reporting Using Natural Language Processing
Guanren Zhou, Khalid M. Mosalam
Natural Hazards Review, 2025
paper /
product
We build an event-centric multimodal system that uses fine-tuned RoBERTa models and VLMs to convert web text and event graphics into provenance-preserving records. The shared records support source-grounded LLM synthesis and longitudinal analysis as new information arrives.
|
|
|
A Large Language Model for Disaster Structural Reconnaissance Summarization
Yuqing Gao, Guanren Zhou, Khalid M. Mosalam
18th World Conference on Earthquake Engineering (WCEE), 2024
paper
A hierarchical LLM pipeline combines CNN-extracted visual attributes with structured metadata to generate per-instance reports and cross-instance summaries from multimodal observations.
|
Target-Conditioned Event Grounding in Multi-Event Documents
Ongoing work
We formulate target-conditioned event grounding at the proposition level, distinguishing statements about the target event, target-related context, other events, or no event instance. We are developing a two-pass annotation and evaluation framework that separates grounding from proposition-segmentation errors.
|
|