Guanren Zhou

I am a Ph.D. candidate at UC Berkeley, advised by Prof. Khalid M. Mosalam in the STAIR Lab. My research focuses on grounded multimodal reasoning and generation across vision-language and NLP/LLM systems.

I am seeking a Summer 2027 internship in AI/ML research, machine learning engineering, or software engineering for AI systems.

Email  /  Scholar  /  GitHub  /  LinkedIn

Profile photo of Guanren Zhou

Research

My work asks how models can reason over complex visual and textual inputs while grounding their outputs in identifiable image regions and source passages.

Entity-Guided Expert-Grounded Visual Diagnosis for Structures in the Wild
Guanren Zhou, Xiaolei Chu, Ziqi Wang, Khalid M. Mosalam
Manuscript in preparation

We combine hierarchical SAM 3 grounding with context-aware filtering and DINOv3-based specialist models to structure visual inputs for VLM reasoning. Citation-aligned masks and relational cues connect generated claims to the image regions that support them, without new end-to-end diagnosis annotations or VLM fine-tuning.

Caption-Conditioned Next-Best-View Policies for Generalizable 3D Reconstruction
Xiaolei Chu, Guanren Zhou, Khalid M. Mosalam, Ziqi Wang
Manuscript in preparation
code

GeoScout combines a 3D occupancy belief with refined VLM geometry captions to condition a PPO next-best-view policy. The captions provide semantic priors over unobserved shape, helping the policy reconstruct held-out object instances with fewer views than an otherwise matched no-caption policy.

Reinforcement Learning for Intelligent Optimization of Steel Structures: A Review
Yuqing Gao, Yuhang Lu, Meiyu Du, Wei Wang, Guanren Zhou
Smart Construction, 2026(3), 0014
paper

This review contrasts case-specific optimization with reusable RL policies, highlighting cross-instance generalization, constraint satisfaction, and verifiable evaluation.

Social Amplification Dominates Collective Hazard Response
Xiaolei Chu, Guanren Zhou, Marco Broccardo, Didier Sornette, Khalid M. Mosalam, Ziqi Wang
arXiv preprint, 2026
arXiv

Fine-tuned BERTweet estimates state-level stress prevalence from social-media posts to calibrate an interpretable network model, revealing that social influence outweighed direct exposure in over 80% of U.S. states.

Automated Virtual Earthquake Reconnaissance Reporting Using Natural Language Processing
Guanren Zhou, Khalid M. Mosalam
Natural Hazards Review, 2025
paper / product

We build an event-centric multimodal system that uses fine-tuned RoBERTa models and VLMs to convert web text and event graphics into provenance-preserving records. The shared records support source-grounded LLM synthesis and longitudinal analysis as new information arrives.

A Large Language Model for Disaster Structural Reconnaissance Summarization
Yuqing Gao, Guanren Zhou, Khalid M. Mosalam
18th World Conference on Earthquake Engineering (WCEE), 2024
paper

A hierarchical LLM pipeline combines CNN-extracted visual attributes with structured metadata to generate per-instance reports and cross-instance summaries from multimodal observations.

Ongoing Research

Target-Conditioned Event Grounding in Multi-Event Documents
Ongoing work

We formulate target-conditioned event grounding at the proposition level, distinguishing statements about the target event, target-related context, other events, or no event instance. We are developing a two-pass annotation and evaluation framework that separates grounding from proposition-segmentation errors.


Website design borrows from Jon Barron