Guanren Zhou

I am a Ph.D. candidate at UC Berkeley, advised by Prof. Khalid M. Mosalam in the STAIR Lab. My research focuses on grounded multimodal reasoning and generation across vision-language and NLP/LLM systems.

I am seeking a Summer 2027 internship in AI/ML research, machine learning engineering, or software engineering for AI systems.

Email  /  Scholar  /  GitHub  /  LinkedIn

Profile photo of Guanren Zhou

Research

My work asks how models can reason over complex visual and textual inputs while grounding their outputs in identifiable image regions and source passages.

Entity-Guided Expert-Grounded Visual Diagnosis for Structures in the Wild
Guanren Zhou, Xiaolei Chu, Ziqi Wang, Khalid M. Mosalam
Manuscript in preparation

Hierarchical SAM 3 retrieval and DINOv3 specialist classifiers ground a frozen VLM’s fine-grained diagnoses in segmentation masks, enabling traceability without new task-specific annotations or VLM fine-tuning.

GeoScout: Caption-Conditioned Next-Best-View Policies for Generalizable 3D Reconstruction
Xiaolei Chu, Guanren Zhou, Khalid M. Mosalam, Ziqi Wang
Manuscript in preparation
code

GeoScout conditions a PPO next-best-view policy on refined VLM geometry captions, using language as a prior over unobserved shape to reconstruct unseen object instances with fewer views than a matched no-caption policy.

Reinforcement Learning for Intelligent Optimization of Steel Structures: A Review
Yuqing Gao, Yuhang Lu, Meiyu Du, Wei Wang, Guanren Zhou
Smart Construction, 2026(3), 0014
paper

This review contrasts case-specific optimization with reusable RL policies, highlighting cross-instance generalization, constraint satisfaction, and verifiable evaluation.

Social Amplification Dominates Collective Hazard Response
Xiaolei Chu, Guanren Zhou, Marco Broccardo, Didier Sornette, Khalid M. Mosalam, Ziqi Wang
arXiv preprint, 2026
arXiv

Fine-tuned BERTweet estimates state-level stress prevalence from social-media posts to calibrate an interpretable network model, revealing that social influence outweighed direct exposure in over 80% of U.S. states.

Automated Virtual Earthquake Reconnaissance Reporting Using Natural Language Processing
Guanren Zhou, Khalid M. Mosalam
Natural Hazards Review, 2025
project page / paper

An event-triggered pipeline combines fine-tuned RoBERTa, VLM extraction from event graphics, and citation-constrained LLM synthesis to turn heterogeneous web sources into source-attributed briefings and temporal analyses.

A Large Language Model for Disaster Structural Reconnaissance Summarization
Yuqing Gao, Guanren Zhou, Khalid M. Mosalam
18th World Conference on Earthquake Engineering (WCEE), 2024
paper

A hierarchical LLM pipeline combines CNN-extracted visual attributes with structured metadata to generate per-instance reports and cross-instance summaries from multimodal observations.

Ongoing Research

Target-Conditioned Event Attribution in Multi-Event Documents
Ongoing work

This project formulates proposition-to-event grounding as a four-way target-conditioned classification task that exposes false attribution hidden by document-level relevance; benchmark construction is underway.


Website design borrows from Jon Barron