🔥 Newest
TL;DR
An evaluation framework where an off-the-shelf VLM drives a full 3D human body by issuing atomic skill commands, so its decisions are tested with real physical consequences.
Hi, I am Jiawei Gu, a PhD student in Computer Science at the University of Washington, advised by Ranjay Krishna. I also work closely with Yejin Choi and Manling Li.
My research began with scaling language models as a foundation for intelligence. I now build on this by unifying modalities to teach models to think in words and images, and act in the world. My focus is scaling native multimodal reasoning, including understanding, world modeling, and embodied reasoning.
* denotes equal contribution
🔥 Newest
An evaluation framework where an off-the-shelf VLM drives a full 3D human body by issuing atomic skill commands, so its decisions are tested with real physical consequences.
ICLR 2026
ThinkMorph enables true multimodal chain-of-thought reasoning where text and vision complement each other, achieving emergent intelligence and strong generalization with just 24K training samples.
ICML 2025 Oral
We contribute EMMA (Enhanced MultiModal reAsoning), a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding. SOTA models struggle big time.
ACL 2025 Oral + 🏆 Best Paper Honorable Mention
We propose Speculative Reward Models that dramatically improve LLM decision-making efficiency while maintaining high accuracy.
A comprehensive survey of using LLMs as evaluators: how to build reliable LLM judges, how to evaluate them, and where they are applied.
EMNLP 2024
We propose CMR Scaling Law to predict critical mixture ratios for continual pre-training, enabling efficient domain adaptation while preventing catastrophic forgetting.