Our Projects

  • ViewAgent

    ViewAgent

    ViewAgent studies whether VLM agents can plan camera moves to find a target view in real 3D scenes. Its ViewSuite benchmark exposes a multi-turn view-planning gap and shows how on-policy view graph distillation improves spatial planning.

  • RAGEN-2

    RAGEN-2

    RAGEN-2 studies reasoning collapse in agentic RL — how multi-turn LLM agents drift into template and low-entropy failure modes — and how to keep reasoning diverse and input-grounded. Oral (Top 0.7%) at ICML 2026 and Best Paper (Top 1%) at the CVPR 2026 MMRAgI Workshop.

  • MindCube

    MindCube

    MindCube probes whether VLMs can build a spatial mental model from limited views — tracking object position and orientation and mentally simulating unseen viewpoints. ICLR 2026; Best Paper at the ICCV 2025 SP4V Workshop and selected as The Best of ICCV.

  • ENACT

    ENACT

    ENACT benchmarks the embodied cognition of VLMs by having them world-model egocentric interaction — predicting the forward and inverse dynamics of first-person manipulation. ICLR 2026, with two Outstanding Paper Awards at ICLR 2026 workshops.

  • Theory of Space

    Theory of Space

    Theory of Space asks whether foundation models can construct, revise, and exploit a spatial belief through active exploration — integrating partial observations into a globally consistent map. ICLR 2026.

  • RAGEN

    RAGEN

    We introduce RAGEN to train LLM reasoning agents via RL in multi-turn, stochastic environments. RAGEN is formulated with MDP and optimized through Reasoning-Interaction Chain Optimization (RICO). RAGEN-0.5B is trained across three agentic tasks, showing intriguing reasoning patterns.

  • VAGEN

    VAGEN

    VAGEN is an RL framework improving VLM agent training with the TRICO algorithm. By selectively focusing on critical tokens and enhancing cross-turn credit assignment, TRICO outperforms prior methods on visual agentic tasks.

  • Embodied Agent Interface

    Embodied Agent Interface

    Current evaluations of LLMs in embodied AI lack standardization and detailed error analysis. Our introduce a unified interface (Embodied Agent Interface) for diverse tasks and LLM modules (planning, decomposition, etc.) and fine-grained metrics (identifying hallucination, affordance errors, etc.). This enables systematic assessment, pinpointing specific LLM limitations and strengths to inform more effective integration into embodied agents.

  • Long Video Haystack

    Long Video Haystack

    We introduce LongVideoHaystack, a 480-hour video temporal search dataset with 15,092 human-annotated instances, where SOTA scores 2.1% Temporal F1.