Agent Evaluation · RLVRA01
Agents' Last Exam
A large-scale benchmark for measuring frontier agents on long-horizon, economically valuable tasks. ALE covers 1K+ tasks across 55 subdomains and 13 industry clusters, running agents in real OS environments and scoring their outputs with automated evaluators.
I helped build the end-to-end benchmark and evaluation system: executable task environments, computer-use and tool interfaces, scalable execution, and programmatic graders that enable verifiable rewards and RLVR.
Robot Learning · DataR02
Robot-Learning Data Benchmark
An ongoing effort with Zhuo Xu and Prof. Masayoshi Tomizuka to build a large-scale benchmark for robot learning. I work on cleaning and curating heterogeneous egocentric manipulation and human-object interaction data, including EgoVerse, EgoDex, and HOI4D.
The goal is to turn diverse human demonstrations into reliable, usable resources for training and evaluating embodied systems.
Ongoing research
Computer Vision · UAVV03
UAV-to-UAV Detection & Tracking
Research with Prof. Avideh Zakhor on zero-shot UAV-to-UAV detection and multi-object tracking under high-speed motion and strong ego-motion. I develop PyTorch pipelines that integrate optical-flow-based tracking with detection and tracking baselines.
This work targets reliable perception when objects are small, motion is fast, and background dynamics are substantial.
CodeTargeting CVPR submission Embodied Systems · VLAE04
Robot Delivery with VLA
At Starbot, I worked on integrating vision-language-action models and agentic workflows for robot planning and control in a real delivery system. The system was demonstrated at CES 2026.
This experience connected foundation-model reasoning with the constraints of embodied execution and deployment.