Embodied Agent Arena

Are frontier VLM agents ready to be robot generalists?

An empirical study across perception, reasoning, and action.

1 HKUST (Guangzhou)2 The Chinese University of Hong Kong3 Knowin AI

01 The arena

From seeing a scene
to completing a task.

One arena tests what agents can perceive, reason about, and physically achieve across five robotic capabilities.

1,000evaluation cases
32 + 1sources + GeoProbe
5capability domains
7frontier VLM agents

Introducing GeoProbe

Separate object motion
from camera motion.

168 new cases isolate camera motion, object motion, depth, and scale.

How the arena is built. Task examples and the five-domain composition sit alongside GeoProbe's estimation targets and the workflow connecting model actions to evaluation.

02 A deep dive of Astra

Where Astra leads. Where it falls short.

Compare all seven agents

Performance across five capability domains
AgentGeometrySpatial reasoningAffordancePlanningManipulation
Rotation ↓degreesTranslation ↓cmPass@1 ↑%AbsRel ↓relative errorMask IoU ↑overlapContact ↑%Pass@1 ↑%Task success ↑%

Bold marks the best score; arrows show the better direction. ± is the standard deviation across runs. Manipulation uses the baseline setting.

03 From scores to behavior

Watch the decisions
behind the scores.

Citation