Foundation vision-language models recognize objects, interpret instructions, and reason about spatial relations. Yet a robot also needs to know where to move, how to orient its gripper, and whether an action actually worked. Inferring these details from RGB images alone leaves fundamental uncertainty about depth, contact, and 3D spatial relationships. We identify accessible perceptual evidence as the critical missing piece between VLM competence and manipulation capability.
Robo-Harness K1 exposes perception as tools. The agent queries calibrated depth, inspects visual anchors with persistent object identities, evaluates grasp hypotheses, and selects generic motions from the returned evidence. Visual reference lines overlay calibrated measurements onto the VLM's existing RGB views, making 3D object relations, orientations, and distances visible without a depth encoder. The K1 harness supplies calibrated spatial evidence, not task-specific pick/place primitives or hidden object poses. The VLM remains responsible for choosing the target, the motion, and the recovery strategy.
The same interface supports both frontier models without fine-tuning and training compact student models: tool-call traces naturally align with next-token prediction, so successful interactions become supervision for what to observe, what to measure, and which tool to call next.