Robo-Harness K1

Harnessing Robot-Use Agents
via Perception Augmentation

Overview

From VLM Understanding to Manipulation Capability

Foundation vision-language models recognize objects, interpret instructions, and reason about spatial relations. Yet a robot also needs to know where to move, how to orient its gripper, and whether an action actually worked. Inferring these details from RGB images alone leaves fundamental uncertainty about depth, contact, and 3D spatial relationships. We identify accessible perceptual evidence as the critical missing piece between VLM competence and manipulation capability.

Robo-Harness K1 exposes perception as tools. The agent queries calibrated depth, inspects visual anchors with persistent object identities, evaluates grasp hypotheses, and selects generic motions from the returned evidence. Visual reference lines overlay calibrated measurements onto the VLM's existing RGB views, making 3D object relations, orientations, and distances visible without a depth encoder. The K1 harness supplies calibrated spatial evidence, not task-specific pick/place primitives or hidden object poses. The VLM remains responsible for choosing the target, the motion, and the recovery strategy.

The same interface supports both frontier models without fine-tuning and training compact student models: tool-call traces naturally align with next-token prediction, so successful interactions become supervision for what to observe, what to measure, and which tool to call next.

In action

Successful demonstrations

Gemini 3.7 Flash with Robo-Harness K1. Current-view perception overlays show what the agent observed, alongside its function calls.

LIBERO-PRO

Franka Panda
Spatial · Tools + execution

Bowl onto plate

Locate the black bowl between the plate and the ramekin, then place it on the plate.

Object · Tools + execution

Ketchup into basket

Pick the target object and transfer it into the basket.

Goal · Tools + execution

Open the middle drawer

Ground the drawer handle and control contact while pulling the drawer open.

RoboSuite

Across robot arms
UR5e · Tools + execution

Nut assembly

Grasp the square nut, align it above the peg, and lower it into place.

IIWA · Tools + execution

Cube stacking

Pick up a cube and stack it on the other cube.

Dual Panda · Tools + execution

Two-arm lifting

Lift the pot with two arms. Native success requires the pot bottom to rise more than 10 cm above the table.

Decision observations pause briefly to show the recorded perception overlays and tool calls (key arguments shown). Motion segments retain every recorded control step at simulator speed; overlays are not carried onto changed views. Each clip ends with a two-second hold of the successful terminal state. API waiting time is omitted. Use fullscreen to read the annotations.

The method

Perception as Tools

K1 connects a vision-language model to calibrated perception, motion, and memory tools through the VLM's existing image-and-language interface. The agent selects what to measure, acts on the evidence, and observes what actually happened.

Region grounding, depth and geometry measurement, anchor tracking, and grasp candidates on existing camera images
Ground regions, measure geometry, track visual anchors, and inspect grasp hypotheses in the existing camera views.
01

Ground and Measure

A text query returns labeled candidate regions with persistent identifiers for the agent to inspect. Calibrated depth queries and region geometry convert RGB-D observations into metric surface positions, principal directions, and relative displacements. Evidence is projected back into the existing external and wrist-camera views as visual reference lines. The agent reads 3D relations directly from annotated images without interpreting raw point clouds or depth arrays.

02

Act and Verify

The agent selects a grasp hypothesis or specifies a motion target. Generic tools control translation, rotation, and gripper opening at adaptive granularity: fine-grained bounded steps for precise alignment, and direct target moves for longer transits that reduce model calls and keep the context history clean. Every action reports achieved motion and residual error, so a requested pose is not confused with a completed movement.

03

Maintain Spatial Evidence

Each detected region receives a persistent identity anchor that survives across observations. Anchor tracking provides frame-to-frame correspondences lifted into world coordinates, separating apparent camera motion from actual surface displacement. Validity flags distinguish live measurements from uncertain or historical references. A sliding context window, searchable full history, and agent-maintained progress records carry information through longer episodes.

04

Train from Tool-Call Traces

Teacher interactions are filtered into executable next-tool-call examples, paired with the observations and tool evidence available at each decision. Qwen3.5-9B is fine-tuned with language-side LoRA and a frozen vision encoder. The student uses the same frozen harness at evaluation without a separate action head. Tool-call traces align directly with the VLM's native next-token prediction objective, enabling a training paradigm with better sample efficiency and generalization than action regression.

Measurements describe observed surfaces, not hidden geometry. Tracking can become uncertain, and a grasp candidate is a hypothesis, not a guarantee of collision-free motion or secure contact. Re-observation and execution feedback remain part of the control loop.

Experiments

Results

LIBERO-PRO

On matched configurations, K1 augments GPT-6 Astra from 61.1% to 88.9%, a gain of 27.8 percentage points. Gemini 3.7 Flash with K1 reaches 77.8%, surpassing RGB-only Astra despite being a weaker model. This shows that perception tools and model capability are complementary: a weaker model equipped with K1 can outperform a stronger model without it, and K1 further lifts an already strong model.

LIBERO-Pro accuracy: K1 with GPT-6 Astra 88.9%, K1 with Gemini 3.7 Flash 77.8%, RGB-only GPT-6 Astra 61.1%, plus published reference results
The top three rows use matched task configurations. Hatched bars are published references with different evaluation protocols, not matched reruns.

Agentic Post-Training

At epoch 5, Qwen3.5-9B with K1 reaches 51.2% on training-matched configurations (A), 44.2% on new initial states (B), and 13.9% on held-out task conditions (C), versus OpenVLA's 20.9%, 30.2%, and 0.0% respectively. K1 RUA is the only method to succeed on held-out tasks where every VLA baseline scores 0.0%.

Learning curves comparing Qwen with K1, RGB-only Qwen, OpenVLA, pi0.5, and Qwen VLA
All methods share the same 107 source episodes. B evaluates new initial states of seen conditions; C holds out entire task conditions from fine-tuning. Generalization on B and C does not degrade significantly despite the limited training data. Further scaling of demonstration diversity is a natural next step.

RoboSuite Transfer

Gemini with K1 transfers to RoboSuite without any target-environment fine-tuning. On the four shared tasks (cube lifting, restacking, stacking, and nut assembly):

90.0%Panda

88.8%UR5e

86.2%IIWA

The small spread across arms supports separating decisions from actuation: the agent reasons about measured geometry, while embodiment-specific adapters control each arm. All arms use PandaGripper to isolate changes in arm kinematics. The broader Panda evaluation includes wiping and two-arm tasks, reaching 72.1% overall. Two-arm lifting reaches 75.0%, while handover reaches only 5.0%, exposing coordination limits beyond geometric target-reaching.