A PHYSICAL BENCHMARK NETWORK FOR ROBOT MANIPULATION

Your policy.
Our robot.
One comparable result.

RobotReplica is a network of organizations that host standardized robot manipulation benchmarks. Find a site with the same robot you use, send us your policy, and we evaluate it on maintained real-world tasks.

RobotReplica network connecting physical robot evaluation sites
A DISTRIBUTED PHYSICAL BENCHMARK NETWORKMULTIPLE ROBOTS

OVERVIEW

A shared evaluation service
for real robots.

Robotics results are difficult to compare when every lab uses a different robot, scene, and protocol. RobotReplica keeps physical benchmark sites running so the community can evaluate on consistent hardware and tasks.

01

Physical benchmark sites

Each partner maintains a robot, workspace, task objects, cameras, and an evaluation protocol.

02

Robot-matched evaluation

You choose the site with the same robot as your system. The site runs your policy for you.

03

Verified leaderboards

Every robot and site has its own board, keeping scores comparable and evidence traceable.

THE ROBOTREPLICA MODEL

Researchers do not need to rebuild the full benchmark. The benchmark stays at the host site; policies travel to it.

CURRENT SITES

Start with the robot
you already use.

Two organizations are building the first RobotReplica sites. Each site owns its hardware, tasks, evaluation process, and robot-specific leaderboard.

OpenArm standardized evaluation cellSITE 02
In developmentSan Francisco, California

General Intelligence Labs

OpenArm

RobotReplica OpenArm

A new hosted benchmark for the open-source OpenArm platform, extending the network to larger, bimanual, contact-rich manipulation.

  • Bimanual platform
  • Open-source hardware
  • Protocol in development

HOW IT WORKS

From your model
to a verified score.

1

Find your robot

Choose a site that operates the same robot embodiment as your model or policy.

2

Contact the site

Share your policy, interface requirements, and the benchmark track you want to enter.

3

We run the evaluation

The host executes your policy on its maintained setup under a standardized protocol.

4

Compare the result

Your verified score is added to the leaderboard for that site and robot.

Evaluation details—policy interface, checkpoints, task coverage, and reporting—are coordinated directly with the selected host site.

LEADERBOARDS

Results grouped by
robot and site.

A score is meaningful only when the embodiment and physical protocol match. RobotReplica keeps those contexts explicit.

SO-101VLA-ReplicaIRVL @ UT DallasLIVEView leaderboard
OpenArmRobotReplica OpenArmGeneral Intelligence LabsIN DEVELOPMENTComing soon

EXPAND THE NETWORK

Operate a robot?
Host a site.

We are looking for organizations that can maintain a reproducible robot setup and evaluate community policies. Help bring a new embodiment into RobotReplica.

Propose a site