
Adversarial scene variation
Perturb object pose and contact conditions across parallel task instances.
Loading...
BreakBot probes manipulation policies against physics shifts, control degradation and randomized edge cases—then returns measurable evidence before hardware is at risk.
32 seeds
rollouts per run
0—100
comparable score
≈75 sec
typical runtime
Franka
live task
The stress-test terminal
A production-shaped queue sends versioned jobs to a private RTX worker. Isaac Lab executes deterministic rollouts, renders the evidence, and returns metrics without exposing worker credentials.
Mass, friction, pose and actuator limits
Repeatable rollouts expose brittle behavior
Joint position, absolute IK and relative IK
Robustness, collisions, timing and replay
Native NVIDIA Isaac Lab worker on private compute
Shareable H.264 output attached to every run
Normal friction. Full actuator authority. Low latency. The Franka policy secures the cube, clears the table and completes the trajectory across the expected operating range.
MASS
0.35 kg
MOTOR
92%
LATENCY
18 ms
Increase payload, cut friction, weaken the motors or delay control. BreakBot maps the boundary where a safe-looking policy becomes a hardware incident.
Find the breaking pointExtensible test coverage
BreakBot currently exposes one small model for public tests. The worker contract is designed for richer manipulation benchmarks when more compute is available.

Perturb object pose and contact conditions across parallel task instances.

Probe grasp tolerance, small-object control and placement sensitivity.

Turn randomized failures into a reproducible robustness surface.

Compare joint-space and inverse-kinematics behavior near operating limits.

A worker-ready contract for additional small manipulation models.
Public tests run on privately operated GPU capacity and may wait in a queue. No model upload is required or accepted in this version.
The intelligence assurance layer
Autonomy is moving from controlled demos into factories, hospitals and infrastructure. BreakBot is building the verification terminal between a trained policy and an expensive machine: configure the stress, execute at scale, preserve the evidence, ship with confidence.
Attack the operating envelope.
Quantify robustness across seeds.
Inspect physics-derived evidence.
Turn failures into better policies.
Today: a public Franka benchmark on private RTX compute. Next: more policies, more task families and deeper adversarial evaluation—without changing the submission or evidence architecture.