Proof · benchmark

The trajectories come out 8–37× smoother

On DROID, FMB and RH20T (public robot datasets with encoder-grade ground truth), motion reconstructed from ordinary monocular video carries far less jerk than Encord's 2D-keypoint annotation on every moving-camera task tested, at p < 1e-6. (The off-the-shelf splatting arm produces no trajectory to compare; fixed-camera cells show a far smaller margin.) Smoothness is the property that decides how well a policy trains.

  • H3 · data qualityDecisive win8–37× smoother trajectories, every task, p < 1e-6
  • H1 · accuracyCell-dependentwins moving-camera pick-place, ties insertion, loses close-range stacking
  • Arm C · SplaticaVisualization-gradeno trajectory at all, ghosts the moving object, ≈4× the cost
  • H2 · absolute scaleHonest negativeno method hits usable absolute metres on this data - one reference fixes it
Median action jerk, lower is smootherOpenRealityEncord
Pick-and-placeDROID · wrist cam
108.9
13.2
8.2×smoother
InsertionFMB · wrist cam
269.6
11.7
23×smoother
StackingRH20T · in-hand
104.5
2.85
37×smoother

m/s³ · median over episodes

Pooled paired test · 69 episodes

Action jerk −106.3 m/s³ (95% CI [−157.8, −66.5]), velocity noise −30.6 m/s², both p < 1e-6. And every arm stays above the encoder-truth smoothness floor: 0 / 24 cells over-smoothed, so it recovers real motion, it doesn’t blur it away.

Position accuracy is cell-dependent. We report the losses too

End-effector error on moving-camera capture. We win where it counts for monocular capture, tie on insertion, and lose close-range contact-rich stacking. Pooled across every moving-camera cell it’s a statistical tie (−2.6 mm, p = 0.157). The baselines were tuned as hard as our own pipeline.

  • winPick-and-placeDROID · moving camera34.9 mmvs≈79 mm≈2.2× tighter · paired −18.5 mm, p ≈ 0.01
  • tieInsertionFMB · moving camera18.8 mmvs18.6 mmstatistical tie · p ≈ 0.4
  • lossStackingRH20T · in-hand, contact-rich19.5 mmvs9.6 mmbaseline wins close range · p ≈ 0.07–0.10

Splatica's off-the-shelf 3D Gaussian splatting costs ≈4× more per clip ($0.31 vs $0.075), produces no trajectory at all, and ghosts the moving object (visualization-grade, not training-grade). Ground truth: DROID · RH20T · FMB (public robot datasets). No absolute-metric claim. OpenReality ships scale-normalized metres.

Read the full benchmark →