SmolVLA / paired closed-loop experiment
Change the view.
Check the policy.
Same task. Same starting state. Different light or camera position. Watch a fresh policy rollout in every condition, then inspect where its behavior changes.
A net success rate can hide changes in both directions. Select a group, then open any paired seed in the replay.
“Not completed” includes step limits and terminated runs; execution errors remain in the attempt ledger. Groups use the complete set of paired seeds. These are observed outcomes on one task, not a significance test or a general robustness score.
Check an exported report independently
With Robot Reel 0.8.0+ installed: robot-reel stress stress-lab --paired-report paired-outcomes.json. Replace stress-lab with your extracted experiment folder. The command checks the full experiment before comparing every group and count. From a source checkout, use python3 -m robot_reel.cli stress docs/stress --paired-report paired-outcomes.json.
Reference controls
Source trace ↓Selected controls
Source trace ↓Controls are normalized commands in [−1, 1], not joint angles. Each panel holds its own final observation when its run ends. A held sample is labelled explicitly.
Turn a moment into a review.
Export the selected pair with its recorded observations, controls and inference timing. Import the JSON into this lab to compare its facts and restore the selection. Your note stays separate from the recorded evidence.
Files are processed in this browser. Notes are included only in your downloads; they are not uploaded or saved between visits. Exports use a relative replay link and retain the complete experiment's counts.
To check a downloaded review against your local lab: python3 -m robot_reel.cli stress docs/stress --review review.json. Requires Robot Reel 0.6.0+ or the current checkout; replace the paths with your lab folder and downloaded JSON.
Small experiment. Complete record.
These are new closed-loop runs with native MuJoCo lighting and camera changes applied before the first policy observation. Initial qpos/qvel, renderer settings and inference noise hashes are checked across every pair.
Success is the environment's recorded predicate. “Step limit” means the task was not completed within this experiment's fixed budget. This is a single-task diagnostic, not an official LIBERO or LIBERO-plus benchmark score. Ten paired seeds do not establish general robustness.
Read the exact conditions, clocks and limitations ↗Take the evidence with you.
Every camera, applied action, inference timing, initial state and result. The offline folder embeds its telemetry and needs no service.
MCAP contains JSON observation and inference channels, with episode-relative nanosecond timestamps. Videos are separate MP4s. No model weights are bundled.
Inspect every recorded result and source file
| Trial | Outcome | Actions | Episode | Policy calls | Evidence |
|---|
Attempt history
| Trial / attempt | Status | Started UTC | Result or error |
|---|