📖 This is the concise version (~3 min). For the full engineering details (design decisions, algorithm / reward, diagnostics, reproduce commands), read the deep-dive →
From “can record” to “can train & eval”
The intro post wired SO-101, MuJoCo, and LeRobot on AMD ROCm so you can teleoperate in sim and collect expert trajectories.
v0.1.3 (release-v0.1.3) closes the next loop — Lab 01 pick-and-place: sim demos → ACT / SmolVLA training → MuJoCo closed-loop eval, plus downloadable Hub reference assets.
Demos work with keyboard, Joy-Con, or leader on the same LeRobot v3.0 record path. The published Lab 01 reference set was collected with a leader arm; you can run the same lab scripts with keyboard or Joy-Con locally.
What imitation learning is doing here
Imitation learning (behavior cloning) does not hand-code a grasp planner. A policy learns observation → action from expert demos: a human teleoperates successful pick-and-place episodes; the network fits those actions given joint state and camera frames; then you run closed-loop rollouts in sim and score success.
Unlike RL, there is no reward function to tune — data quality and coverage set the ceiling. Lab 01 packages the usual industry starter loop in three steps:
| Stage | Role |
|---|---|
| 1. Data collection | Teleop expert trajectories → LeRobot v3.0 dataset (state + images + actions) |
| 2. Policy training | Behavior cloning: match predicted actions to demos (ACT / SmolVLA) |
| 3. Eval | Policy drives the arm in MuJoCo closed loop; count success / failure |
teleop demos → dataset → train (BC) → closed-loop eval
One Lab 01 line
| Stage | What |
|---|---|
| Record | labs/lab01_pnp/record.cmd (or your teleop config) |
| Train | train.cmd (SmolVLA) / train_act.cmd (ACT) |
| Eval | Single eval.cmd + lab YAMLs (full-range / demo / fixed pose) |
| Hub | Dataset + ACT / SmolVLA reference weights (override to your HF id) |
Knobs live in _env.sh (LAB01_* overrides). Shared lab conventions: labs/README.md.
Reference results (same 50-episode set)
On MI300X 50K checkpoints with full-range random cube spawn (details in the deep-dive):
| Policy | Success (indicative) |
|---|---|
| ACT | 32/50 (64%) |
| SmolVLA | 11/50 (22%) |
Fixed-pose demos raise ACT further; SmolVLA stays brittle — “trainable” ≠“classroom-stable.” Success criteria and failure modes are in the lab runbook §6.
Eval footage (ACT fixed pose)
Closed-loop MuJoCo GUI clip (trimmed, ~1.5Ă— speed; full mp4 under assets/videos/):
ACT pick-and-place eval in MuJoCo

Where to start
git clone --recursive https://github.com/rocPAI-Forge/so101-simstudio.git
cd so101-simstudio
git checkout release-v0.1.3
make rocm-sync && source .venv-rocm/bin/activate
# Read the runbook, or download Hub assets and eval (Lab 01 §7)
./labs/lab01_pnp/eval.cmd
- Code / Release: rocPAI-Forge/so101-simstudio · v0.1.3
- Lab runbook: labs/lab01_pnp/lab01_pnp.md
- Deep-dive: README-details.md
đź“– Want more? Full engineering details in the deep-dive.