📖 This is the concise version (~3 min). For the full engineering details (design decisions, algorithm / reward, diagnostics, reproduce commands), read the deep-dive →

From “can record” to “can train & eval”

The intro post wired SO-101, MuJoCo, and LeRobot on AMD ROCm so you can teleoperate in sim and collect expert trajectories.

v0.1.3 (release-v0.1.3) closes the next loop — Lab 01 pick-and-place: sim demos → ACT / SmolVLA training → MuJoCo closed-loop eval, plus downloadable Hub reference assets.

Demos work with keyboard, Joy-Con, or leader on the same LeRobot v3.0 record path. The published Lab 01 reference set was collected with a leader arm; you can run the same lab scripts with keyboard or Joy-Con locally.

What imitation learning is doing here

Imitation learning (behavior cloning) does not hand-code a grasp planner. A policy learns observation → action from expert demos: a human teleoperates successful pick-and-place episodes; the network fits those actions given joint state and camera frames; then you run closed-loop rollouts in sim and score success.

Unlike RL, there is no reward function to tune — data quality and coverage set the ceiling. Lab 01 packages the usual industry starter loop in three steps:

StageRole
1. Data collectionTeleop expert trajectories → LeRobot v3.0 dataset (state + images + actions)
2. Policy trainingBehavior cloning: match predicted actions to demos (ACT / SmolVLA)
3. EvalPolicy drives the arm in MuJoCo closed loop; count success / failure
teleop demos  →  dataset  →  train (BC)  →  closed-loop eval

One Lab 01 line

StageWhat
Recordlabs/lab01_pnp/record.cmd (or your teleop config)
Traintrain.cmd (SmolVLA) / train_act.cmd (ACT)
EvalSingle eval.cmd + lab YAMLs (full-range / demo / fixed pose)
HubDataset + ACT / SmolVLA reference weights (override to your HF id)

Knobs live in _env.sh (LAB01_* overrides). Shared lab conventions: labs/README.md.

Reference results (same 50-episode set)

On MI300X 50K checkpoints with full-range random cube spawn (details in the deep-dive):

PolicySuccess (indicative)
ACT32/50 (64%)
SmolVLA11/50 (22%)

Fixed-pose demos raise ACT further; SmolVLA stays brittle — “trainable” ≠ “classroom-stable.” Success criteria and failure modes are in the lab runbook §6.

Eval footage (ACT fixed pose)

Closed-loop MuJoCo GUI clip (trimmed, ~1.5Ă— speed; full mp4 under assets/videos/):

ACT pick-and-place eval in MuJoCo

ACT 50K training loss (MI300X)

Where to start

git clone --recursive https://github.com/rocPAI-Forge/so101-simstudio.git
cd so101-simstudio
git checkout release-v0.1.3
make rocm-sync && source .venv-rocm/bin/activate

# Read the runbook, or download Hub assets and eval (Lab 01 §7)
./labs/lab01_pnp/eval.cmd

đź“– Want more? Full engineering details in the deep-dive.