Towards Embodied AI with MuscleMimic: Unlocking full-body musculoskeletal motor learning at scale
Chengkun Li, Cheryl Wang, Bianca Ziliotto, Merkourios Simos, Jozsef Kovecses, Guillaume Durandau, Alexander Mathis
CVBW @ CVPR 2026 (Workshop Accepted Paper)
arXiv:2603.25544 (Mar. 2026)
MyoFullBody: 416 muscles · 123 joints · 17 mimic sites
1
BACKGROUND
Why simulate with muscles?
Fig 1. MyoBimanualArm (top) · MyoFullBody (bottom) — front/back/side
2
PROBLEM
Two bottlenecks in prior work
Bottleneck 1 — Validation gap
Bottleneck 2 — Compute cost
10 days → 20 hours (1B steps, single H100)
3
OVERVIEW
MuscleMimic at a glance
1
Two validated bodies
MyoBimanualArm (126 muscles) ·
MyoFullBody (416), full collisions
2
Retargeting pipeline
SMPL mocap → MSK morphology
(GMR-Fit, constraint-preserving)
3
GPU-native training
MuJoCo Warp + JAX,
8,192 parallel environments
4
Algorithmic finding
Single-epoch PPO beats
multi-epoch for muscles
5
Biomech validation
Kinematics · kinetics · GRF · EMG
vs. human experimental data
13,598
sim steps/sec (8,192 envs, 1× H100)
92.6 / 99.5%
test success (full-body / bimanual)
r = 0.90
walking kinematics vs. human data
4
METHOD 1/4
Two embodiments — model specs
| MyoBimanualArm | MyoFullBody |
Base | Fixed thorax + both arms | Free-root full body (toes to fingertips) |
Joints / DoF | 76 / 54 | 123 / 72 |
Muscles | 126 (64 w/o fingers) | 416 (354 w/o fingers) |
Mimic sites | 6 + thorax reference | 17 (full-body landmarks) |
Mass | 9.3 kg (both arms) | 84.3 kg (adult human) |
Collisions | arm–thorax, arm–arm | environment + full self-collision |
Fig 9. 17 mimic sites — shared anchors for retargeting, reward, termination
5
METHOD 1/4
Muscle actuators — activation dynamics
∂act/∂t = (ctrl − act) / τ(ctrl, act)
τ = τact · (0.5 + 1.5·act) (activation, ctrl > act)
τ = τdeact / (0.5 + 1.5·act) (deactivation, ctrl ≤ act)
Why this equation matters
The delay & nonlinearity between command and effect is the cause of the E = 1 finding coming up.
6
METHOD 2/4
Motion dataset & retargeting
Fig 10. Pre-process (SMPL fit · scaling) → IK → post-process (interp · artifact fixes)
Retargeting quality (KINESIS 972, Tab. 4)
Joint-limit violations 12.26% → 0.27%
Tendon-jump rate 30.14% → 3.20%
RMSE 0.039 → 0.025 m
Speed (per frame) 0.076 → 0.251 s (Mocap-Body 3× faster)
7
METHOD 3/4
GPU-native training system
Fig 2. Parallel envs 16 → 8,192: throughput 174 → 13,598 SPS (~78×)
8
METHOD 4/4
Observations · actions · reward
Fig 11. State history + current & future goals (0.2 s × 5) + previous action
9
KEY FINDING
For muscles, re-use hurts — single-epoch PPO
Fig 3. (A) E=10 learns fastest early (B) E=1 wins asymptotically; E=3 collapses at ~2.7B (C) KL: E>1 explodes to 10¹⁰, E=1 stays below 10⁻¹
Cause: delayed, nonlinear activation dynamics (Eq. 1) amplify small policy changes → extreme off-policy sensitivity. High GPU throughput compensates for one-gradient-per-sample, making truly on-policy training feasible. (+ batch 128 more stable than 32, Fig 4)
10
ARCHITECTURE
Gated residual policy — depth, gradually
Fig 12. Normalize → [Dense–LN–SiLU ×2 + gated residual] × N blocks → action
11
RESULTS 1/3
Imitation — one policy, hundreds of motions
92.62%
MyoFullBody test success (108 unseen motions)
99.46%
MyoBimanualArm test success (312)
6.63°
joint angle error (full-body test)
2.28 cm
relative site position error
Fig 6. walk / run / turn / dance / jump / kick-twist
Fig 5. bimanual
12
RESULTS 2/3
Biomechanical validation — does it walk like us?
Fig 7a. Walking 1.2 m/s: left-leg joint angles, moments, GRF vs. two human datasets (9 subjects each)
13
RESULTS 3/3
EMG comparison — an honest limitation
Fig 8. Per-muscle activation patterns vs. human EMG — correlations 0.2–0.6, on par with static-optimization baseline
Why not a perfect match?
MSK redundancy: many muscle combinations produce the same joint trajectory → a policy imitating kinematics may find a valid but non-human strategy
14
DISCUSSION
Limitations & takeaways
Limitations (stated by authors)
Takeaways for our lab