Humanoid Robot Dancing and Running
1 September 2026
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan
Inclusive, Self-Reliant, Sustainable
Universitas Gadjah Mada Team
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
Running
Result
The Methode
The Controller
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Methode: Deep Reinforcement Learning
WHAT is Deep Reinforcement Learning?
WHY Deep Reinforcement Learning
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Architecture
Training hyperparameters
lr = 3×10⁻⁴, 4 epochs, 4 minibatches, gradient clip 0.5, horizon 64, 256 parallel environments.
Observation: 18 dimensions
Action: 6-dimensional residual
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
PPO Algorithm
Advantage is estimated with GAE(λ):
Clipped objective function:
Total Loss:
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Reward Design
Velocity tracking (exponential)
Flight-phase bonus
Energy penalty
Posture stability penalty
Total Reward per Steps:
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
Curriculum Learning & Training
Curriculum Learning
Promotion gate:
Training
PHASE 1: PRE-TRAINED WALK
PHASE 2: SPRINT CURICULLUM
Finish Time: 120s
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Controller
Feedback-based PID Control
Central Pattern Generator - Gait Generator
DCM Capture Step
PPO Algorithm
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
Dancing
The Methode
The Controller
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Result: Tutting Dance
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Result: Tari Kecak Bali
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Result: Ateez San at City vs ATM
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Methode: PoseSeq
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
The Methode: PoseSeq
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
DanceController: Motion Playback & Balance Control
Purpose
CSV Motion Trajectory ⟶ Time-Based Interpolation ⟶ Joint Position Reference ⟶ Balance Feedback ⟶ Corrected Joint Target ⟶ Humanoid Robot
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
CSV Loading & Joint Trajectory Interpolation
The controller finds two neighboring motion frames and calculates the intermediate joint position
The interpolated position is sent to each joint as its target position:
ioBody->joint(i)->q_target() = target;
Why Interpolation?
It prevents sudden jumps between motion frames and produces a smoother joint trajectory during simulation.
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
IMU-Based Balance Feedback
Sensors Used
The correction is limited to prevent excessive ankle movement and unstable feedback.
Tilit Estimation
pitchTilt = atan2(-a.x(), a.z());
rollTilt = atan2(a.y(), a.z());
PD Feedback
Correction=KP×Tilt+KD×AngularVelocity
Correction Limitation
maxCorr = 0.35 rad ≈ 20°
This is a simplified IMU-based ankle strategy, not a full ZMP preview controller. It is mainly intended to compensate for small balance disturbances
ugm.ac.id
Merakyat, Mandiri, Berkelanjutan | Inclusive, Self-Reliant, Sustainable
“THANKYOU”
“TERIMAKASIH”