Go2 RL GYM
🌎 English | 🇨🇳 中文 | 📄 Paper [RSS 2026]
This repository builds on unitree_rl_gym to train the Unitree Go2 quadruped with reinforcement learning.
For the IsaacLab-based version, see go2_rl_robotlab.
📦 Installation
Follow the step-by-step setup guide in setup.md.
🛠️ Usage Guide
1. Train
Run the following command to launch training:
python legged_gym/scripts/train.py --task=xxx --headless
⚙️ Arguments
--task: Required. Options includego2,go2_cts,go2_moe_cts,go2_moe_ng_cts,go2_mcp_cts,go2_ac_moe_cts,go2_dual_moe_cts;go2_moe_ctsis the paper's final version.--headless: Render viewer by default; set totrueto disable rendering for higher throughput.--resume: Resume training from a chosen checkpoint in the logs.--experiment_name: Experiment folder to save/load from.--run_name: Run subfolder name to save/load from.--load_run: Name of the run to load (defaults to the most recent run).--checkpoint: Checkpoint index to load (defaults to the latest file).--num_envs: Number of parallel simulated environments.--seed: Random seed.--max_iterations: Maximum training iterations.--sim_device: Physics simulation device. Use--sim_device=cputo force CPU.--rl_device: RL computation device. Use--rl_device=cputo force CPU.--robogauge: Enable RoboGauge evaluation tool; disabled by default. Evaluation results are saved asresults_{it}.yamlinlogs/{exp_name}/{date}/robogauge_resultsand logged to TensorBoard.--robogauge_port: RoboGauge server port; default is 9973.
RoboGauge evaluation requires a separate server to be started. Refer to the RoboGauge documentation.
Default checkpoint path: logs/<experiment_name>/<date_time>_<run_name>/model_<iteration>.pt
Model Evaluation
The trained model above was evaluated using the RoboGauge framework via Sim2Sim. The models in the table below are the best models after 150k training steps. All released checkpoints are hosted on Hugging Face: wty-yy/go2_rl_gym_data.
| Model | Score | Tracking | Safety | Quality | Level | Download |
|---|---|---|---|---|---|---|
| go2_moe_cts (Ours) | 0.6713 | 0.6669 | 0.7857 | 0.7392 | 7.85 | ckpt |
| go2_ac_moe_cts | 0.6509 | 0.6442 | 0.7644 | 0.7149 | 7.52 | ckpt |
| go2_mcp_cts | 0.6399 | 0.6355 | 0.7542 | 0.7058 | 7.41 | ckpt |
| go2_moe_ng_cts | 0.6519 | 0.6447 | 0.7639 | 0.7186 | 7.56 | ckpt |
| CTS vanilla | 0.5786 | 0.5755 | 0.7066 | 0.6624 | 6.83 | ckpt |
| HIM | 0.5379 | 0.5453 | 0.6476 | 0.6050 | 6.19 | ckpt |
| DreamWaQ | 0.5054 | 0.5105 | 0.6149 | 0.5730 | 5.74 | ckpt |
In the downloaded ckpt files,
*.ptis used for Python deployment, and*.onnxis used for C++ deployment. The models above were all trained with self-collision disabled. In later tests, we found that enabling self-collision can also achieve strong results; see go2_moe_cts_164k_0.6715 - exported with complete model weights - model_164000.pt.
2. Play
Visualize policies inside Gym with:
python legged_gym/scripts/play.py --task=xxx
Notes
- Play launches on randomized terrain with difficulty between 7 and 9.
- It automatically loads the latest checkpoint inside the experiment folder.
- You can specify another model via
experiment_name,load_run, andcheckpoint, for example:python legged_gym/scripts/play.py --task=go2_moe_cts --num_envs 100 --experiment_name go2_cts_hard_terrain --load_run Mar21_22-54-5-46_ --checkpoint 100000
💾 Policy Export
Play exports the Actor network to logs/{experiment_name}/exported/policies:
policy.pt: TorchScript model for Sim2Sim.policy.onnx: ONNX model for Sim2Real.policy.pkl: Raw weights.
Demonstration
3. Sim2Sim (Mujoco)
Run policies in the Mujoco simulator:
python deploy/deploy_mujoco/deploy_go2.py
Connect an Xbox-compatible gamepad to enable teleoperation; otherwise, the agent keeps a default forward command.
- Swap the policy: The default checkpoint is
deploy/pre_train/go2/go2_cts_150k.pt. Replacepolicy_pathin the YAML config with your ownlogs/{experiment_name}/exported/policies/policy.pt. - Swap terrains: Default terrain is
resources/robots/go2/stairs.xml. Alternatives includeflat.xml,race_track.xml,cross_stairs.xml, andcross_slope.xml. Generate new terrains with windigal - mujoco_terrains
Results
| Flat | Stairs | Race Track |
|---|---|---|
![]() |
![]() |
![]() |
4. Sim2Real
4.1 Python Deployment
# Onboard Jetson: pick Python by JetPack version
# JetPack 6: Python 3.10
# JetPack 5: Python 3.8
conda create -n deploy python=3.10
conda activate deploy
# Install the matching PyTorch wheel for your Jetson
# https://forums.developer.nvidia.com/t/pytorch-for-jetson/72048
git clone https://github.com/unitreerobotics/unitree_sdk2_python.git
cd unitree_sdk2_python
pip3 install -e .
In the Unitree app, open Device → Service, disable mcf/*, and enable the ota_box service.
Assuming the interface to the low-level controller is eth0:
cd deploy/deploy_real
python deploy_real_go2.py eth0
Press start to stand and A to engage the controller.
4.2 C++ Deployment
Follow the usage described in unitree_cpp_deploy.
Demonstration
| Python Deploy | C++ Deploy |
|---|---|
![]() |
![]() |
C++ Deployment: Policy 1/2/4 trained by go2_rl_gym, Policy 3 trained by go2_rl_robotlab.
https://github.com/user-attachments/assets/b72e10f2-ffdb-407d-bb1f-9d545e7f9f63
🎉 Acknowledgements
This repository would not exist without the following open-source projects:
- unitree_rl_gym: Unitree's core RL training framework.
- legged_gym: Base locomotion environment.
- rsl_rl: Reinforcement learning algorithms.
- mujoco: High-performance CPU physics simulator.
- unitree_sdk2_python: Python hardware interface for deployment.
- unitree_sdk2: C++ hardware interface for deployment.
Related publications implemented in this repo:
Contributors:
- @windigal: CTS algorithm reproduction, terrain generation, video editing
- @wertyuilife2: CTS algorithm reproduction
📄 Citation
If you find our work helpful, please cite:
@inproceedings{wu2026robogauge,
title={Toward Reliable Sim-to-Real Predictability for MoE-based Robust Quadrupedal Locomotion},
author={Tianyang Wu and Hanwei Guo and Yuhang Wang and Junshu Yang and Xinyang Sui and Jiayi Xie and Xingyu Chen and Zeyang Liu and Xuguang Lan},
booktitle={Proceedings of Robotics: Science and Systems},
year={2026}
}
🔖 License
New contributions follow the MIT License; the original unitree_rl_gym remains under the BSD 3-Clause License.
See the complete LICENSE file for details.







