9.3 KiB
Go2 RL GYM
🌎 English | 🇨🇳 中文
This repository builds on unitree_rl_gym to train the Unitree Go2 quadruped with reinforcement learning.
📦 Installation
Follow the step-by-step setup guide in setup.md.
🛠️ Usage Guide
1. Train
Run the following command to launch training:
python legged_gym/scripts/train.py --task=xxx --headless
⚙️ Arguments
--task: Required. Options includego2,go2_cts,go2_moe_cts,go2_moe_ng_cts,go2_mcp_cts,go2_ac_moe_cts,go2_dual_moe_cts;go2_moe_ctsis the paper's final version.--headless: Render viewer by default; set totrueto disable rendering for higher throughput.--resume: Resume training from a chosen checkpoint in the logs.--experiment_name: Experiment folder to save/load from.--run_name: Run subfolder name to save/load from.--load_run: Name of the run to load (defaults to the most recent run).--checkpoint: Checkpoint index to load (defaults to the latest file).--num_envs: Number of parallel simulated environments.--seed: Random seed.--max_iterations: Maximum training iterations.--sim_device: Physics simulation device. Use--sim_device=cputo force CPU.--rl_device: RL computation device. Use--rl_device=cputo force CPU.--robogauge: Enable RoboGauge evaluation tool; disabled by default. Evaluation results are saved asresults_{it}.yamlinlogs/{exp_name}/{date}/robogauge_resultsand logged to TensorBoard.--robogauge_port: RoboGauge server port; default is 9973.
RoboGauge evaluation requires a separate server to be started. Refer to the RoboGauge documentation.
Default checkpoint path: logs/<experiment_name>/<date_time>_<run_name>/model_<iteration>.pt
Model Evaluation
The trained model above was evaluated using the RoboGauge framework via Sim2Sim. The models in the table below are the best models after 150k training steps.
| Model | Score | Tracking | Safety | Quality | Level | Download |
|---|---|---|---|---|---|---|
| go2_moe_cts (Ours) | 0.6713 | 0.6669 | 0.7857 | 0.7392 | 7.85 | ckpt |
| go2_ac_moe_cts | 0.6509 | 0.6442 | 0.7644 | 0.7149 | 7.52 | ckpt |
| go2_mcp_cts | 0.6399 | 0.6355 | 0.7542 | 0.7058 | 7.41 | ckpt |
| go2_moe_ng_cts | 0.6519 | 0.6447 | 0.7639 | 0.7186 | 7.56 | ckpt |
| CTS vanilla | 0.5786 | 0.5755 | 0.7066 | 0.6624 | 6.83 | ckpt |
| HIM | 0.5379 | 0.5453 | 0.6476 | 0.6050 | 6.19 | ckpt |
| DreamWaQ | 0.5054 | 0.5105 | 0.6149 | 0.5730 | 5.74 | ckpt |
In the downloaded ckpt files,
*.ptis used for Python deployment, and*.onnxis used for C++ deployment. The models above were all trained with self-collision disabled. In later tests, we found that enabling self-collision can also achieve strong results; see go2_moe_cts_self_0.6669 - ckpt.
2. Play
Visualize policies inside Gym with:
python legged_gym/scripts/play.py --task=xxx
Notes
- Play launches on randomized terrain with difficulty between 7 and 9.
- It automatically loads the latest checkpoint inside the experiment folder.
- You can specify another model via
experiment_name,load_run, andcheckpoint, for example:python legged_gym/scripts/play.py --task=go2_moe_cts --num_envs 100 --experiment_name go2_cts_hard_terrain --load_run Mar21_22-54-5-46_ --checkpoint 100000
💾 Policy Export
Play exports the Actor network to logs/{experiment_name}/exported/policies:
policy.pt: TorchScript model for Sim2Sim.policy.onnx: ONNX model for Sim2Real.policy.pkl: Raw weights.
Demonstration
3. Sim2Sim (Mujoco)
Run policies in the Mujoco simulator:
python deploy/deploy_mujoco/deploy_go2.py
Connect an Xbox-compatible gamepad to enable teleoperation; otherwise, the agent keeps a default forward command.
- Swap the policy: The default checkpoint is
deploy/pre_train/go2/go2_cts_150k.pt. Replacepolicy_pathin the YAML config with your ownlogs/{experiment_name}/exported/policies/policy.pt. - Swap terrains: Default terrain is
resources/robots/go2/stairs.xml. Alternatives includeflat.xml,race_track.xml,cross_stairs.xml, andcross_slope.xml. Generate new terrains with windigal - mujoco_terrains
Results
| Flat | Stairs | Race Track |
|---|---|---|
![]() |
![]() |
![]() |
4. Sim2Real
4.1 Python Deployment
# Onboard Jetson: pick Python by JetPack version
# JetPack 6: Python 3.10
# JetPack 5: Python 3.8
conda create -n deploy python=3.10
conda activate deploy
# Install the matching PyTorch wheel for your Jetson
# https://forums.developer.nvidia.com/t/pytorch-for-jetson/72048
git clone https://github.com/unitreerobotics/unitree_sdk2_python.git
cd unitree_sdk2_python
pip3 install -e .
In the Unitree app, open Device → Service, disable mcf/*, and enable the ota_box service.
Assuming the interface to the low-level controller is eth0:
cd deploy/deploy_real
python deploy_real_go2.py eth0
Press start to stand and A to engage the controller.
4.2 C++ Deployment
Follow the usage described in unitree_cpp_deploy.
Demonstration
| Python Deploy | C++ Deploy |
|---|---|
![]() |
![]() |
🎉 Acknowledgements
This repository would not exist without the following open-source projects:
- unitree_rl_gym: Unitree's core RL training framework.
- legged_gym: Base locomotion environment.
- rsl_rl: Reinforcement learning algorithms.
- mujoco: High-performance CPU physics simulator.
- unitree_sdk2_python: Python hardware interface for deployment.
- unitree_sdk2: C++ hardware interface for deployment.
Related publications implemented in this repo:
Contributors:
- @windigal: CTS algorithm reproduction, terrain generation, video editing
- @wertyuilife2: CTS algorithm reproduction
📄 Citation
If you find our work helpful, please cite:
@article{wu2026robogauge,
title={Toward Reliable Sim-to-Real Predictability for MoE-based Robust Quadrupedal Locomotion},
author={Tianyang Wu and Hanwei Guo and Yuhang Wang and Junshu Yang and Xinyang Sui and Jiayi Xie and Xingyu Chen and Zeyang Liu and Xuguang Lan},
year={2026},
journal={arXiv preprint arXiv:2602.00678},
url={https://arxiv.org/abs/2602.00678},
}
🔖 License
New contributions follow the MIT License; the original unitree_rl_gym remains under the BSD 3-Clause License.
See the complete LICENSE file for details.







