Skip to content

Repository files navigation

rl_training

IsaacSim Isaac Lab RSL-RL Discord Python Linux platform License

Tutorial Videos

We've released the following tutorials for training and deploying a reinforcement learning policy. Please check it out on Bilibili or YouTube!

Overview

rl_training is a RL training library for deeprobotics robots, based on IsaacLab. The table below lists all available environments:

Robot Model Environment Name (ID) Screenshot
Deeprobotics Lite3 Rough-Deeprobotics-Lite3-v0 Lite3
Deeprobotics M20 Rough-Deeprobotics-M20-v0 deeprobotics_m20
Deeprobotics DR02 Amp-Flat-Deeprobotics-DR02-v0 deeprobotics_dr02

Note

If you want to deploy policies in mujoco or real robots, please use the corresponding deploy repo in Deep Robotics Github Center.

Contribution

Everyone is welcome to contribute to this repo. If you discover a bug or optimize our training config, just submit a pull request and we will look into it.

Installation

  • Install Isaac Sim 5.1.0 and Isaac Lab 2.3.2 with Python 3.11 using the installation guide. These are the versions tested with this repository; select these versions when following the guide. The commands below assume Linux and Bash with an activated Isaac Lab Conda environment.

  • Clone this repository separately from the Isaac Lab installation (i.e. outside the IsaacLab directory):

    git clone --recurse-submodules https://github.com/DeepRoboticsLab/rl_training.git
    cd rl_training
    # Also repairs an existing clone made without --recurse-submodules.
    git submodule update --init --recursive
  • Activate your Isaac Lab environment (replace deep-robotics-humanoid with your environment name), then install the library. This also installs RSL-RL 5.0.1 and PyBullet 3.2.7, required by DR02 AMP training.

    conda activate deep-robotics-humanoid
    python -m pip install -e source/rl_training
  • Configure the environment's C++ runtime before launching Isaac Sim:

    python scripts/tools/setup_conda_runtime.py
    conda deactivate
    conda activate deep-robotics-humanoid

    The script checks that Conda's libstdc++.so.6 provides CXXABI_1.3.15 and installs activation/deactivation hooks in that environment. This prevents Isaac Sim from loading an older system runtime that breaks Conda's ICU/SQLite imports. It preserves any existing LD_PRELOAD value and restores it on deactivation. Re-running the script is safe. If it reports an outdated or missing runtime, run conda install -c conda-forge "libstdcxx-ng>=15", then retry the script. System libraries are unchanged.

  • Verify that the extension is correctly installed by running the following command to print all the available environments in the extension:

    python scripts/tools/list_envs.py
  • Verify DR02 with a short training run before starting a full run:

    python scripts/reinforcement_learning/rsl_rl/train.py \
      --task=Amp-Flat-Deeprobotics-DR02-v0 --headless --num_envs=16 --max_iterations=1

    It should complete one learning iteration and save model_0.pt under logs/rsl_rl/dr02_amp/<run>/.

Setup as Omniverse Extension (Optional, click to expand)

We provide an example UI extension that will load upon enabling your extension defined in source/rl_training/rl_training/ui_extension_example.py.

To enable your extension, follow these steps:

  1. Add the search path of your repository to the extension manager:

    • Navigate to the extension manager using Window -> Extensions.
    • Click on the Hamburger Icon (☰), then go to Settings.
    • In the Extension Search Paths, enter the absolute path to rl_trainingb/source
    • If not already present, in the Extension Search Paths, enter the path that leads to Isaac Lab's extension directory directory (IsaacLab/source)
    • Click on the Hamburger Icon (☰), then click Refresh.
  2. Search and enable your extension:

    • Find your extension under the Third Party category.
    • Toggle it to enable your extension.

Try examples

Deeprobotics Lite3:

# Train
python scripts/reinforcement_learning/rsl_rl/train.py --task=Rough-Deeprobotics-Lite3-v0 --headless

# Play
python scripts/reinforcement_learning/rsl_rl/play.py --task=Rough-Deeprobotics-Lite3-v0 --num_envs=10

Deeprobotics M20:

# Train
python scripts/reinforcement_learning/rsl_rl/train.py --task=Rough-Deeprobotics-M20-v0 --headless

# Play
python scripts/reinforcement_learning/rsl_rl/play.py --task=Rough-Deeprobotics-M20-v0 --num_envs=10

Deeprobotics DR02:

# Train
python scripts/reinforcement_learning/rsl_rl/train.py --task=Amp-Flat-Deeprobotics-DR02-v0 --headless

# Play
python scripts/reinforcement_learning/rsl_rl/play.py --task=Amp-Flat-Deeprobotics-DR02-v0 --num_envs=10

Note

If you want to control a SINGLE ROBOT with the keyboard during playback, add --keyboard at the end of the play script.

Key bindings:
====================== ========================= ========================
Command                Key (+ve axis)            Key (-ve axis)
====================== ========================= ========================
Move along x-axis      Numpad 8 / Arrow Up       Numpad 2 / Arrow Down
Move along y-axis      Numpad 4 / Arrow Right    Numpad 6 / Arrow Left
Rotate along z-axis    Numpad 7 / Z              Numpad 9 / X
====================== ========================= ========================
  • Record video of a trained agent (requires installing ffmpeg), add --video --video_length 200
  • Play/Train with 32 environments, add --num_envs 32
  • Play on specific folder or checkpoint, add --load_run run_folder_name --checkpoint model.pt
  • Resume training from folder or checkpoint, add --resume --load_run run_folder_name --checkpoint model.pt

Multi-gpu acceleration

  • To train with multiple GPUs, use the following command, where --nproc_per_node represents the number of available GPUs:

    python -m torch.distributed.run --nnodes=1 --nproc_per_node=2 scripts/reinforcement_learning/rsl_rl/train.py --task=<ENV_NAME> --headless 
    python -m torch.distributed.run --nnodes=1 --nproc_per_node=2 scripts/reinforcement_learning/rsl_rl/train.py --task=Rough-Deeprobotics-Lite3-v0 --headless --distributed --num_envs=2048
  • Note: each gpu will have the same number of envs specified in the config, to use the previous total number of envs, devide it by the number of gpus.

  • To scale up training beyond multiple GPUs on a single machine, it is also possible to train across multiple nodes. To train across multiple nodes/machines, it is required to launch an individual process on each node.

    For the master node, use the following command, where --nproc_per_node represents the number of available GPUs, and --nnodes represents the number of nodes:

    python -m torch.distributed.run --nproc_per_node=2 --nnodes=2 --node_rank=0 --rdzv_id=123 --rdzv_backend=c10d --rdzv_endpoint=localhost:5555 scripts/reinforcement_learning/rsl_rl/train.py --task=<ENV_NAME> --headless --distributed

    Note that the port (5555) can be replaced with any other available port. For non-master nodes, use the following command, replacing --node_rank with the index of each machine:

    python -m torch.distributed.run --nproc_per_node=2 --nnodes=2 --node_rank=1 --rdzv_id=123 --rdzv_backend=c10d --rdzv_endpoint=ip_of_master_machine:5555 scripts/reinforcement_learning/rsl_rl/train.py --task=<ENV_NAME> --headless --distributed

Tensorboard

To view tensorboard, run:

tensorboard --logdir=logs

Export Policy to ONNX (without Isaac Sim)

Export a trained checkpoint to ONNX directly from the .pt file — no Isaac Sim or environment setup required:

# Lite3
python scripts/tools/export_onnx_fast.py \
    --checkpoint_path logs/rsl_rl/deeprobotics_lite3_rough/<run>/model_5000.pt \
    --robot lite3 \
    --output_path exported/lite3_policy.onnx

# M20
python scripts/tools/export_onnx_fast.py \
    --checkpoint_path logs/rsl_rl/deeprobotics_m20_rough/<run>/model_5000.pt \
    --robot m20 \
    --output_path exported/m20_policy.onnx

Robot metadata (joint names, stiffness/damping, default positions, action scales) is embedded in the ONNX file as model properties. Add --no_metadata to skip this.

Compare Training Runs

Diff the agent.yaml and env.yaml configs between two runs (saved automatically to params/ by train.py):

python scripts/tools/compare_runs.py \
    logs/rsl_rl/deeprobotics_lite3_rough/<run1> \
    logs/rsl_rl/deeprobotics_lite3_rough/<run2>

Troubleshooting

DR02 startup errors

  • CXXABI_1.3.15 not found, omni.kit has no attribute test, or cannot import name tests: the latter errors can cascade from the C++ runtime failure. Run python scripts/tools/setup_conda_runtime.py in your activated Conda environment, then deactivate/reactivate it and restart training. An already running Python process cannot pick up the new runtime.
  • Missing DR02-pro.urdf or DR02-pro_fix_joints.urdf: run git submodule update --init --recursive from the repository root.
  • Missing pybullet_utils or incompatible RSL-RL: re-run python -m pip install -e source/rl_training in the environment used for training. The package declares the tested versions. The training scripts use the repository's config converter for RSL-RL 5, so the missing Isaac Lab handle_deprecated_rsl_rl_cfg helper is not required.

Some startup warnings remain with the tested setup. A missing viewport is expected in headless mode. The DR02 URDF contains fixed sensor links without inertia; the importer assigns small inertias and adjusts joint axes. These messages do not prevent training, but changing the robot's inertial properties should be based on measured model data. CPU powersave and GPU peer-to-peer messages concern machine performance, not Python imports.

Pylance Missing Indexing of Extensions

In some VsCode versions, the indexing of part of the extensions is missing. In this case, add the path to your extension in .vscode/settings.json under the key "python.analysis.extraPaths".

Note: Replace <path-to-isaac-lab> with your own IsaacLab path.

{
    "python.languageServer": "Pylance",
    "python.analysis.extraPaths": [
        "${workspaceFolder}/source/rl_training",
        "/<path-to-isaac-lab>/source/isaaclab",
        "/<path-to-isaac-lab>/source/isaaclab_assets",
        "/<path-to-isaac-lab>/source/isaaclab_mimic",
        "/<path-to-isaac-lab>/source/isaaclab_rl",
        "/<path-to-isaac-lab>/source/isaaclab_tasks",
    ]
}

Clean USD Caches

Temporary USD files are generated in /tmp/IsaacLab/usd_{date}_{time}_{random} during simulation runs. These files can consume significant disk space and can be cleaned by:

rm -rf /tmp/IsaacLab/usd_*

Acknowledgements

The project uses some code from the following open-source code repositories:

About

RL_Training Repo Based on Isaaclab

Resources

Stars

271 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages