Deep Reinforcement Learning for Custom Quadruped Locomotion
Proprioception-only locomotion policies transferred from simulation to a custom quadruped across flat and rough terrain.

- My contribution
- Locomotion policy development, Isaac Gym training, and simulation-to-robot transfer.
- Outcome
- Proprioception-only locomotion, trained in simulation and transferred to a custom quadruped.
Project overview
| Area | Details |
|---|---|
| Platform | Custom quadruped robot |
| Training | Parallel reinforcement learning in Isaac Gym |
| Observation | Proprioception only; no external terrain sensing |
I developed blind locomotion policies in simulation and transferred them to a custom quadruped. The project focused on producing stable behavior from proprioceptive observations while bridging the gap between parallel simulation and the physical robot.
Policy variants
| Policy | Inputs | Terrain | Deployment |
|---|---|---|---|
| Blind flat-terrain walking | Proprioception only | Flat ground | Deployed on custom robot |
| Blind rough-terrain walking | Proprioception only | Rough terrain | Deployed on custom robot |
Result
The published demonstration shows the blind rough-terrain policy handling 15 cm stairs, dynamic terrain, and complex terrain. This is a demonstration result; a repeat-trial success rate and standardized robustness benchmark are not available in this project record.
Architecture
The training architecture uses a privileged critic during simulation while keeping the deployed actor dependent on observations available on the robot. This supports learning with richer simulation signals without requiring those signals at runtime.
Demonstration
The project media above shows the learned policy on the physical quadruped and the training architecture behind it.
Evidence & demonstration
Demos & project media

Project demonstration
Loads content from LinkedIn when you choose to play.Video description
The custom quadruped walks across physical test terrain using a policy based on proprioception. The video shows stairs, loose material, and changing support surfaces. The accompanying architecture diagram describes how training uses a privileged critic while the deployed actor uses robot observations.
