Autonomous Navigation and Collision Avoidance for Drones

A simulated forest, a sampling-based expert planner, and a student policy trained to fly through it from depth images alone.

2024 · CSCI 5561 Computer Vision, University of Minnesota · advised by Junaed Sattar · created with Marcus Berg and Vaagishwaran Sivakumar

The dataset generator flying a trajectory through a procedural forest, with the depth camera the policy sees inset.

A reimplementation of "Learning High-Speed Flight in the Wild". A custom C++ and OpenGL renderer generates procedural forest scenes and the paired depth imagery for them. A Metropolis-Hastings sampler produces multi-modal expert trajectories through those scenes. A convolutional student policy is then trained to predict trajectories directly from depth, without access to the ground-truth map the expert used.

My role

Wrote the OpenGL dataset generator, covering scene synthesis, depth rendering, and octree collision, and the training pipeline for the student policy.

Media

The naive straight-line path from start to end cuts through the trees, and the collision-free trajectory follows the corridor.
The naive straight-line path from start to end cuts through the trees, and the collision-free trajectory follows the corridor.
The course the trajectories are planned across.
The course the trajectories are planned across.
The same generator in an indoor scene. The red path is the expert trajectory, and the inset is the rendered depth.
The dataset generator's viewer, showing a procedural forest scene.
The dataset generator's viewer, showing a procedural forest scene.
The same scene with the rendered depth image the policy consumes.
The same scene with the rendered depth image the policy consumes.
An indoor scene used to vary the geometry the generator produces.
An indoor scene used to vary the geometry the generator produces.
Procedurally generated environment used for training rollouts.
Procedurally generated environment used for training rollouts.
Rendered depth image, the only input the student policy receives.
Rendered depth image, the only input the student policy receives.
Minimum-snap trajectory through the generated environment.
Minimum-snap trajectory through the generated environment.
Metropolis-Hastings sampling of multi-modal expert trajectories.
Metropolis-Hastings sampling of multi-modal expert trajectories.
Student policy architecture.
Student policy architecture.
Network diagram from the implementation notes.
Network diagram from the implementation notes.
Samples from the generated dataset.
Samples from the generated dataset.
Training loss.
Training loss.
Validation loss.
Validation loss.
Loss curve from the training run.
Loss curve from the training run.

Documents