Entropy Regularised Deterministic Optimal Control: From Path Integral Solution to Sample-Based Trajectory Optimisation

by   Tom Lefebvre, et al.
Ghent University

Sample-based trajectory optimisers are a promising tool for the control of robotics with non-differentiable dynamics and cost functions. Contemporary approaches derive from a restricted subclass of stochastic optimal control where the optimal policy can be expressed in terms of an expectation over stochastic paths. By estimating the expectation with Monte Carlo sampling and reinterpreting the process as exploration noise, a stochastic search algorithm is obtained tailored to (deterministic) trajectory optimisation. For the purpose of future algorithmic development, it is essential to properly understand the underlying theoretical foundations that allow for a principled derivation of such methods. In this paper we make a connection between entropy regularisation in optimisation and deterministic optimal control. We then show that the optimal policy is given by a belief function rather than a deterministic function. The policy belief is governed by a Bayesian-type update where the likelihood can be expressed in terms of a conditional expectation over paths induced by a prior policy. Our theoretical investigation firmly roots sample based trajectory optimisation in the larger family of control as inference. It allows us to justify a number of heuristics that are common in the literature and motivate a number of new improvements that benefit convergence.


page 1

page 7


Optimal Control as Variational Inference

In this article we address the stochastic and risk sensitive optimal con...

Stochastic Optimal Control as Approximate Input Inference

Optimal control of stochastic nonlinear dynamical systems is a major cha...

HJB Optimal Feedback Control with Deep Differential Value Functions and Action Constraints

Learning optimal feedback control laws capable of executing optimal traj...

Smoothing-Averse Control: Covertness and Privacy from Smoothers

In this paper we investigate the problem of controlling a partially obse...

Direct Policy Optimization using Deterministic Sampling and Collocation

We present an approach for approximately solving discrete-time stochasti...

GuSTO: Guaranteed Sequential Trajectory Optimization via Sequential Convex Programming

Sequential Convex Programming (SCP) has recently seen a surge of interes...

Solar-Sail Trajectory Design of Multiple Near Earth Asteroids Exploration Based on Deep Neural Network

In the preliminary trajectory design of the multi-target rendezvous prob...

Please sign up or login with your details

Forgot password? Click here to reset