When a correct reward is not enough: testing PPO against an exact broker strategy
This paper asks a simple question: can a common reinforcement learning method rediscover a known, exact trading strategy from the equations that define it? The authors place a learning agent that uses proximal policy optimization (PPO) into a continuous-time broker–trader model that has an analytical solution. The PPO agent replaces the broker and chooses trading speed while interacting with an informed trader and random, uninformed client orders.
To make the comparison fair, the team derived a discrete, finite-step reward that matches the broker’s continuous-time payoff. They checked that the discrete reward is correct by refining the time grid and by using an exact one-step identity. They then trained two kinds of agents: a feed‑forward neural actor (PPO–FFNN) and a recurrent actor using long short‑term memory (PPO–LSTM). The experiments address three questions: recovering the analytical policy when all state variables are observed (full information), doing well when some variables are hidden (partial information), and adapting a frozen analytical policy when market costs change.
The results are mixed. When there is no random uninformed flow, the PPO–FFNN found actions that approach the analytical reference. But when the uninformed flow is stochastic, both PPO–FFNN and PPO–LSTM gave inaccurate actions. The authors checked that the neural networks could in principle represent the right actions via supervised learning. The deeper problem was the agent’s critic — the component that estimates which actions are better — which did not reliably rank nearby actions in Monte Carlo diagnostics. The authors also tried potential‑based reward shaping, a technique meant to help learning, but it did not provide reliable improvement.
Under partial information, a simple certainty‑equivalent controller performed well. “Certainty‑equivalent” here means a causal controller that estimates the hidden variables from observed history and then applies the analytical rule. That controller stayed close to the reference strategy. PPO agents, by contrast, showed larger errors and achieved lower payoffs when important state variables were hidden.