So, basic idea is to reformulate the reinforcement learning problem so
that you can perform a limited amount of operations.
