An old attempt to train a model-based RL system for solving the inverted pendulum problem. Also contains a custom environment implementation with graphics. The model encodes uncertainty which is to be reduced over time. There are almost certainly some mistakes on the theoretical side.