nodus.examples.rl
View MarkdownReinforcement learning in one file: teach a small model to spell words backwards.
Run it with python -m nodus.examples.rl. To train on your own task, copy this file and change tasks and reward:
a task is a prompt and the answer only reward sees, and reward scores one completion against that answer.
main() -> Nonereward
Section titled “reward”reward(completion, answer)1.0 for the exact reversed word, 0.0 for a wrong one, None when the completion holds no answer.