LLM Post Training
Modern post-training can involve supervised finetuning (SFT), preference optimisation, RL training using various policy optimisation algorithms and distillation. Fun commentary on the meme by Nathan Lambert A good analogy I think of different stages is Pretraining: Learning language, knowledge and task representations through next-token prediction. SFT: Adapting a pretrained model to imitate desired responses, follow instructions and produce task-specific output formats.A Preference optimisation: Learn which responses should be preferred over others. RL or verifiable RL: It takes one step further, optimizing model behaviour using rewards, preferences or verifiable outcomes rather than only imitating reference responses. Distillation: Learn from the behaviour of a stronger teacher model, rather than only from fixed target responses or scalar rewards. The Smol Training Playbook ...