Welcome to my small corner of the internet where I share things I’m learning, building, and exploring.
- I work as MLOps Engineer at Fuzzy Labs
- You can find me on the following socials.
Welcome to my small corner of the internet where I share things I’m learning, building, and exploring.
Modern post-training can involve supervised finetuning (SFT), preference optimisation, RL training using various policy optimisation algorithms and distillation. Fun commentary on the meme by Nathan Lambert A good analogy I think of different stages is Pretraining: Learning language, knowledge and task representations through next-token prediction. SFT: Adapting a pretrained model to imitate desired responses, follow instructions and produce task-specific output formats.A Preference optimisation: Learn which responses should be preferred over others. RL or verifiable RL: It takes one step further, optimizing model behaviour using rewards, preferences or verifiable outcomes rather than only imitating reference responses. Distillation: Learn from the behaviour of a stronger teacher model, rather than only from fixed target responses or scalar rewards. The Smol Training Playbook ...
Learning programming languages is fun. My usual path is a couple of years writing the code and building projects in the particular language to get comfortable. It includes learning the best practices and understanding the different ways something could have been implemented. Then, occasionally, I get curious about what is happening under the hood sort of like peeling a layer of onion. Eventually, peeling back enough layers, it reaches the compiler and finally the assembly or machine code. This code is then executed by the CPU. ...
Measuring vLLM cold-start bottlenecks on GKE and evaluating ways to reduce time to first request.
Architecture components of Docker
Implementing and profiling pipeline parallelism techniques using PyTorch