Article Series

Deep-dive series that build understanding across multiple articles, from foundations to production.

Policy Optimization for LLMs: From Fundamentals to Production

From PPO fundamentals through GRPO and GDPO to a map of the whole GRPO variant family — the complete policy optimization series for aligning language models with reinforcement learning.

5 parts · ~138 min total reading time

View series

Adaptive Optimization at Scale: Contextual Bandits from Theory to Production

A 5-part journey from decision frameworks and regret theory through algorithm implementations to production deployment of contextual bandit systems.

5 parts · ~113 min total reading time

View series