topic · 2 notes
Reinforcement Learning, as it ships.
Engineering notes on Reinforcement Learning by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
How a 4B model trained with RL produced query plans faster than Postgres
A 4B Qwen model post-trained with SFT and a custom GRPO variant hit a 1.81x geometric mean speedup and 44.7% lower summed latency on 113 join-heavy queries.
Cognition SWE-2 posts 92.8 on the Terminal-Bench 2.1 benchmark, 27.3 on Terminal-Bench 4
Cognition's SWE-2, a Kimi K3 post-train, scores 92.8% on Terminal-Bench 2.1 but 27.3% on Terminal-Bench 4, cuts steps 127 to 53, and ships only in Devin.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.