topic · 2 notes
LLM Inference, as it ships.
Engineering notes on LLM Inference by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
GLM 5.3 Flash after a month on Wagtail: 1B of 2B tokens, $68, 4 kWh
GLM 5.3 Flash took 1B of Wagtail's 2B September tokens for $68 and 4 kWh. A wrong-model run, provider limits and R&D, not the model, took the rest.
What the Jeeves decision model's reasoning buys: 5 points for 11x the latency
The Jeeves decision model's reasoning lifts dev accuracy from 0.775 to 0.825, but median latency rises from 0.3 s to 3.3 s and p90 hits 17.1 s on one H100.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.