topic · 2 notes
Benchmarks, as it ships.
Engineering notes on Benchmarks by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
August 25, 2026 · Coding Agents · Benchmarks · Evaluation
Only 28 of 520 agent runs actually completed a whole-repo stack migration
SWE Refactor Bench audits whether a migration happened, not just that tests pass. Only 28 of 520 runs passed all three stages, and 13 of 20 tasks got nothing.
August 7, 2026 · AI Cost Management · Agentic Coding · LLM Serving
Databricks cut AI coding spend 70% by chasing the efficiency frontier
Databricks reports a 70% cut in AI coding spend. The named mechanisms are internal evals, a meta-harness called Omnigent, and an AI Gateway. Savings figures are self-reported and directional.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.