topic · 2 notes
Model Evaluation, as it ships.
Engineering notes on Model Evaluation by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
GPT-6 Astra code review cost evaluation: 4% more bugs at 2.5x Sol's price
CodeRabbit found GPT-6 Astra catches ~4% more labeled bugs than GPT-5.6 Sol overall and 20% more cross-file, at $10/$50 per 1M tokens versus Sol's $4/$20.
DeepMind open sources WeatherNext Cyclones, a 1,000-member ensemble at 28km
WeatherNext Cyclones runs 1,000-member ensembles at 28x28km, produces a 15-day forecast in under a minute on a TPU, and is now open source. What to verify before you build on it.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.