topic · 2 notes
Local Inference, as it ships.
Engineering notes on Local Inference by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
August 20, 2026 · Quantization · LLM Serving · GGUF
Unsloth Dynamic 3.0 puts its accuracy gains in the small GGUFs
Unsloth Dynamic v3.0 GGUFs for Qwen3.8-27B claim up to +10% top-1 accuracy at equal size, a new imatrix calibration set, and MTP stripped below UD-Q2_K_XL.
August 10, 2026 · Local Inference · Open Weights · Agentic Coding
Meta's 30B Muse Glimmer fits a local coding agent under 20GB, Apache 2.0
Meta's Muse Glimmer is a 30B open-weights agentic model under Apache 2.0, quantized to ~4-bit to fit under 20GB with room for KV cache, vision encoder and a speculative drafter.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.