topic · 4 notes
Quantization, as it ships.
Engineering notes on Quantization by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
August 23, 2026 · LLM Serving · Quantization · vLLM
The local serving defaults that make an open-weight model feel dumber than it is
A Level1Techs post argues sampler settings, chat templates, attention backends and undocumented KLD claims explain why local open-weight models underperform.
August 20, 2026 · Quantization · LLM Serving · GGUF
Unsloth Dynamic 3.0 puts its accuracy gains in the small GGUFs
Unsloth Dynamic v3.0 GGUFs for Qwen3.8-27B claim up to +10% top-1 accuracy at equal size, a new imatrix calibration set, and MTP stripped below UD-Q2_K_XL.
August 14, 2026 · LLM Serving · Open Weights · Agent Infrastructure
Qwen3.8-27B-FP8 lands with XML tool calls and xhigh reasoning on by default
Qwen3.8-27B-FP8 published 2026-08-14. FP8 27B implies roughly 27GB of weights; the card's chat template mandates XML tool calls, xhigh reasoning by default, and raises on tool-only turns.
August 10, 2026 · Local Inference · Open Weights · Agentic Coding
Meta's 30B Muse Glimmer fits a local coding agent under 20GB, Apache 2.0
Meta's Muse Glimmer is a 30B open-weights agentic model under Apache 2.0, quantized to ~4-bit to fit under 20GB with room for KV cache, vision encoder and a speculative drafter.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.