topic · 11 notes
LLM Serving, as it ships.
Engineering notes on LLM Serving by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
August 23, 2026 · LLM Serving · Quantization · vLLM
The local serving defaults that make an open-weight model feel dumber than it is
A Level1Techs post argues sampler settings, chat templates, attention backends and undocumented KLD claims explain why local open-weight models underperform.
August 22, 2026 · LLM Serving · Agents · Inference Cost
GPT-5.6 Sol drops 20% per token, and the agent stacks that won't feel it
OpenAI cut GPT-5.6 Sol pricing 20%. The class-level breakdown (cached input, output, reasoning) isn't confirmed, and both agent stacks in the sources bill via Codex seats, not tokens.
August 20, 2026 · Quantization · LLM Serving · GGUF
Unsloth Dynamic 3.0 puts its accuracy gains in the small GGUFs
Unsloth Dynamic v3.0 GGUFs for Qwen3.8-27B claim up to +10% top-1 accuracy at equal size, a new imatrix calibration set, and MTP stripped below UD-Q2_K_XL.
August 19, 2026 · Mojo · GPU Kernels · Open Source
Mojo's compiler goes Apache 2.0 while MAX stays source-available
Mojo 1.0's compiler and toolchain are now Apache 2.0 with LLVM exceptions. MAX is source-available with device restrictions dropped, and kernel work still needs a prebuilt compiler.
August 18, 2026 · GPU Serving · Rust · LLM Serving
Rust Gets Native GPU Offload via rustc and LLVM Backends
Drehwald et al. present a Rust-native GPU offload framework using LLVM Offload infrastructure. It matches CUDA/HIP kernel performance on RAJAPerf without unsafe pointers or vendor DSLs.
August 14, 2026 · LLM Serving · Open Weights · Agent Infrastructure
Qwen3.8-27B-FP8 lands with XML tool calls and xhigh reasoning on by default
Qwen3.8-27B-FP8 published 2026-08-14. FP8 27B implies roughly 27GB of weights; the card's chat template mandates XML tool calls, xhigh reasoning by default, and raises on tool-only turns.
August 13, 2026 · Agent Frameworks · LLM Serving · Tool Calling
DeepSeek's agent harness is public, but the sampling config isn't in the README
DeepSeek open-sourced dsh (MIT, 12,293 commits) on the same day V4 Pro 0813 hit OpenRouter at $0.435/$0.87 per 1M and 1M context. Here is what the release actually specifies.
August 12, 2026 · Apple Silicon · Metal · Inference Engines
What a single-model Metal engine buys you over llama.cpp on Apple Silicon
h3.c ships hand-written Metal for MiniMax-H3 only; a macOS VM capability shim made llama.cpp 11.08x faster at prompt processing. Two data points on portability's price.
August 11, 2026 · LLM Serving · AI Security · Reasoning Models
The encrypted reasoning block is the side channel, not token timing
Stolen Thoughts decoded 315,320 hidden reasoning blocks from 6,708 public agent trajectories by replaying encrypted thinking blocks into weaker same-provider models, recovering 704 secrets.
August 10, 2026 · Local Inference · Open Weights · Agentic Coding
Meta's 30B Muse Glimmer fits a local coding agent under 20GB, Apache 2.0
Meta's Muse Glimmer is a 30B open-weights agentic model under Apache 2.0, quantized to ~4-bit to fit under 20GB with room for KV cache, vision encoder and a speculative drafter.
August 7, 2026 · AI Cost Management · Agentic Coding · LLM Serving
Databricks cut AI coding spend 70% by chasing the efficiency frontier
Databricks reports a 70% cut in AI coding spend. The named mechanisms are internal evals, a meta-harness called Omnigent, and an AI Gateway. Savings figures are self-reported and directional.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.