topic · 2 notes
Apple Silicon, as it ships.
Engineering notes on Apple Silicon by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
September 2, 2026 · LLM Serving · Apple Silicon · Mixture of Experts
Run 104GB Qwen3.8-Flash-Next on a 48GB Mac via slotstream at ~12 tok/s
slotstream keeps a 3.8GB dense trunk resident and streams 68GB of experts from SSD into a 20.1GB slot pool for ~12 tok/s decode on a 48GB M5 Pro.
August 12, 2026 · Apple Silicon · Metal · Inference Engines
What a single-model Metal engine buys you over llama.cpp on Apple Silicon
h3.c ships hand-written Metal for MiniMax-H3 only; a macOS VM capability shim made llama.cpp 11.08x faster at prompt processing. Two data points on portability's price.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.