topic · 2 notes
KV Cache, as it ships.
Engineering notes on KV Cache by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
Flash-dLLM pairs an I/O-aware KV cache with self-drafting decode, and reports no numbers
Flash-dLLM (arXiv 2609.26796v1, 22 September 2026) fuses an I/O-aware KV-cache kernel with self-drafting decode, but reports no speedup or memory figure.
DeepSeek v4.1 Flash pricing, benchmarks, and 890 bytes per token
DeepSeek V4.1-Flash is a 552B MoE with 8B active params at prefill, 16B at decode, 890 bytes per token KV cache, and off-peak rates at 50% of peak.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.