topic · 3 notes
Agent Infrastructure, as it ships.
Engineering notes on Agent Infrastructure by Samir Sengupta - each one read from primary sources on the day it happened, with what it changes for people building on it.
August 14, 2026 · LLM Serving · Open Weights · Agent Infrastructure
Qwen3.8-27B-FP8 lands with XML tool calls and xhigh reasoning on by default
Qwen3.8-27B-FP8 published 2026-08-14. FP8 27B implies roughly 27GB of weights; the card's chat template mandates XML tool calls, xhigh reasoning by default, and raises on tool-only turns.
August 11, 2026 · LLM Serving · AI Security · Reasoning Models
The encrypted reasoning block is the side channel, not token timing
Stolen Thoughts decoded 315,320 hidden reasoning blocks from 6,708 public agent trajectories by replaying encrypted thinking blocks into weaker same-provider models, recovering 704 secrets.
August 8, 2026 · Agent Infrastructure · Web Crawling · Rate Limiting
What an agent fleet did to Hugging Face, and what to cap before yours does it
OpenAI's Black Hat timeline of the Hugging Face incident plus one site's 35,000:1 crawl-to-referral ratio, and the concurrency, caching and backoff limits to build into a fetch pipeline.
Hiring for AI or ML?
I am open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. New York based, shipping worldwide.