Open to workNew YorkGet in touch

Coding Agents

GLM 5.3 Flash after a month on Wagtail: 1B of 2B tokens, $68, 4 kWh

Wagtail tried to run a month of agentic coding on one cheap open model and got halfway. The model held up; budgets, providers and model selection did not.

Published
October 3, 2026
Read
6 min
Author
Samir Sengupta
Wagtail's month on GLM 5.3 Flash: 1B of 2B tokens on the target model for $68 and about 4 kWh
Note 081 / 081Daily note · Written from 2 sources

the short version

  • A flash-tier open model handled core Wagtail work, sites, UI tasks, documentation and evals, but only 1B of the month's 2B tokens actually went to it.
  • One overnight build of Wagtail's MCP server on the non-Flash GLM 5.3 used 450M tokens, $150 and 5 kWh. A Hacker News reply says that one session was 150% of the monthly budget.
  • GLM 5.3 Flash performance degraded on capacity-limited providers, forcing switches to DeepSeek V4.1 Flash and Qwen 3.8 Flash.
  • Wagtail's October plan treats constant, local measurement of spend and energy, plus an explicit experimentation budget, as part of the setup.

Wagtail set out to run all of September's coding work on GLM 5.3 Flash and got halfway. 1B of 2B tokens went to the target model, for $68, about 4 kWh and 365 grams of carbon emissions. The model handled Wagtail core work, sites built on Wagtail, UI tasks, documentation and evals. The other 1B tokens went elsewhere for three reasons: an MCP server build that ran on the pricier non-Flash GLM 5.3, provider capacity limits, and R&D across many models. The sources give no head-to-head numbers against frontier models.

The Wagtail post does not show the model failing at the work. It shows that model selection, provider capacity and experimentation decided where the tokens went. Total energy use for the month came to about 35 kWh against a 10 kWh target. The post does not break that gap down by cause.

Where GLM 5.3 Flash held up

The author calls the model excellent and cites three properties from Wagtail's earlier models comparison as the reason it suited a single-model month. The task mix was broad: Wagtail itself, sites built with it, UI tasks, AI R&D, some documentation writing and a lot of evals. The author's summary is that it was "really polyvalent."

  • 1M context window, which the author says makes any extended coding task possible.
  • Vision support, so it can build off screenshots or do visual QA.
  • Availability across a wide range of providers, so users see the benefits of healthy competition.

The author wants to continue the challenge, but only for the 50% or more of usage that is normal day-to-day production work, not R&D. The post calls it totally viable to focus day-to-day developer work on one or two flash-tier cheap models. It proposes that the majority of AI inference work should run on such efficient models, measured in cost or energy use rather than tokens.

One wrong-model session cost 450M tokens

The largest single miss came from building Wagtail's experimental MCP server, which the team describes as a vibe-coded prototype. The post says the 'wrong' model consumed 450M tokens, $150 and 5 kWh almost overnight. A Hacker News commenter replying in the first person as the author explained what happened. The non-Flash GLM 5.3 had run for the whole build without anyone noticing. Because of the price difference between the two models, that one session came to a quarter of the month's spend and 150% of the budget.


We could have achieved similar results for most likely 5x less cost with not that much more effort.
Wagtail blog, One month coding with GLM 5.3 Flash

For teams building agent setups, the implication is that model choice has to be enforced, not assumed. When two models share a family name and differ sharply in price, one long autonomous session can exceed a monthly budget. The MCP server does work, and the author says it now serves as a demo of its capabilities.

Provider capacity limits forced fallback models

The second hurdle was infrastructure. Wagtail's Pareto frontier of candidate models is filtered to open models it could confirm were available in a European data center. The providers it chose work most of the time, but the post says they are popular and do not have the same capacity as the big labs. Performance of GLM 5.3 Flash in particular degraded. The author's likely explanation is that the model sits high on the Pareto frontier for Wagtail's work.

The fix was to switch to similar models, DeepSeek V4.1 Flash and Qwen 3.8 Flash. The author says the switch was very simple but unexpected. Teams with similar data-residency constraints have a smaller set of models and providers to choose from. Naming fallback models in advance is a reasonable precaution.

The harness needs measurement, budgets and agent roles

Usage tracking ran through AgentsView, which Wagtail lists among its Agentic engineering recommendations. Energy figures came from Neuralwatt. On Hacker News, a commenter replying as the author said these measurements cover GPU energy use only, and that they make model efficiency much more visible than token counts do. The third source of off-target usage was deliberate: benchmarking models on Wagtail tasks, which needs data across a wide range of models. The post includes only a sneak peek of that benchmark.

For October, the post lists the changes it expects will make a cheap-model month work:

  1. Constant, local measurement and reporting of tokens, energy use and spend, and ideally of whether usage leads to concrete positive outcomes.
  2. Budgeting for experimentation, not just day-to-day tasks, with more deliberate decisions about which prototypes are worth building and how.
  3. Better prompt selection and multi-agent techniques: orchestrator, scout, implementer and reviewer agents with bounded goals.
  4. Continued pursuit of more efficient techniques and models, including Jev-style decision diffusion models and the latest flagship models.

Wagtail also plans to make leaner models more viable on its tasks through agent skills and a new CLI prototype intended to work well with agents. On efficiency, the author writes that Jev-style decision diffusion models look very promising if they can run as efficiently as hoped.

What Hacker News commenters push back on

Commenters split on where the Flash model fits in an agent pipeline. One said it is a pretty decent coder but should be paired with a good planner and reviewer, naming astra low for planning and sol 6.1 medium for reviews. Another, a z.ai subscriber who does not self-host, took the opposite view. They find it good enough for planning and reviewing, but too slow for execution despite the name. Neither claim comes with measurements, and the Wagtail post does not report latency.

The energy figures drew the most discussion. One commenter calculated that energy is about 1% of the total cost. Another asked whether the figures account for the full lifecycle, including data center construction, cooling, chip manufacturing and idle chips. A commenter replying as the author raised a different concern: two months ago they used 10x fewer tokens and probably not much more than 5 kWh on inference, against about 30 kWh this month. That figure differs from the post's roughly 35 kWh, and the thread does not reconcile the two.

Another commenter said that, according to the article, full GLM 5.3 sits in a "10x to 100x" energy range relative to the Flash figures. The post's text does not state that ratio directly. One commenter runs Qwen3.8-Flash-Next locally on a 395+ machine and says it does most properly planned tasks, with speed as the limitation. Another asked how the Flash and non-Flash versions compared and whether they were used for different tasks. The thread gives no answer.

Frontier gap and self-hosting remain unmeasured

The sources do not say how far GLM 5.3 Flash trails frontier models on Wagtail tasks. Wagtail's benchmark exists, but the post's text gives no scores. The post also says nothing about self-hosting GLM 5.3 Flash. All reported usage went through inference providers, so hardware requirements, quantization behavior and throughput on owned GPUs are open questions. Latency is also unreported, which matters given the commenter who found the model too slow for execution.

Three things to watch: Wagtail's October attempt to keep most inference on one or two flash-tier models, publication of the Wagtail task benchmark, and the update the author promises at Wagtail Space 2026 in November. Until those numbers exist, the supportable claim is narrow. A cheap open model can carry day-to-day agentic coding on a real codebase. That holds only if the setup enforces model choice, measures spend and energy continuously, and has fallback models ready.

Questions this raises

Is GLM 5.3 Flash good enough for daily coding agent work?

Wagtail calls it excellent and polyvalent after using it for Wagtail core work, sites, UI tasks, documentation and evals. The author says it is totally viable to focus day-to-day developer work on one or two flash-tier cheap models. The sources give no head-to-head numbers against frontier models.

How much did GLM 5.3 Flash cost for a month of coding?

Wagtail ran 1B tokens through GLM 5.3 Flash for $68 and about 4 kWh. By contrast, one session that accidentally used the non-Flash GLM 5.3 consumed 450M tokens, $150 and 5 kWh.

Why did Wagtail not run the whole month on GLM 5.3 Flash?

Three reasons: an MCP server build that ran on the pricier non-Flash GLM 5.3, provider capacity limits that degraded GLM 5.3 Flash performance, and R&D benchmarking across many models. The post does not show the model failing at the work.

These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.

More notes

Building something on this?

I ship production LLM, RAG and agentic systems for a living - the infrastructure behind the things these notes are about. Open to roles, contract work and research collaboration.