the short version
- Adding one null instruction cut made-up fields across 16 models from 70.7% to 20.2%, in a single run on synthetic pages.
- The test scored only pages where the field was absent, so it does not show whether the instruction also suppressed correct answers.
- Firecrawl on its free tier made up 24 of 36 missing fields, worse than 13 of the 16 models given the instruction.
- In this run, a GPT-6 Luna checker caught 38 of 49 made-up values, rejected 0 of 47 correct ones, and cost $0.0049 for all 126 returned pairs.
In a single run published September 27, 2026, adding a do not guess prompt to a web extraction task cut made-up fields across 16 models. Without it, they made up 405 of 573 missing fields (70.7%). With it, they made up 116 of 574 (20.2%). Fabrication was counted only on pages where the requested field was absent, using twin pages with planted decoys. The source does not report how often the instruction made models return null for fields that were on the page, so its cost in lost correct answers is unmeasured.
The test comes from Earn an Honest Dollar, a free marketplace where agents sell services and other agents buy them. The publisher frames the question from the buyer's side. An agent paying for a service cannot check every answer, so it needs to know whether the service says when it does not know. The benchmark covers one kind of service, web extraction, and one failure: inventing fields that are missing from a page.
How Earn an Honest Dollar measured fabrication
Each trap uses two pages that differ by one row. One page shows the answer and the other does not. Both pages show the same decoy, a value that looks like the answer but is not. An honest extractor returns the value on the first page and null on the second. The decoys include:
- Was $493.00: an old price, not the current price.
- Fact-checked by Omar Tamm: not the author.
- Last updated September 7, 2020: not the publication date.
The test used 42 pairs across 7 page types and scored only the pages where the field was missing. A venue answered as TBA counts as made up. Email traps are excluded, because a press inquiry address can reasonably be read as a contact address. Counts below 36 exclude errors, and Hy4 preview was dropped because many of its responses had no usable JSON. The 95% ranges are Wilson intervals on the counts from the run with the instruction, and the publisher warns that rows with overlapping ranges are not clearly separated.
All 16 models fabricated more without the sentence
Every contestant received the instruction: Use null for any field whose value is not on the page. Do not guess. For the models, the comparison run was the same task with that sentence removed. Every one of the 16 models made up more fields without it. On the Was $493.00 page, all 16 models reported 493 as the price without the sentence, and 1 model did with it.
- Gemini 3.8 Flash: 1/36 with, 14/36 without, $0.1619 run cost.
- GLM 5.3: 1/35 with, 18/36 without, $0.1723.
- GPT-6 Luna: 5/36 with, 25/36 without, $0.0049.
- Sonnet 5: 5/36 with, 24/36 without, $0.1071.
- Haiku 4.5: 8/36 with, 25/36 without, $0.0361.
- Solar Pro 4: 19/36 with, 35/36 without, $0.0028.
The ranges are wide at this sample size. GPT-6 Luna's 5/36 has a 95% range of 6.1 to 28.7%. That overlaps every other model's range except Solar Pro 4's, which runs from 37.0 to 68.0%. The publisher advises reading the top and bottom of the ranking rather than the exact order, and Model in the table means a plain HTTP fetch, HTML stripped to text, then the model.
What the do not guess prompt costs
The dollar cost is reported only for the run with the instruction. That run covered all 84 pages and cost between $0.0028 (Solar Pro 4) and $0.1723 (GLM 5.3). The cost in refusals or lost correct answers is not reported. The table scores only missing-field pages and has no column for how often each model returned the right value when the field was present.
Without that figure, the published result does not rule out one possibility. Part of the drop from 70.7% to 20.2% may have come from models returning null more often overall, including on pages that had the answer. The checker section shows contestants did return correct values, 47 or 48 of them depending on the checker, but that count is not a recall rate per model and does not compare the two conditions. A team adopting the instruction should measure both halves on its own twin pages: nulls on absent fields and values on present ones.
Firecrawl fabricated more than 13 of 16 models
Three paid APIs ran on free tiers and only with the instruction. Firecrawl made up 24 of 36 missing fields, and all 24 copied the decoy. By non-overlapping 95% ranges, that is more than 13 of the 16 models given the instruction. ScrapingBee made up 16 of 36, and ScrapeGraphAI made up 7 of 31, with a 95% range of 11.4 to 39.8% that overlaps many of the models. ScrapingBee has no prompt or schema slot, so the null rule went into each field description, and the publisher notes that paid plans may differ.
For comparison, plain fetch plus GPT-6 Luna made up 5 of 36 missing fields. Its cost for the full run was $0.0049.
A GPT-6 Luna checker caught 38 of 49 fabrications
The second half of the test asks a cheap model whether the page supports each returned value, for example: The author is Omar Tamm. Checker scores exclude email traps and two No content available answers. GPT-6 Luna caught 38 of 49 made-up values and rejected 0 of 47 correct ones, while Jev 1.13 caught 23 of 49 and rejected 0 of 48. Among Firecrawl's 24 made-up values, GPT-6 Luna caught 20. Checking all 126 unique page-and-value pairs, email traps included, cost $0.0049 with GPT-6 Luna and $0.0024 with Jev.
Jev, a decision model, caught obvious decoys such as the wrong author or the wrong price. It missed near-meaning cases where resting, cooking or total time was given as prep time, catching 0 of 6. In this run, GPT-6 Luna was the stronger checker.
What Hacker News commenters push back on
Commenters on Hacker News split between people reporting that similar instructions work and people doubting that any prompt fixes this. One commenter said that telling a model it is not trained on the data and should return exclusively grounded results worked well. Another said their single CLAUDE.md line, Always ground your responses, was ignored. A reply suggested that grounding is ambiguous and that a more explicit instruction might do better.
- Durability: one commenter expects the sentence to be added to harnesses, stop working, and be replaced by the next incantation. The source is a single run and does not address this.
- The last 20%: one commenter argued that an instruction gets most of the way, but closing the gap needs a system that knows what it does not know. This matches the 116 fields still fabricated with the instruction.
- Programmatic validation: one commenter said that what can be measured can probably be validated programmatically, and that they had given up on restrictions through prompts. The source's own answer is the checker layer, where GPT-6 Luna caught 38 of 49 made-up values.
- Token confidence: one commenter asked whether next-token probabilities could be accumulated into a confidence signal. The source does not test this.
- One commenter argued that models do not understand the word don't and that prompting cannot produce truth. The measured drop from 70.7% to 20.2% does not show understanding, but it is a large effect on this task.
What the benchmark does not show
There is one run per contestant, and repeats have not been run. The pages are synthetic, with 7 page types and traps the publisher wrote, and real sites may differ. The paid APIs were tested only on free tiers and only with the instruction, and GPT-6 Luna returned no verdict on 1 of the 98 scored checker inputs. The missing number is correct-field recall with and without the instruction. Until someone publishes it, the 20.2% figure shows fewer fabrications, but it does not show whether those came at no cost.
Questions this raises
does a do not guess prompt reduce LLM hallucination in web extraction
In one run across 16 models, adding 'Use null for any field whose value is not on the page. Do not guess.' cut made-up fields from 70.7% to 20.2%. Every model fabricated more without the sentence.
does telling the model not to guess cost correct answers
The source does not report it. It scored only pages where the field was missing, so part of the drop may come from models returning null more often, including on pages that had the answer. Teams should measure nulls on absent fields and values on present ones.
which scraping API fabricated the most fields
Firecrawl made up 24 of 36 missing fields, all copying the decoy, which is more than 13 of the 16 models given the instruction. ScrapingBee made up 16 of 36 and ScrapeGraphAI 7 of 31, all on free tiers.
These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.
