Flat vector pipeline illustration: three web pages tagged 'Stale', 'Wrong region' and 'Blocked' feed into a 'Model' desk panel that passes them through faithfully, ending in an output report with a red flag on one row.

Last Thursday a fractional CFO showed me a competitor pricing report her research agent had assembled overnight. Five competitors, full pricing tables, clean formatting, ready by 7 AM. She'd already dropped two of the numbers into a client deck when she happened to check one against the competitor's live site.

Off by 40%. The next one too. Three of the five tables were wrong.

Her first move was to open settings and look for a smarter model.

Wrong layer.

The autopsy took eleven minutes.

We traced the run together. The agent had visited five competitor sites. One served it a cached page from last March, old pricing still on it. One served the European version of the site, so every number included VAT. One refused the visit entirely, and the agent fell back to a third-party "pricing roundup" blog post from two years ago.

Then the model did exactly what models do. It read what it was handed and summarized it faithfully. Confidently. In perfect formatting.

The reasoning was fine. The reading list wasn't.

Blame travels to the visible layer.

When an agent gets something wrong, operators blame the part they can see: the model, the prompt. So they upgrade the subscription, switch providers, rewrite the instructions. Then the new model reads the same stale page and produces the same wrong number, just phrased better.

Here's the order of operations inside every research agent: it fetches first and thinks second. Most failures happen in the first step, where nobody is looking. The thinking step gets all the scrutiny because it's the step with a brand name attached.

This is Context Debt showing up at runtime. You may have written down who your clients are and what your voice sounds like. But if you never wrote down which sources count as truth, the agent decides for itself. And it decides badly, because to a model, a cached page and a live page look identical.

Add a sources section to the onboarding document.

The fix is structural, and it's small. In the same document where you tell the agent who it works for, add a section that tells it what it's allowed to treat as a source:

  1. The approved list. The five to ten places numbers may come from. Official pricing pages, your own data exports, the industry report you actually pay for. Name them.

  2. The citation rule. Every figure in every output carries its source URL and the date it was retrieved. No exceptions. A number without both gets flagged, not shipped.

  3. The freshness rule. Anything older than your tolerance (30 days for pricing, pick your own for everything else) gets marked stale in the output, in plain text, where you'll see it.

Then run a ten-minute spot check once a week: pick three numbers from the agent's output at random and verify them by hand. Track your hit rate. Until spot-checked figures verify at 100% for a month straight, nothing the agent produces goes in front of a client unreviewed.

Same model. Same prompt. Different reading list.

The CFO didn't switch models. She added a 14-line sources section to her agent's context document and turned on the citation rule. The next morning's report had one flagged figure in it: a stale cache, caught and labeled by the agent itself before she ever opened the file.

The model was never the problem. It just summarized whatever walked in the door. Decide what walks in the door.

Free Context Stack → promptsquad.beehiiv.com — the 6-layer context document that tells your AI what good looks like, including what counts as a source. Subscribe and the framework comes to you.

Forward this to a friend who just upgraded their model subscription because the numbers came back wrong.

— Chris

PromptSquad · AI strategies that hold up in the real world.