CrawlQStudio

Product field note · Current as of 28 July 2026

Research doesn’t become wrong. It becomes unverifiable first.

Somewhere in your strategy deck is a finding from eighteen months ago that everyone now treats as fact. It might still be true. The uncomfortable part is that nobody can tell you — the evidence behind it was discarded when the deck shipped, and re-checking it would cost about what the original study cost. So it stands, unexamined, holding up decisions.

This piece is about the shelf life of research findings, not about how to run AI-assisted research. If you want the practical framework for the latter, that is a different piece: market research using generative AI.

How a finding decays

Nothing goes wrong at any single step. That is what makes it hard to catch.

  • Week 1

    Fully trusted

    The finding is fresh, the analyst is in the room, and everyone knows which data it came from.

  • Month 3

    Trusted, unverifiable

    The finding is quoted in a strategy doc. The underlying evidence is in a folder somebody would have to go looking for.

  • Month 9

    Quoted, unexamined

    It has become an internal fact. New work is built on top of it. Nobody has re-checked whether the market moved.

  • Month 18

    Load-bearing and unfalsifiable

    It underpins a positioning decision. The analyst has changed roles. Checking it would mean redoing the work, so nobody does.

The cost is not the wrong finding

It is tempting to frame this as a risk of acting on stale data, and that risk is real. But the larger cost is quieter: teams stop asking. When re-checking a finding is expensive, the rational move is to accept it, and an organisation slowly accumulates a set of beliefs nobody is incentivised to test. The research function keeps producing, and the stock of things everyone “knows” keeps growing, and the two are increasingly unrelated.

The second cost is duplication. Work gets redone because nobody can establish whether it was already done, or under what conditions. The analyst who would have known has moved teams. The deck exists but does not answer the question being asked now.

What has to survive alongside a finding

A finding on its own is a claim. What makes it checkable later is the material around it — and that material has to be captured at the moment of the work, because it cannot be reconstructed afterwards.

  • The evidence

    Which sources actually supported the claim — not a bibliography, the specific passages.

    Without it

    The finding becomes an assertion with a confident tone and no way to test it.

  • The method

    What was asked, of what corpus, under which constraints and exclusions.

    Without it

    You cannot tell whether a later contradictory finding is new information or a different question.

  • The date and version

    What was true when, and which version of the corpus it drew on.

    Without it

    Stale findings and current findings look identical in a slide.

  • The decision path

    Which alternatives were considered and why they were set aside.

    Without it

    Teams re-litigate settled questions, or repeat an analysis someone already ran.

Why AI raises the stakes

AI-assisted synthesis makes research faster, and that is a real gain. It also changes the arithmetic. When a team ran four studies a year, one analyst could plausibly hold the context for all of them in their head — the evidence-discarding problem existed, but a human patched over it. At forty synthesis passes a quarter, that patch stops working. Volume converts a tolerable weakness into a structural one.

There is a sharper version of the problem too. When a model synthesises across a corpus, the reasoning step — why these passages support this conclusion — is exactly the part that is not written down anywhere unless something deliberately captures it. The output reads as confident prose either way. This is the same failure we describe as context decay in why AI forgets your brand, arriving in the research function rather than the content one.

What a durable record looks like

Grounding matters first: synthesis that runs against your own corpus, with the retrieved passages recorded, is checkable in a way that synthesis against the open internet is not. That discipline is the subject of the brand intelligence and market research framework, and the structure underneath it is described on the knowledge graph page.

For findings that carry real consequence — the ones underpinning a positioning decision or an external claim — there is a stronger tier. Governed decisions get a cryptographic receipt: canonical JSON, a Merkle tree, an ed25519 signature. That receipt is anchored to the public Sigstore Rekor transparency log, run by the Linux Foundation rather than by us. Those two layers are shipped today. Further provenance layers — a richer evidence-record format, and embedded content credentials travelling inside a media file — are roadmap rather than shipped. The full ledger of which is which sits in Article 50: why compliance requires replayable memory.

The boundary on that, stated precisely rather than flatteringly: the export originates from your own CrawlQ tenant, so that first step is not independent of us. What is genuinely CrawlQ-free is the step after the anchor — checking a receipt in the public log uses infrastructure we do not run and tooling we do not control. The post-anchor verification step is independent of CrawlQ; the end-to-end journey is not. The controls are listed on our trust page.

The honest limits

Three things this does not do. It does not make a finding correct — a well-recorded conclusion drawn from a thin sample is still thin, and no amount of provenance repairs weak method. It does not tell you when the market moved; it tells you what you believed and on what basis, which is what makes a re-check cheap enough to actually run. And it does not remove the judgement call about whether a finding still applies — software can hold the evidence and show its age, but someone still has to decide.

What changes is narrower and worth having: the question “is this still true?” stops being a research project and becomes a lookup. Teams re-check things they would otherwise have accepted, because re-checking got cheap.

Questions we get asked

What is AI market research memory?
It is keeping the evidence, method, date and decision path attached to a research finding, rather than keeping only the finding itself. The practical difference shows up when somebody asks 'is this still true?' about a conclusion from last year. With memory, you re-open the record: these were the sources, this was the question, this was the corpus version, these alternatives were considered. Without it, the honest answer is that re-checking costs as much as redoing the research — which is why, in most organisations, nobody re-checks.
Does AI make this problem better or worse?
Both, and it is worth being precise about which. AI makes producing research dramatically faster, which is genuinely valuable. But volume is exactly what turns evidence-discarding into a structural problem: when a team could run four studies a year, an analyst could plausibly hold the context for all of them. When the same team generates forty synthesis passes a quarter, that stops being possible, and the fraction of findings nobody can re-derive rises quickly. Faster production without a matching record does not give you more knowledge — it gives you more claims.
Is this not what a research repository already does?
A repository is a real improvement and often the right first step — it solves findability, which matters. The gap, as repositories are typically configured, is that they store outputs: decks, reports, summaries. What they usually do not store is the relationship between a specific claim and the specific evidence that supported it, the constraints in force at the time, or the corpus version it drew on. So a repository will reliably find you the deck from March. It will not usually tell you whether the third bullet on slide nine still holds, which is the question people actually have.
How is this different from just citing sources?
Citation is the right instinct and it covers part of the problem. Two things it tends not to survive, as citation is usually practised. First, citations point outward at documents but rarely capture the reasoning step — why this passage supports this conclusion rather than a weaker one. Second, a citation is a snapshot: on its own it does not tell you the source has since been superseded, or that the corpus it belonged to has moved on. Attribution answers 'where did this come from'. Memory answers 'does it still stand, and how would I know'.
Can an outsider verify a research record without CrawlQ's help?
Partly, and the boundary matters. The export originates from your own CrawlQ tenant — you produce the record from the platform, so that first step is not independent of us. What is independent is everything after the anchor: governed decisions get a cryptographic receipt, and once that receipt is anchored to the public Sigstore Rekor transparency log, its presence can be checked against infrastructure we do not operate with tooling we do not control. Those two layers — the receipt and the public anchor — are shipped today. Further provenance layers, including a richer evidence-record format and embedded content credentials, are roadmap rather than shipped. So the accurate claim is narrow: the post-anchor verification step is CrawlQ-free, not the whole chain.
Does this apply if our research is mostly qualitative?
Arguably more, not less. Quantitative findings carry some of their own context — a number has a sample size and a date attached almost by convention. Qualitative synthesis is where the reasoning step does the heavy lifting and is least likely to be written down: which interviews supported a theme, which contradicted it, what the analyst discounted and why. That is precisely the material that evaporates when the analyst moves on, and precisely what a record has to hold for a qualitative finding to stay checkable.

Research that stays checkable.

The brand intelligence framework covers grounding synthesis in your own corpus — and what a defensible insight looks like when the board asks where it came from.