Home / The moat / Information Gain

Information Gain

The only content moat AI cannot copy

Search systems score how much a document adds beyond what they have already read. Almost every content programme on earth is optimised to add nothing. We build the other kind — and we are betting the firm on it.

5 / 5

engines sampled nightly, so the data is ours before it is anyone else’s

412

live prompts in the panel behind every claim we publish

< 40

Gain Score below which we refuse to publish a page

Why novelty is the scoring signal

I

Redundancy is the industry default

The standard playbook is to read the top ten and write a better version. To a system scoring marginal novelty, a better version of what it already holds is worth close to nothing — however well written it is.

II

Models cite what they cannot infer

An LLM only needs a source when the claim is not already in its weights. Original numbers, named practitioners and tests it has never seen force the citation. A good summary does not.

III

AI made volume free, so volume stopped being a moat

Anything a competitor can generate in an afternoon cannot defend you. The asymmetry has moved to information you had to go out and collect.

IV

Gain compounds, rankings decay

A dataset only you could gather gets quoted for years across every engine, without a refresh cycle. That is a balance-sheet asset, not a campaign.


On the record

Google holds a patent for estimating the information gain of a document relative to what a user has already been shown — scoring the marginal value of reading one more page on the same topic.

US 11,354,342 B2 — Contextual estimation of link information gain. Filed 2019, granted 2022. The mechanism generalises: every answer engine has a reason to prefer the source that adds something.

US 11,354,342 B2

Filed 2019 · Granted 2022

Marginal-value scoring

The Gain Score, and how we grade it

Five signals, scored out of 100 before a page is briefed and again after it ships. Below 40 we do not publish it.

01

Original data

Numbers you generated: panels, surveys, logs, your own funnel. Not a chart re-plotted from someone else’s study.

30 pts

02

First-hand testing

You ran the thing and reported what happened, including the failures. Method visible enough to be checked.

25 pts

03

Named expertise

A real practitioner on record, saying something specific enough that it could be wrong.

20 pts

04

Net-new framing

A model, a term or a distinction the topic did not have before you published it.

15 pts

05

Proprietary benchmark

A measure you defined and maintain, that others start quoting by name.

10 pts


Graded side by side

The same query, three pages, scored against the rubric above.

Graded side by side Same query · three pages
Original data None
First-hand testing Thin
Named expertise Byline only
Net-new framing Derivative
Proprietary benchmark None
12 A careful rewrite of the consensus. Everything on it, the engine has already read somewhere it trusts more.
Original data None
First-hand testing None
Named expertise None
Net-new framing Averaged
Proprietary benchmark None
05 Fluent, complete, and made entirely of what the model already knew. Zero gain by construction.
Original data 412-prompt panel
First-hand testing Run nightly
Named expertise Practitioners on record
Net-new framing Ours first
Proprietary benchmark Share of Answer
88 Numbers, tests and named practitioners that exist nowhere else. The engine has to cite us to say it.

Five instruments that manufacture gain

Gain is not a writing style. It is an operations problem: you have to go and get information that does not exist yet. These are the five ways we get it.

01

Original data

The 412-prompt panel

Five engines, sampled nightly on the real questions your buyers ask, held as a time series. We publish what moved and why.

Why nobody copies it: it is a year of nightly readings. You cannot start one retroactively.

02

First-hand

Teardowns run in public

We test the claim ourselves — schema variants, answer formats, citation patterns — and show the runs, including the ones that failed.

Why nobody copies it: requires clients, budget and a willingness to publish losses.

03

Expertise

Practitioner interviews

Named operators inside your category, on record, saying things a model has never seen written down anywhere.

Why nobody copies it: access is earned over years. It does not scrape.

04

Benchmark

Proprietary metrics

Share of Answer and Citation Rate: measures we defined, documented, and now get quoted for by name.

Why nobody copies it: whoever defines the measure owns the citation for it.

05

Yours alone

Internal ops data

Your funnel, pricing history and support archive, turned into publishable evidence with the sensitive parts stripped.

Why nobody copies it: it is literally your company. No competitor has it.

How a gain asset gets built

01

Redundancy audit

We score your existing library against the rubric and find where you are paying to repeat the internet.

Week 1

02

Gain gap map

Which questions in your category have no original source yet. Those are the assets worth building.

Week 1–2

03

Collection

The unglamorous part: running the panel, booking the interviews, pulling the ops data, doing the test.

Week 2–6

04

Publish for citation

Structured so an engine can lift the finding cleanly and has to attribute it. Answer-first, entity-clean, schema-backed.

Week 6+

05

Defend

Track who quotes the finding, refresh the dataset on a cycle, and extend the lead before anyone starts collecting.

Ongoing


Find out what your content actually adds

Send us ten URLs. We grade each one against the rubric above and hand back the Gain Score, the missing signals, and the three assets we would build first. Free, one week.