How PageNews measures coverage
When several publishers cover the same event, PageNews compares the language each of them used and shows where they differed. This page explains exactly how those numbers are produced, and what they do and do not mean.
It is not a language model
The analysis is rule-based. PageNews holds curated lists of terms associated with particular kinds of framing — the vocabulary of confrontation, of negotiation, of policy mechanics, of household cost — and counts which of them a report used, weighted by how strongly each term discriminates. There is no model, nothing is generated, and the same text always produces the same scores.
That is a deliberate choice rather than a limitation we regret. The evidence behind every score is the actual terms counted, so any figure on the site can be checked against the report it came from. A model would read nuance this cannot, and would not be able to show you why it decided anything.
What is actually read
PageNews does not hold publishers’ article text and does not fetch it. The analysis reads the headline and, where the publisher supplies one in their feed, their own short summary — usually one or two sentences. Every analysed report records which of these it had, and the event page says so.
- Headline only — roughly a dozen words. Confidence is capped well below full.
- Headline and the publisher’s summary — the usual case.
- Full article text — not used, because PageNews does not hold it.
Scores and confidence are separate numbers
Every measurement carries two figures. The score, from 0 to 100, is how strongly that kind of framing was present. The confidence, also a percentage, is how much text the score was read from and how many terms agreed. They are never combined: a score of 80 measured with 41% confidence is shown as exactly that, because folding the second number into the first would quietly turn uncertainty into a measurement.
What the measures mean
Framing measures apply to almost any report. Domain measures are calculated only where the subject genuinely warrants them — a football report gets no inflation score, and the absence is not recorded as a zero, because a zero would drag down an average and claim an observation that was never made.
- Conflict and cooperation — whether the report reached for the vocabulary of confrontation or of agreement.
- Policy, personality, institutional — whether the actor in the story is a measure, a named individual, or a body.
- Consequence, uncertainty, urgency — how far the report deals in outcomes, hedges and immediacy.
- Economic measures — inflation concern, consumer cost focus, market risk, employment and growth emphasis.
- International measures — escalation, diplomacy, national security and humanitarian emphasis.
Tone
Where a person or organisation is named, PageNews measures the tone of the language immediately around each mention, from −100 to +100. This describes one report’s word choice about that name. It is not a view of the person, it is not a judgement of accuracy, and it is not a description of the publisher.
What markets did afterwards
On events whose coverage is about markets, economies, business or security, PageNews shows how gold, Brent crude, the S&P 500 and the two crypto pairs moved in the hours after the story was first indexed here. Readings are taken every fifteen minutes and kept for thirty days.
This is a co-movement, not a cause. Prices move for many reasons at once, a price series cannot establish which one, and PageNews does not claim this story moved anything. Where a market was closed across the window the instrument is omitted rather than reported as unchanged, and moves too small to distinguish from ordinary drift are not shown at all.
What PageNews will not say
The analysis measures observable properties of text. It cannot establish whether a report is accurate, fair, or written in good faith, and it is never aggregated into a claim about a publisher’s politics. A high conflict score means a report used conflict vocabulary — and events genuinely are conflicts, so reporting one as such is correct, not evidence of anything.
You will not find labels such as “left-wing”, “biased”, or “pro-” or “anti-” anything on this site. Those are claims about intent, and nothing here measures intent.
When a comparison is shown at all
A measure appears on an event page only when at least two publishers were analysed on it, the scores differ by a meaningful margin, and confidence clears a floor. Six outlets scoring alike is a real finding about the event and a useless row in a comparison, so it is not shown.
Limitations, plainly
- One or two sentences of publisher summary is thin material, and the confidence figures reflect that honestly.
- Term lists are curated by hand and will miss vocabulary, particularly in specialist subjects.
- Irony, quotation and negation are not detected. A report quoting someone else’s aggressive language scores as aggressive.
- English only.
- The analysis can simply be wrong. Where it matters, read the original reports — every one is linked.
Versioning
Every stored analysis records which analyser, ruleset version and schema version produced it, along with a hash of the exact text read. When any of those change, the affected reports are analysed again, so a figure on the site reflects the current method rather than a historical one.
Related
See also our editorial policy, how the site works, and the publishers we collect from.