NEWSNYOUSEE HOW IT RIPPLES
METHODOLOGY

How we make our numbers

Every score on this site is either measured — traceable to a published source you can open — or editorial — a judgment we make and own. This page tells you which is which, for every number we show.

A score you can't interrogate is decoration. Every analysis here is built from sources that were fetched and read, and every figure in one carries a citation you can follow. Where a number is our judgment rather than someone else's measurement, we say so on the page it appears. Nothing here is written to make our numbers look more rigorous than they are.

The two buckets

Confusing these two is how a trust mechanic collapses. A reader who discovers that a precise-looking number was actually a guess stops believing the numbers that weren't.

MEASURED

Someone else published it

Gas flows, energy prices, trade volumes, industrial output. These exist independently of us. They get fetched, dated and cited — never typed in by hand. If we can't cite it, we don't call it measured.

EDITORIAL

We made the call

Impact Score, Trust Index, Confidence. No institution publishes these — they're our reading of the evidence. That's the product, not a defect. But it means they're only as good as the method behind them, so the method is public.

No data pipeline will ever make an editorial score "real." There is no API that returns an Impact Score. Wiring up live data sources will fix our measured figures; it will not turn a judgment into a measurement. Anyone promising otherwise is selling you a number.

Where each metric stands today

Every analysis on this site is now sourced: each figure in one is taken from a citation you can follow, and the Trust Index is counted from those citations rather than typed. Region health on five hubs — Russia, Europe, the Middle East, North America and East Asia — is an assessed read: still editorial, because no source publishes those numbers, but each states its reasoning and lists the indicators behind it. The other five hubs carry placeholder figures, flagged as placeholders directly on the numbers rather than passed off as considered. This table is the honest inventory:

MetricBucketStatusWhat it actually is right now
Influence ScoreEditorialCOMPUTEDA defined formula over six hand-entered dimension ratings. The arithmetic is real and reproducible; the inputs are our estimates.
Impact ScoreEditorialCOMPUTEDA defined formula over five hand-entered dimension ratings. The arithmetic is real and reproducible; the inputs are our estimates.
Trust IndexMeasuredCOMPUTEDCounted from the cited sources: agreement = supports ÷ (supports + disputes). No citations means no Trust Index. All eight analyses carry one, from 38 citations in total — every source behind them was fetched and read, and you can follow each.
ConfidenceEditorialJUDGMENTA hand-assigned label — Confirmed, Likely or Uncertain — attached to each claim in the Five-Question Framework. It reads the cited evidence but it is not counted from it, and it never will be: no source publishes how sure we should be.
Region healthEditorialASSESSED · 5 OF 10Stability, Economy and Conflict Risk, scored 0–100 against what is normal for that region. Russia, Europe, the Middle East, North America and East Asia carry a considered read: the score is ours, but each states its reasoning and links the indicators behind it. The other five hubs show placeholder figures, labelled as placeholders on the numbers — scaffolding, not a guess passed off as assessed.
Reader voteMeasuredCOMPUTEDYour own call, stored in your browser with the real date you made it. We show no reader aggregate — see below.
"Updated Xh ago"MeasuredCOMPUTEDA real ISO instant recording when the content was written, rendered relative to your clock at view time. Hover any timestamp for the exact date.

How an analysis gets sourced

An analysis here is a traced causal chain, not a scenario we found plausible. Before a page can call itself sourced it goes through six stages, and the two that take the time are reading and being wrong:

StageWhat happens
PickA causal chain worth tracing. If we can't state the chain in one line, it isn't ready to write.
SourceFind and read real sources, deliberately including ones that cut against the chain. Each becomes a citation with a real publisher, URL and date. This is where the hours go.
DraftAnswer the framework questions twice — once plainly, once in full. Same claim, two altitudes.
ScoreImpact dimensions, confidence per question, probability estimates. Judgments, made deliberately rather than to fill the panel.
CheckRead the rendered page and follow every source link on it.
ShipPublish, then keep revisiting the open estimates as events move.

That is roughly six to nine hours per analysis. We publish the figure because it's what makes a cadence claim honest: at that cost, one analysis a fortnight is sustainable and four in a launch week is a promise we'd break.

The sourcing is not a formality, and here is the proof. Three of the first five analyses on this site had their causal chain contradicted by their own sources once we went and read them. The gas piece ran the wrong way round entirely — it had Russia cutting supply to Europe, when the record shows the EU legislating the phase-out itself, with dated deadlines. The migration piece had displacement driving policy to harden; the numbers were falling while the regime hardened anyway. The sanctions piece asserted two steps that its own data denied, one of which no source supported at all. Each was rewritten to match the record and the Impact Score moved with it.

We're telling you this because it's the strongest thing we can say about the method: the intuitive chain — the one that sounds right before you check — was wrong more often than it was right. Analysis that never surprises its author isn't analysis. It's the reader's job to catch the fourth one, and the source list is there so you can.

Influence Score

Six dimensions of national power, each rated 0–100, combined into one composite. The weights encode what we think power actually leans on — economic and military weight count for more than soft power:

Military × 1.2 Economic × 1.3 Diplomatic × 1.0 Technology × 1.1 Energy × 1.0 Soft Power × 0.9 composite = Σ(rating × weight) ÷ 6.5 (6.5 = sum of weights)

Worked example — Russia, rated [84, 42, 55, 58, 90, 45] across those six axes in order:

(84×1.2)+(42×1.3)+(55×1.0)+(58×1.1)+(90×1.0)+(45×0.9) = 100.8 + 54.6 + 55 + 63.8 + 90 + 40.5 = 404.7 ÷ 6.5 = 62.3 → 62

That's the whole method. You can check our arithmetic, and you should be able to argue with the weights — a different weighting is a different theory of power, which is a legitimate disagreement to have with us.

The inputs are estimates. The formula is honest; the six ratings feeding it are our editorial judgment, not measurements. Five of the nine actors we score (United States, China, Japan, Iran, Brazil) are reference nodes rather than regions with a hub of their own.
The Influence Meter rates actors, never regions. An actor is an entity with a single foreign policy — there is one thing to score. A region like Africa, East Asia or Central Asia is not: averaging 54 countries' militaries, or China's position with Japan's, produces a composite describing nobody on the map. That is a category error rather than merely an unsourced number, and labelling it "editorial" would not repair it. Regions that are arenas rather than actors get a Who competes here panel instead, which names the powers in play and what each is contesting. The scores on that panel are each actor's global rating — we do not have a figure for an actor's weight inside a given region, and we will not estimate one.

Impact Score

A 0–100 read on how far an event's consequences travel, across five dimensions: Geopolitical, Economic, Energy, Security and Humanitarian. Each is rated 0–100 and the headline is their unweighted mean:

composite = Σ(rating) ÷ 5

Worked example — the European gas analysis, rated [78, 71, 84, 62, 34] across Geopolitical, Economic, Energy, Security and Humanitarian:

(78 + 71 + 84 + 62 + 34) ÷ 5 = 329 ÷ 5 = 65.8 → 66

No dimension is privileged over another. Impact asks how far consequences travel, so weighting one domain above the rest would be a claim we can't currently defend — and an unweighted mean puts no thumb on the scale. If we later adopt weights, they'll be published here the way the Influence Score weights are, before they're applied.

What changed, and why you're reading about it. This headline once read 78 while the dimensions beside it averaged 73.6 — it had been typed in by hand rather than calculated, and each rating was separately hardcoded into the bar widths, so the two could drift apart silently. The ratings are now the single source of truth, and the headline, the ring and the bars are all derived from them: 78 → 74. Then the analysis was sourced, the reading of the event changed, and the ratings changed with it: 74 → 66. Both moves are the method working. We could have kept 78 by reverse-engineering weights that happened to produce it; that would have been picking the answer first and calling it a method.
The ratings are judgments about sourced facts, not about nothing. Impact asks a question no source answers — how far do the consequences travel — so the five numbers are ours. But the event underneath them isn't ours: it's the cited record. The gas analysis rates Energy at 84 because the EU wrote a Russian gas phase-out into law with dated deadlines, and Humanitarian at 34 because nothing in that record shows a humanitarian dimension worth more. When the record moved, the ratings moved. A rating that never moves when the sources move was never reading them.
The inputs are estimates. As with the Influence Score, the formula is honest but the five ratings feeding it are our editorial judgment, not measurements.

Trust Index

It answers one narrow question: how much do independent sources actually agree about the underlying facts? Not whether we're confident — that's the Confidence rating, and it's a different claim. Trust is about the evidence, not about us.

agreement = supports ÷ (supports + disputes) sources = how many we cited no citations → no Trust Index

It is counted from the source list at the bottom of each analysis — the same list you can click through. It is not a number anyone types; the build refuses to accept a hand-entered one, and it refuses to call an analysis sourced if nothing on it is cited.

Every analysis on the site now carries one. The spread is the useful part:

AnalysisTrustSourcesWhy it lands there
US–Iran military escalation83%6Five support, one disputes — the highest on the site, on the most-cited page.
Strait of Hormuz80%5Four sources support the reading, one disputes it.
Europe's Russian gas ban80%5Four support, one disputes.
Ukraine's fifth year80%5Four support, one disputes. Nothing here contests the casualty or territorial figures — the single dispute is CSIS on why the strike campaign escalated, which is exactly where that analysis marks itself Uncertain.
EU–China tech sovereignty60%5Three support, two dispute.
EU migration politics60%5Three support, two dispute.
Russia–China sanctions50%4Two support, two dispute. The sources genuinely split on what the trade figures mean, and the score is supposed to say so rather than round the disagreement away.
UK–India FTA and Russian oil33%3One support, two dispute — the lowest on the site, and correct. Three read sources beat five with padding, and a minority reading that we still think is right is precisely what this number is for.
A low Trust Index is not a broken one. 33% doesn't mean we one-third-believe the piece; it means the published record is evenly divided about the underlying facts, which is a thing worth knowing before you read our reading of them. If every analysis here scored 90% you should assume we were choosing sources rather than counting them.
It used to read 72%, and 72% was invented. The figure was hand-entered, and for a while paired with the names of real institutions that had never been consulted. A trust score with nothing behind it is worse than none, because it looks like it's counting something. It is now structurally impossible to type one: the number exists only if the citations do.

The rule the sourcing runs on: fetch the source, never cite a search result about it. It has caught real errors on this site. A summary claimed oil "surged to $120" during the Hormuz crisis; the article it was summarising said $76.58. Another offered a forecast price that would have been published here as an actual one. A third cited a sanctions designation to a page that predated it by four years. Every citation on this site was opened and read.

Confidence rating

Each claim in the Five-Question Framework carries one of three labels. These are deliberately coarse — a false precision like "71% confident" would imply arithmetic we haven't done:

LabelWhat we mean by it
ConfirmedThe underlying fact is directly observable and not seriously disputed. Disagreement is about what it means, not whether it happened.
LikelyOur reading of the evidence, which competent analysts could reasonably contest. Most causal claims live here.
UncertainWe're extrapolating. Treat as a hypothesis, not a finding.

These labels are assigned by hand. That's appropriate — a judgment shouldn't pretend to be a computation — but it does mean the label is only worth as much as our track record, which is the next section.

Timestamps

Every "updated" figure records a real instant — when that content was actually last written. The page stores the machine-readable date and derives the human phrase from it against your own clock, so it ages on its own. Hover any timestamp to see the exact date behind it.

These used to be frozen prose. "updated 2h ago" was typed into the page as a literal string. It was arguably true for one hour and false every hour afterwards — and by the time we replaced it, the content it described was four hours old, not two. A freshness claim that ages into a lie is worse than showing no date, because staleness is exactly what it purports to rule out.

The three connections also claimed to have been updated 2 hours, 1 day and 3 days ago respectively. They were all written within the same minute. The spread was decoration — the visual signature of a busy newsroom, with nothing behind it. They now carry the one real instant they share.

If we don't know when something changed, we won't show a date for it. An invented timestamp is the cheapest possible lie and the easiest one to get caught in.

Reader votes

Each scenario asks you to commit to a falsifiable call before reality answers. Your vote is stored in your own browser along with the date you made it, so when the scenario resolves you can see what you actually predicted rather than what you'd prefer to remember.

We don't show you what other readers think, and that's deliberate. This widget used to report that 61% of readers agree with you — a hardcoded number printed identically whether you voted Yes or No. It wasn't just invented; it was incoherent. 61% agreeing with the Yes voters and 61% agreeing with the No voters is 122% of readers. It told everyone the crowd was on their side, whatever they'd just clicked.

We have no vote store yet, so we have no aggregate to report, so we report none. When one exists, the counts shown will be the counts we actually have — never seeded with a plausible-looking starting number to make the feature feel alive.

There's a reason this one stung more than the other placeholders. A platform that exists to show you the whole board shouldn't run a widget whose only function is to tell you that you were right.

Track record

The honest answer to "how often are you right?", counted from the same data the analyses are built from — never typed in by hand:

0 RESOLVED PREDICTIONS

We haven't published a dated prediction that has since resolved. We have no track record, so we're not going to show you one.

How we'll be scored

We are committing to the scoring rule now, while the record is still empty. That ordering is the entire point — picking how you'll be judged after you can see your results is how every flattering track record ever gets built.

Brier score = mean( (probability we published − what happened) ² ) what happened: 1 if it occurred, 0 if it did not lower is better
ScoreWhat it means
0.00Perfect. Not going to happen.
0.25The bar. This is what you score by saying "50%" to everything. If we don't beat it, our estimates are worth nothing and you should ignore them.
1.00Confidently, maximally wrong.

We use Brier rather than a "% correct" hit-rate deliberately. A hit-rate lets you look brilliant by only ever predicting near-certainties, and it rounds a published 44% into a yes/no we never actually claimed. Brier punishes confident wrongness harder than hedged wrongness, which is exactly the incentive we want pointed at ourselves.

Why this section exists while it's empty. A track record link that quietly goes nowhere, or leads to a page of selectively remembered wins, is worth less than admitting the number is zero. Every prediction resolves here — including the ones we get wrong.

Revisions are part of the record. An estimate that moves shows its trail on the analysis itself — "was 52% → now 44%" — because a revision that quietly overwrites the original isn't a correction, it's a rewrite.

Corrections

When we get something wrong, the fix is logged on the analysis itself rather than silently applied. A changed number without a changed-number notice is indistinguishable from never having been wrong, and a platform that's never visibly wrong is a platform nobody should trust.

This policy is easy to state while nothing is at stake and hard to keep when a call ages badly. Hold us to it.

Data layer, voice layer

The neutral data layer and the Founder's Lens are different things and always look different. The Lens is one person's opinion on its own cream panel, clearly marked. When we're reporting we're reporting; when we're arguing you'll see the panel change. If you ever can't tell which layer you're reading, that's a bug — tell us.