Confidence Ladders vs. Gut Feel: Why "Probably True" Needs a Number

2026-08-11 · Philip Choo · AI9OS

"Probably true" is not a finding. It is a feeling with a job title — and it is unusable the moment anyone asks what it rests on. A confidence ladder replaces that phrase with a rung: a named level, with a stated evidential requirement, that a reader can check and disagree with. This post sets out the four rungs we use, what each one actually demands, and why the top one is deliberately the only thing on the platform a human must certify by hand.

The problem with the phrase

An investigator writing "probably" is usually doing something reasonable — signalling honest uncertainty. The trouble is what the word does downstream. It is not calibrated: your "probably" and your client's counsel's "probably" are not the same probability, and neither of you finds that out until the position has been taken. It is not auditable: nothing about it says what would have to be false for this to be wrong. And it is not comparable: a report with nine "probably"s gives a reader no way to tell which one is load-bearing.

A rung fixes all three, because a rung is a claim about evidence, not about how the investigator feels.

The four rungs

RAW — one source, no corroboration. The finding exists. That is the entire claim. A single social profile, one record, one screenshot. Raw is not an insult; most findings start here and a great many honestly stay here. What matters is that the report says so, rather than dressing a single sighting in confident prose.

CORROBORATED — two or more sources agree. Agreement, and nothing more. This rung carries a trap severe enough that we treat it as the ladder's weakest point: agreement between sources that are not actually independent is not corroboration, it is one source echoing. Three sites republishing the same wire copy are one source. This is the failure we have written about at length in the five distortions of OSINT verification — consensus is the distortion that most reliably fools careful people.

CROSS-VERIFIED — the sources are independent, and that independence has been checked. The step from corroborated to cross-verified is not "more sources". It is a deliberate act of tracing each source back toward its origin and asking whether they are genuinely separate — different collection methods, different custody, no common upstream. Most findings that feel confirmed are sitting here, and this rung is where the real work happens.

CONFIRMED — cross-verified, and it has survived a structured attempt to break it. Independence alone is not enough, because independent sources can be independently wrong. The finding must be challenged: what would have to be true for this to be false, and has anyone looked? Only then does it reach the top.

Why the top rung is human-gated on purpose

An engine can do most of this. It can group evidence, spot contradictions, notice that two sources share a domain, and propose that a finding move up. Our platform does all of that.

What it does not do — by design, not by limitation — is certify a finding as CONFIRMED. That step is reserved for a licensed human, and the refusal is structural rather than a setting someone can switch off.

The reason is narrow and worth stating exactly. An automated promotion to CONFIRMED would be a machine asserting that it has adequately imagined the ways it might be wrong. That is precisely the capability it does not have. Everything below the top rung is a statement about what was collected, which software can establish. The top rung is a statement about what was considered and rejected, which it cannot. Automating that step would not accelerate the judgment; it would delete it while leaving the label behind — the same overreach that separates a platform that speeds up collection from one that quietly inflates conclusions.

What this looks like in a report

Rungs travel with findings, not in a methodology appendix nobody reads. Each item carries its rung, the sources behind it, and — for anything above RAW — the independence check that justified the climb. Counsel can then do something they usually cannot: read the exhibit list and see immediately which findings would survive challenge and which are one contradiction away from collapsing.

The most useful consequence is one clients do not expect. A ladder makes it cheap to say a finding did not climb. When the only vocabulary available is confident prose, an investigator under pressure to deliver has an incentive to write "probably". When the vocabulary is a rung, leaving something at RAW is an ordinary, defensible act rather than an admission of failure — and the pressure to overstate goes away. That is the same instinct behind our case-acceptance gate: the professional answer is often the smaller one.

The test to apply to any report, including ours

Pick the finding the conclusion depends on most and ask three questions. What rung is it on? Which sources put it there, and has anyone checked they are independent of each other? What would have to be false for it to be wrong?

If the report cannot answer all three, it is not a confidence problem. It is a record problem — and the record is what you would have needed if the matter had gone anywhere.

General information for practitioners, not legal advice. AI9OS is an open-source-intelligence technology platform; investigation services are conducted solely by licensed agencies under Singapore's Private Security Industry Act.

AI9OS turns public information into verified, chain-of-custody findings for licensed investigation agencies, law firms and corporate risk teams.

Request a demo