Category: Enterprise Risk Tags: risk intelligence, analytic tradecraft, sourcing, confidence levels, estimative language, screening, traceability

In 1951 a national intelligence estimate told the United States government that a Soviet attack on Yugoslavia should be considered a serious possibility. The estimate was carefully written and the analysts agreed on it. Some time later Sherman Kent, who chaired the board that produced these estimates, went around asking readers what odds they had taken "serious possibility" to mean. The answers ranged from roughly one chance in five to roughly four chances in five. Everyone had read the same sentence. Nobody had received the same message, and the spread between the readings was the difference between ignoring a threat and mobilizing against it. Kent wrote up the problem in an essay on estimative language that is still worth reading, and the intelligence community spent the following decades doing something risk management never has: standardizing how uncertainty is expressed, and writing the standard down where anyone can check it against the work.

I think about that essay every time I look at a risk register with a column that says High. High is our serious possibility. It is a word that survives every review because it is unfalsifiable, it means something different to the person who wrote it and the executive who reads it, and no one can reconstruct how it was reached. We have been running the 1951 experiment continuously for seventy-five years without ever running Kent's follow-up poll, which is the only part that would tell us what our reports actually communicate.

The Standard Is Published, and You Can Read It

What makes this more embarrassing than it needs to be is that the fix is not proprietary. The Office of the Director of National Intelligence publishes its analytic standards, in Intelligence Community Directive 203, as a public document. It is a few pages long. It requires analysis to describe the quality and credibility of the sources it rests on, to express and explain uncertainty rather than hiding it in adjectives, to distinguish clearly between the underlying reporting and the analyst's own assumptions and judgments, and to consider alternative explanations rather than presenting a single line as inevitable. There is a companion directive on sourcing. None of it is exotic. All of it is enforceable, because a reviewer can hold a finished product against the standard and see which requirement was skipped.

Now hold a normal enterprise risk report against those four requirements. Source quality is usually absent, because the source is a workshop. Uncertainty is not expressed at all, because the rating is a single integer with no interval and no confidence attached. The line between evidence and assumption is invisible, since the same cell holds both. And alternatives are not considered, because the format has room for one number. A risk report fails every one of the standards a much harder discipline wrote down and published, and the reason is not that risk practitioners are less rigorous people. It is that nobody ever made rigor a format requirement.

Secrecy About What You Know Is Not Secrecy About How You Know It

The organization that ran America's reconnaissance satellites managed to hold both ideas at once, which is the part commercial risk intelligence gets backwards. The National Reconnaissance Office was created in 1961 and its existence was not officially acknowledged until 1992. For three decades the fact of the agency was classified. And yet within the walls, the discipline was the opposite of vague: collection had to be tasked and accounted for, imagery carried its provenance, and analysis that reached a decision maker was expected to say what it rested on. When the Corona satellite imagery was declassified by executive order in 1995, the historical record that came out was not a pile of assertions. It was product with a lineage.

Commercial risk intelligence inverts that arrangement completely. The existence of the product is the most public thing about it, marketed hard, while the method is the secret. You are sold a country score of 3.4 and the composition of the 3.4 is proprietary, which means the one thing you cannot do with a number you are paying for is defend it. I wrote about that trade in risk intelligence you cannot audit, and the tradecraft comparison is what sharpens it: an agency that would not confirm it existed still documented its sources internally, because analysis without provenance is not intelligence, it is opinion with better clearance. A vendor that will not tell you what went into a rating has not protected a method. It has removed your ability to be accountable for using it.

Confidence Is Not Severity

The single most portable idea in the tradecraft is the separation of the judgment from the confidence in the judgment. An intelligence assessment says what it believes and, separately, how much weight the evidence bears: a high-confidence judgment rests on corroborated reporting from reliable sources, a low-confidence one is plausible but thin, and the reader is told which they are holding. Both can appear in the same product without embarrassment.

Risk registers collapse these two things into one cell, and the collapse destroys information in both directions. A rating of High derived from three years of loss data and a rating of High derived from one senior person's intuition are recorded identically and treated identically in every downstream decision, which flatters the guess and insults the analysis. Worse, the collapse hides where the program should invest next. If a rating carries low confidence, the right response is not to argue about the number, it is to go get better evidence, and you cannot manage that work if the register has no field that admits the number is soft. Adding a confidence dimension is the cheapest upgrade available to most risk programs. It costs one column and it makes the entire register honest about what it does not know, which is the precondition for the traceable scoring I argued for in the data-driven risk assessment.

Screening Is a Tier, Not a Weak Assessment

The other idea worth stealing is structural. Intelligence work distinguishes between broad collection that tells you where to look and finished analysis that tells you what to think. Wide surveillance is not a low-quality assessment. It is a different tier with a different job: cover everything cheaply, find the anomalies, and cue the expensive analytic capacity toward the handful of things that deserve it. Nobody is disappointed that a broad sweep did not produce a finished estimate, because it was never trying to.

Risk programs mostly lack that vocabulary, and the lack causes two opposite failures. Either a quick screen gets promoted into a decision it cannot carry, because a number appeared and numbers are persuasive, or a program refuses to screen at all and insists every location, vendor, or use case receives full assessment, which guarantees the assessments arrive late and the coverage stays partial. Naming the tier fixes both. A screening figure says this site, this vendor, this jurisdiction is worth a closer look, and it should be labeled loudly as what it is. That is why the score on our own map is called a screening number rather than an assessment, and why I would rather say that in the product than let a buyer discover the limit later. Deciding what triggers escalation from tier one to tier two is a policy question a program can actually answer: a threshold breach, a material change, a first entry into a jurisdiction, an asset above a criticality line.

What It Looks Like in a Risk Program

None of this requires adopting government terminology or pretending a GRC team is an intelligence shop. It requires four format changes, each of which is a column or a label rather than a project. Every figure carries its source and the date that source was current, so a reader can go look. Live data, dated snapshots, and human judgment are labeled distinctly rather than blended into one confident-looking output, because a stale snapshot presented as current is worse than no number at all. Confidence is recorded separately from severity. And the tier is stated, so a screening figure cannot quietly become an assessment somewhere downstream in a slide.

The reason to do this is not intellectual tidiness. It is that risk work, like intelligence work, is eventually read by someone who will act on it and later be asked why. Kent's whole argument was that an estimate which cannot be interpreted consistently has failed at its only job, no matter how carefully it was written. A risk rating nobody can reconstruct has failed the same way, and it fails silently, right up until the moment an examiner, a board member, or an incident asks the question that starts with why did you think. The discipline that had the most to lose from getting this wrong published its answer decades ago and left it where anyone can read it. There is no good reason our field still treats sourcing as optional.

PivotRisk is a practitioner-led governance, risk, and resilience practice. Everything published here comes out of programs actually designed, launched, and run inside enterprise software, fintech, and infrastructure companies, not frameworks summarized from a distance.

See sourcing applied to a live picture

The PivotRisk risk intelligence map labels every figure by tier: live feeds say Live, dated snapshots say Snapshot with the date, and editorial judgment says so. Nothing is fetched until you ask, every number names its publisher and links back, and the site score is called a screening number because that is what it is.

Open the Risk Map