Crisis test · How we tested
Prometheus sees losses about as often as they happen.
We took 20 years between 2002 and 2022. For each, we rebuilt Prometheus the December before, using only the data that existed then, let it look a year ahead a million times, and checked where the real year landed.
At a glance
The checks Prometheus had to pass
Think of the test as an exam with four checks, each with a pass mark. Prometheus passed all four. The sections below explain each one.
- Pass mark
- A right model looks this uneven at least 5% of the time: a biggest gap D below 0.294
- Prometheus
- D = 0.145: a right model looks this uneven 74% of the time
- Pass mark
- The count inside the “if right” range, 0 to 4
- Prometheus
- 2 (2008 and 2022)
- Pass mark
- The count inside the “if right” range, 0 to 2
- Prometheus
- 1 (2022)
- Pass mark
- The count inside the “if right” range, 0 to 4
- Prometheus
- 2
Pass marks as this page applies them. The rules, written down and frozen before the test ran, judged the claim on the first three checks and reported the fourth beside them.
With a home. The same four checks also passed for the 60/40 with a US home worth half as much: a right model looks as uneven 84% of the time; losses of 10% or more, 1 seen against 0.54 expected (if right, 0 to 2); losses of 15% or more, none against 0.24 (0 to 1); 2 dots in the worst tenth (0 to 4).
The question
A forecast of chances can’t be judged on one year
Prometheus never says “next year the portfolio loses 14%”. It says how likely each outcome is: a loss this big about one year in ten, a loss that big one year in fifty. One year can’t prove or disprove a statement like that. A forecaster who says “10% chance of rain” isn’t wrong when it rains once.
What can be checked is the long run. If Prometheus says “one year in ten”, then over many years about one in ten should turn out that bad. Too many bad years and it is too rosy; too few and it is too gloomy. The test checks exactly this, on 20 real years, each forecast before it happened.
One year, three steps
How one year is tested
Rebuild
Only the data that existed in December 2007. Nothing later.
Look ahead
A million possible 2008s for a 60/40 portfolio, lined up from worst to best.
Score
The real 2008, January to December: the portfolio lost 14%. 2.3% of the futures ended lower, about 1 in 43.
The key idea
If Prometheus is right, the dots spread evenly
Suppose the million futures Prometheus imagines are an honest picture of what can happen. Then the real year is just one more draw from the same pile, like shuffling a card back into a deck and seeing where it ends up. It is as likely to land in the worst tenth of the line as in the middle tenth or the best tenth. So over many years, roughly a tenth of the dots should land in each tenth of the line: evenly spread.
When a model is wrong, the dots show it, and each kind of mistake leaves its own pattern:
Swings too small
Real years keep falling outside its range, so the dots pile up at both ends.
Too rosy
Real years keep disappointing it, so the dots pile up at the worst end.
Swings too big
Real years look ordinary to it, so the dots crowd the middle.
Our twenty years
Prometheus’s real dots, and what “even” would look like
Each dot below is one real year from 2002 to 2022: the 60/40’s return that calendar year, ranked against the futures Prometheus imagined the December before, as in the three steps above (2010 could not be tested: Prometheus could not be rebuilt in December 2009). The small ticks under the line show where twenty perfectly even dots would sit. Real dots never sit exactly there, because luck always moves them around a little. The question is whether they wander further than luck allows.
Where each real year landed among Prometheus’s futures
2008 · the portfolio lost 14%; 2.3% of the futures ended lower, about 1 in 43
- The worst tenth
- Where twenty even dots would sit
- A real year: point at one to read it
Every year in numbers
| Year | Portfolio | Futures that ended lower |
|---|---|---|
| 2002 | −6.0% | 11% (about 1 in 9) |
| 2003 | +17.8% | 68% |
| 2004 | +8.8% | 46% |
| 2005 | +4.7% | 37% |
| 2006 | +10.1% | 50% |
| 2007 | +7.6% | 45% |
| 2008 | −14.3% | 2.3% (about 1 in 43) |
| 2009 | +12.1% | 63% |
| 2010 | Not tested | December 2009 could not be rebuilt |
| 2011 | +9.2% | 50% |
| 2012 | +10.9% | 62% |
| 2013 | +15.4% | 77% |
| 2014 | +13.5% | 67% |
| 2015 | +1.3% | 31% |
| 2016 | +7.5% | 44% |
| 2017 | +14.5% | 73% |
| 2018 | −2.7% | 21% (about 1 in 5) |
| 2019 | +23.0% | 88% |
| 2020 | +16.2% | 81% |
| 2021 | +15.0% | 76% |
| 2022 | −18.6% | 2.3% (about 1 in 43) |
A simple way to judge them is to pick a zone of the line, count the real dots in it, and ask whether a right model could easily have produced that count. Take the worst tenth. If Prometheus is right, each year has a 10% chance of landing there, so twenty years should put about two dots in it. Luck moves that around: 96 times in 100, a right model gives anywhere from 0 to 4. Call that the “if right” range. We saw 2, inside it, so no warning. Seeing 5 or more would mean bad years turn up more often than Prometheus says.
It works like checking a die. Roll it 60 times and expect ten sixes; 7 or 13 is just luck, but 25 means the die is loaded. The “if right” range marks where luck stops being a believable explanation. The four boxes do the same check on four zones:
If right: 0 to 3 · inside
If right: 0 to 4 · inside
If right: 1 to 7 · inside
If right: 6 to 14 · inside
Where do the “if right” ranges come from?
Suppose Prometheus is exactly right. Then each year has a 10% chance of landing in the worst tenth, whatever the other years did. Most of the time that puts about two dots there, but fewer or more turn up too, just by luck. The chance of each count can be worked out exactly, and the chart shows it. The “if right” range leaves out only the rarest counts at each end, never more than 5% of the chances at either end, so a right model’s count lands inside it at least nine times in ten. Here nothing is left out at the low end, because even no dots at all turns up 12% of the time, so 0 to 4 holds 96% of the chances.
Dots in the worst tenth, if Prometheus is right
The evenness check
One number for “how uneven”: the biggest gap
The four boxes above check four zones. The evenness check looks at every zone at once. For each zone size from 0% to 100%, counted from the bad end, it asks one question: what share of the twenty real years landed in the worst X% of Prometheus’s futures? Plotted against X, that share makes the staircase below.
- Across (x)
- The zone size, counted from the bad end. 10% is the worst tenth, 50% is the bottom half, 100% is the whole line.
- Up (y)
- The share of the 20 real years inside that zone. It steps up by 5% (one year in twenty) each time the zone grows past a real dot.
- Dashed diagonal
- What a right model gives on average: the worst 10% holds 10% of the years, the worst 30% holds 30%, and so on.
The staircase of real years
- Real years: the staircase
- A right model, on average
- The biggest gap, D
Reading the staircase at a few points:
| Zone (x) | Real years in it (y) | A right model, on average | Gap |
|---|---|---|---|
| Worst 10%as a box above | 2 of 20 = 10% | 10% (2 years) | none |
| Worst 20%as a box above | 3 of 20 = 15% | 20% (4 years) | 5 points fewer |
| Worst 30% | 4 of 20 = 20% | 30% (6 years) | 10 points fewer |
| Worst 50%as a box above | 10 of 20 = 50% | 50% (10 years) | none |
| Worst 80.5% | 19 of 20 = 95% | 80.5% (16.1 years) | 14.5 points more, the biggest |
| Everything | 20 of 20 = 100% | 100% (20 years) | none |
Reading the shape. Where the staircase runs above the diagonal, more real years landed in that zone than a right model gives: the dots bunch toward the bad end, so the model is too rosy there. Where it runs below, fewer landed: the model is too gloomy there. A staircase that stays close to the diagonal all the way means the dots are even, as a right model gives.
The check keeps only one number: the biggest gap between the staircase and the diagonal, called D. Here D = 0.145, a gap of 14.5 percentage points, or about three years out of twenty. Is that a lot? With only twenty dots, some gap always appears by luck. Throw twenty truly random dots, measure their biggest gap, and repeat many times: some gaps come out small, some large. This curve shows how often each size turns up for twenty dots:
The biggest gap of twenty dots from a right model
That shaded share, 74%, is the evenness check’s p-value. Read it as: “if Prometheus were exactly right, twenty years would look at least this uneven 74% of the time”. Nothing unusual, so it passes. The pass mark is 5%: below it, the dots would be more uneven than a right model produces one time in twenty, and the model would be called wrong. Two warnings. The p-value is not the chance that Prometheus is right. And a high p-value doesn’t prove it right either: it means the twenty years gave no evidence against it.
Big losses
Losses expected and losses seen
The second check looks only at big losses. It asks one question: over the twenty years, did Prometheus expect about as many years with a loss of 10% or more as really happened? It takes three steps.
Step 1Each December, a chance of a big loss
Each December, the million futures for the year ahead also give a chance of losing 10% or more: the share of them that lose that much. It is counted the same way as the dots. In December 2007, for example, 58,085 of the 1,000,000 futures for 2008 lost 10% or more, so the chance Prometheus gave for 2008 was 5.8%. Each bar below is that chance for one year.
Prometheus’s chance of a loss of 10% or more, each December
Step 2Add the chances to get the number expected
Chances add up to a number of years. It works like a weather forecast: say “10% chance of rain” on each of twenty days and expect 20 × 10% = 2 rainy days. When the chance changes from day to day, add each day’s chance instead of multiplying. Here are Prometheus’s twenty chances, one per year, added up:
- 20024.6%
- 20034.7%
- 20043.9%
- 20054.3%
- 20065.4%
- 20075.4%
- 20085.8%
- 200910.6%
- 20118.7%
- 201210.7%
- 201311.0%
- 20148.9%
- 20159.6%
- 201610.2%
- 20179.4%
- 20188.4%
- 20198.6%
- 20208.8%
- 20218.6%
- 20228.0%
Add all twenty: 155.6%. A total of 155.6% means 1.56 years, so Prometheus expected 1.56 years with a loss of 10% or more. What really happened: 2 such years, 2008 (lost 14.3%) and 2022 (lost 18.6%).
Step 3Is the gap between expected and seen just luck?
Prometheus expected 1.56 big-loss years and 2 happened. Is that close enough, or a sign Prometheus is wrong? It can’t be judged by eye, because luck plays a part. So suppose Prometheus is exactly right, and ask how often each count of big-loss years would turn up.
Each year is then like a die weighted to come up “big loss” with the chance Prometheus gave that year: 5.8% for 2008, 10.6% for 2009, and so on. Even with Prometheus exactly right, the count in twenty years varies, just by luck. The chance of each count can be worked out exactly from the twenty chances; the first chart below shows it.
Real history gave 2, and 2 turns up 27% of the time when Prometheus is right, one of the most common results. So seeing 2 is no evidence against it. 5 or more would be suspicious: that happens only 2% of the time, so seeing it in real history would suggest Prometheus underestimates big losses. The “if right” range, 0 to 4, is made as the worst tenth’s was, and holds 98% of the chances. Anything inside it is ordinary luck.
Losses of 10% or more
How often each count turns up if Prometheus is right, worked out exactly
Losses of 15% or more
How often each count turns up if Prometheus is right, worked out exactly
The 15% chart works the same way with a stricter line. Its chances are smaller, about 2% a year up to 2008 and 3% to 4% after, and they add up to 0.62. Only 2022 lost 15% or more. 2008 lost 14.3%, just short, so it counts as a 10% loss but not a 15% one; the test scores calendar years, and 2008’s worst twelve months, which ran into the next year, lost more. Why is each range lopsided around the expected count? A count can’t go below zero: with 1.56 expected, zero is common and four is possible.
A standard model
A standard bell curve passes too
The same years, judged by a standard bell curve: one for each asset, combined by how the assets moved together, fitted on every year since 1970. It is the comparison the test was built for.
- Prometheus
- Standard bell curve
| Check | Prometheus | Standard bell curve |
|---|---|---|
| Evenness: how often a right model looks this uneven | 74.4% · pass | 9.5% · pass |
| Dots in the worst tenth | 2 (if right 0 to 4) | 3 (if right 0 to 4) |
| Losses of 10% or more: expected · seen | 1.56 · 2 (if right 0 to 4) · inside | 0.91 · 2 (if right 0 to 3) · inside |
| Losses of 15% or more: expected · seen | 0.62 · 1 (if right 0 to 2) · inside | 0.36 · 1 (if right 0 to 2) · inside |
Both pass every check, so this test can’t rank them: it asks whether a model is honest about how often bad years come, and both are, as far as twenty years can tell. The difference that shows: in 20 of 20 years the bell curve’s dot sits lower than Prometheus’s, because it put more of its futures above what really happened. It was a little rosy, and expected fewer big losses (0.91) than turned up (2). With twenty years that is not enough to fail it. That is why the claim is “Prometheus sees losses about as often as they happen”, and never “and a standard model does not”.
Fair test
What made it fair
Rules first
Written down and frozen on 7 October 2026, before this test was run.
Every year
Every December from 2001 to 2021 was tried. None picked, none dropped once a result was seen.
Only the past
Each rebuild learned only from history up to its December: American markets and the economic history of 18 countries.
Checked twice
Automated reviews of the code before any score was read, and every number recomputed separately after.
Nothing hidden
All 20 years are on this page, good and bad.
Bonds earn interest
In every future and every real year, bonds earn a year’s interest as well as their price change. Added on 9 October, after the first scoring; the claim held before and after.
Limits
What it does not show
That Prometheus is exactly right
Passing means no big mistake showed up in twenty years. It does not prove Prometheus exactly right: a smaller mistake can hide in so few years.
That Prometheus beats a standard model
Two standard bell curves passed the same test: one fitted on every year since 1970, and the crisis test’s own. A test both pass cannot rank them, so this shows Prometheus sees losses about as often as they happen, not that it sees them better.
Years it had seen before
Prometheus was designed in 2026, and changes to it were checked on 19 of these 20 years, 2008 among them, with other portfolios. Fixing this test’s rules first cannot undo that hindsight: it is a bit like a student who has seen most of the exam questions beforehand.
Every country and every era
One country, the United States, over 20 years, with only two losses of 10% or more.
2010
December 2009 could not be rebuilt, so 2010 was not tested.
Crashes that span two years
Each year runs January to December, so a crash that spans a year end is split across two: 2008’s worst twelve months ran into 2009 and lost more than calendar 2008 did. The crisis test scores each crash on its own twelve months instead.
The maths
Every formula behind the pictures
Landing spot, and why it should be even
landing spot u =(futures at or below the real year)/ (all futures)If the real year R is drawn from themodel’s own distribution F, thenu = F(R), and for any a from 0 to 1:P(u ≤ a) = P(F(R) ≤ a)= P(R ≤ F⁻¹(a))= F(F⁻¹(a)) = awhich is the definition of an even(uniform) spread.
This is the probability integral transform. It needs nothing about the shape of F: it works for fat tails, skews and crashes alike, which is why it is the standard check of a probability forecast.
The evenness check (Kolmogorov–Smirnov)
sort the n landing spots:u(1) ≤ u(2) ≤ … ≤ u(n)staircase:F̂(x) = (number of u ≤ x) / ndiagonal:xD = the biggest |F̂(x) − x|over every x= max over i ofmax(i/n − u(i), u(i) − (i−1)/n)p = the chance that n dots spreadevenly (uniformly) give a gapof D or morepass if p ≥ 5%;for n = 20 that means D below 0.294
The p-value is computed exactly for twenty dots (the Kolmogorov distribution). It is not the chance that Prometheus is right; it is how surprising this much unevenness would be if it were.
Counting dots in part of the line
K = how many dots land in a partof the line covering a share qK ~ Binomial(n, q)P(K = k) = C(n, k) · q^k· (1 − q)^(n − k)“if right” range: fromthe first k with P(K ≤ k) ≥ 5%to the first k with P(K ≤ k) ≥ 95%
Losses expected and the “if right” range
p(t) = the share of the model’sfutures for year t thatlose 10% or moreexpected = p(1) + p(2) + … + p(n)(linearity of expectation)K = how many years bring such aloss, each year with chance p(t),independently of the othersits chances, multiplied out:start with [1]; for each year,multiply by ((1 − p(t)) + p(t)·x);P(K = k) is the coefficient of x^k(the Poisson-binomial)“if right” range: fromthe first k with P(K ≤ k) ≥ 5%to the first k with P(K ≤ k) ≥ 95%
The expected count needs no assumption at all. The range treats the years as independent, which is how the test was written before it ran.
Rebuilt using only the data available each December. Hypothetical, not investment performance.
Each “if right” range runs from the count at which a right model’s chances first add up to 5% to the count at which they reach 95%, so a right model’s count lands inside it at least nine times in ten.