Crisis test · How we tested

Prometheus sees losses about as often as they happen.

We took 20 years between 2002 and 2022. For each, we rebuilt Prometheus the December before, using only the data that existed then, let it look a year ahead a million times, and checked where the real year landed.

At a glance

The checks Prometheus had to pass

Think of the test as an exam with four checks, each with a pass mark. Prometheus passed all four. The sections below explain each one.

Pass marks as this page applies them. The rules, written down and frozen before the test ran, judged the claim on the first three checks and reported the fourth beside them.

With a home. The same four checks also passed for the 60/40 with a US home worth half as much: a right model looks as uneven 84% of the time; losses of 10% or more, 1 seen against 0.54 expected (if right, 0 to 2); losses of 15% or more, none against 0.24 (0 to 1); 2 dots in the worst tenth (0 to 4).

The question

A forecast of chances can’t be judged on one year

Prometheus never says “next year the portfolio loses 14%”. It says how likely each outcome is: a loss this big about one year in ten, a loss that big one year in fifty. One year can’t prove or disprove a statement like that. A forecaster who says “10% chance of rain” isn’t wrong when it rains once.

What can be checked is the long run. If Prometheus says “one year in ten”, then over many years about one in ten should turn out that bad. Too many bad years and it is too rosy; too few and it is too gloomy. The test checks exactly this, on 20 real years, each forecast before it happened.

One year, three steps

How one year is tested

  1. Rebuild

    Only the data that existed in December 2007. Nothing later.

  2. Look ahead

    A million possible 2008s for a 60/40 portfolio, lined up from worst to best.

  3. Score

    The real 2008, January to December: the portfolio lost 14%. 2.3% of the futures ended lower, about 1 in 43.

The key idea

If Prometheus is right, the dots spread evenly

Suppose the million futures Prometheus imagines are an honest picture of what can happen. Then the real year is just one more draw from the same pile, like shuffling a card back into a deck and seeing where it ends up. It is as likely to land in the worst tenth of the line as in the middle tenth or the best tenth. So over many years, roughly a tenth of the dots should land in each tenth of the line: evenly spread.

When a model is wrong, the dots show it, and each kind of mistake leaves its own pattern:

  • Swings too small

    Real years keep falling outside its range, so the dots pile up at both ends.

  • Too rosy

    Real years keep disappointing it, so the dots pile up at the worst end.

  • Swings too big

    Real years look ordinary to it, so the dots crowd the middle.

Our twenty years

Prometheus’s real dots, and what “even” would look like

Each dot below is one real year from 2002 to 2022: the 60/40’s return that calendar year, ranked against the futures Prometheus imagined the December before, as in the three steps above (2010 could not be tested: Prometheus could not be rebuilt in December 2009). The small ticks under the line show where twenty perfectly even dots would sit. Real dots never sit exactly there, because luck always moves them around a little. The question is whether they wander further than luck allows.

Where each real year landed among Prometheus’s futures

2008 · the portfolio lost 14%; 2.3% of the futures ended lower, about 1 in 43

  • The worst tenth
  • Where twenty even dots would sit
  • A real year: point at one to read it
The 60/40: 60% US shares (S&P 500), 40% ten-year zero-coupon US Treasuries.
Every year in numbers
YearPortfolioFutures that ended lower
2002−6.0%11% (about 1 in 9)
2003+17.8%68%
2004+8.8%46%
2005+4.7%37%
2006+10.1%50%
2007+7.6%45%
2008−14.3%2.3% (about 1 in 43)
2009+12.1%63%
2010Not testedDecember 2009 could not be rebuilt
2011+9.2%50%
2012+10.9%62%
2013+15.4%77%
2014+13.5%67%
2015+1.3%31%
2016+7.5%44%
2017+14.5%73%
2018−2.7%21% (about 1 in 5)
2019+23.0%88%
2020+16.2%81%
2021+15.0%76%
2022−18.6%2.3% (about 1 in 43)

A simple way to judge them is to pick a zone of the line, count the real dots in it, and ask whether a right model could easily have produced that count. Take the worst tenth. If Prometheus is right, each year has a 10% chance of landing there, so twenty years should put about two dots in it. Luck moves that around: 96 times in 100, a right model gives anywhere from 0 to 4. Call that the “if right” range. We saw 2, inside it, so no warning. Seeing 5 or more would mean bad years turn up more often than Prometheus says.

It works like checking a die. Roll it 60 times and expect ten sixes; 7 or 13 is just luck, but 25 means the die is loaded. The “if right” range marks where luck stops being a believable explanation. The four boxes do the same check on four zones:

In the worst 5%2 of 20Expected 1
If right: 0 to 3 · inside
In the worst tenth2 of 20Expected 2
If right: 0 to 4 · inside
In the worst fifth3 of 20Expected 4
If right: 1 to 7 · inside
Below the middle10 of 20Expected 10
If right: 6 to 14 · inside

Where do the “if right” ranges come from?

Suppose Prometheus is exactly right. Then each year has a 10% chance of landing in the worst tenth, whatever the other years did. Most of the time that puts about two dots there, but fewer or more turn up too, just by luck. The chance of each count can be worked out exactly, and the chart shows it. The “if right” range leaves out only the rarest counts at each end, never more than 5% of the chances at either end, so a right model’s count lands inside it at least nine times in ten. Here nothing is left out at the low end, because even no dots at all turns up 12% of the time, so 0 to 4 holds 96% of the chances.

Dots in the worst tenth, if Prometheus is right

If right: 0 to 412%027%129%219%39%43%51%6<1%7<1%8+dots in the worst tenth, out of twenty
If Prometheus is right, the worst tenth gets no dots 12% of the time, 2 dots 29% of the time, and 5 or more only 4% of the time. Real history gave 2: the lilac bar. The other boxes’ ranges come the same way, with a 5%, 20% or 50% chance per year instead of 10%. One thing does stand out: no year landed in the best tenth, which happens 12% of the time (0.9 to the power 20).

The evenness check

One number for “how uneven”: the biggest gap

The four boxes above check four zones. The evenness check looks at every zone at once. For each zone size from 0% to 100%, counted from the bad end, it asks one question: what share of the twenty real years landed in the worst X% of Prometheus’s futures? Plotted against X, that share makes the staircase below.

Across (x)
The zone size, counted from the bad end. 10% is the worst tenth, 50% is the bottom half, 100% is the whole line.
Up (y)
The share of the 20 real years inside that zone. It steps up by 5% (one year in twenty) each time the zone grows past a real dot.
Dashed diagonal
What a right model gives on average: the worst 10% holds 10% of the years, the worst 30% holds 30%, and so on.

The staircase of real years

0%0%20%20%40%40%60%60%80%80%100%100%The worst X% of the futuresShare of the 20 real yearsD = 0.145
  • Real years: the staircase
  • A right model, on average
  • The biggest gap, D
The biggest gap is at 80.5% along: 19 of the 20 real years landed in the worst 80.5% of Prometheus’s futures, where even dots would put 16.1. That gap of 14.5 points is D = 0.145. It comes from the empty best end: no year landed in Prometheus’s best 11.7%; the highest dot is 2019 at 88.3%.

Reading the staircase at a few points:

Zone (x)Real years in it (y)A right model, on averageGap
Worst 10%as a box above2 of 20 = 10%10% (2 years)none
Worst 20%as a box above3 of 20 = 15%20% (4 years)5 points fewer
Worst 30%4 of 20 = 20%30% (6 years)10 points fewer
Worst 50%as a box above10 of 20 = 50%50% (10 years)none
Worst 80.5%19 of 20 = 95%80.5% (16.1 years)14.5 points more, the biggest
Everything20 of 20 = 100%100% (20 years)none

Reading the shape. Where the staircase runs above the diagonal, more real years landed in that zone than a right model gives: the dots bunch toward the bad end, so the model is too rosy there. Where it runs below, fewer landed: the model is too gloomy there. A staircase that stays close to the diagonal all the way means the dots are even, as a right model gives.

The check keeps only one number: the biggest gap between the staircase and the diagonal, called D. Here D = 0.145, a gap of 14.5 percentage points, or about three years out of twenty. Is that a lot? With only twenty dots, some gap always appears by luck. Throw twenty truly random dots, measure their biggest gap, and repeat many times: some gaps come out small, some large. This curve shows how often each size turns up for twenty dots:

The biggest gap of twenty dots from a right model

00.10.20.30.40.5the biggest gap, D74%Prometheus: 0.145Fail line: 0.294
The curve is exact for twenty dots. Prometheus’s gap of 0.145 sits in the busy middle; the shaded part, everything to its right, is 74% of the area. A model would fail only if its gap passed 0.294, a size that right models reach just 5% of the time.

That shaded share, 74%, is the evenness check’s p-value. Read it as: “if Prometheus were exactly right, twenty years would look at least this uneven 74% of the time”. Nothing unusual, so it passes. The pass mark is 5%: below it, the dots would be more uneven than a right model produces one time in twenty, and the model would be called wrong. Two warnings. The p-value is not the chance that Prometheus is right. And a high p-value doesn’t prove it right either: it means the twenty years gave no evidence against it.

Big losses

Losses expected and losses seen

The second check looks only at big losses. It asks one question: over the twenty years, did Prometheus expect about as many years with a loss of 10% or more as really happened? It takes three steps.

Step 1Each December, a chance of a big loss

Each December, the million futures for the year ahead also give a chance of losing 10% or more: the share of them that lose that much. It is counted the same way as the dots. In December 2007, for example, 58,085 of the 1,000,000 futures for 2008 lost 10% or more, so the chance Prometheus gave for 2008 was 5.8%. Each bar below is that chance for one year.

Prometheus’s chance of a loss of 10% or more, each December

0%5%10%4.620024.720033.920044.320055.420065.420075.8200810.620098.7201110.7201211.020138.920149.6201510.220169.420178.420188.620198.820208.620218.02022
Each bar is the chance Prometheus gave, in December, of the 60/40 losing 10% or more in the year ahead. Lilac bars are the years it really happened: 2008 (lost 14.3%) and 2022 (lost 18.6%). The colour only marks those years; every bar goes into the sum in step 2. The bars roughly double from December 2008 on: partly because the history Prometheus learned from then held the 2008 crash, and partly because bonds paid less interest from then on, leaving less to soften a fall.

Step 2Add the chances to get the number expected

Chances add up to a number of years. It works like a weather forecast: say “10% chance of rain” on each of twenty days and expect 20 × 10% = 2 rainy days. When the chance changes from day to day, add each day’s chance instead of multiplying. Here are Prometheus’s twenty chances, one per year, added up:

  1. 20024.6%
  2. 20034.7%
  3. 20043.9%
  4. 20054.3%
  5. 20065.4%
  6. 20075.4%
  7. 20085.8%
  8. 200910.6%
  9. 20118.7%
  10. 201210.7%
  11. 201311.0%
  12. 20148.9%
  13. 20159.6%
  14. 201610.2%
  15. 20179.4%
  16. 20188.4%
  17. 20198.6%
  18. 20208.8%
  19. 20218.6%
  20. 20228.0%

Add all twenty: 155.6%. A total of 155.6% means 1.56 years, so Prometheus expected 1.56 years with a loss of 10% or more. What really happened: 2 such years, 2008 (lost 14.3%) and 2022 (lost 18.6%).

Step 3Is the gap between expected and seen just luck?

Prometheus expected 1.56 big-loss years and 2 happened. Is that close enough, or a sign Prometheus is wrong? It can’t be judged by eye, because luck plays a part. So suppose Prometheus is exactly right, and ask how often each count of big-loss years would turn up.

Each year is then like a die weighted to come up “big loss” with the chance Prometheus gave that year: 5.8% for 2008, 10.6% for 2009, and so on. Even with Prometheus exactly right, the count in twenty years varies, just by luck. The chance of each count can be worked out exactly from the twenty chances; the first chart below shows it.

Real history gave 2, and 2 turns up 27% of the time when Prometheus is right, one of the most common results. So seeing 2 is no evidence against it. 5 or more would be suspicious: that happens only 2% of the time, so seeing it in real history would suggest Prometheus underestimates big losses. The “if right” range, 0 to 4, is made as the worst tenth’s was, and holds 98% of the chances. Anything inside it is ordinary luck.

Losses of 10% or more

How often each count turns up if Prometheus is right, worked out exactly

If right: 0 to 420%033%127%214%35%41%5<1%6+years with such a loss
Expected 1.56, seen 2 (2008 and 2022). If Prometheus is right, 98 times in 100 the count lands between 0 and 4, and 2 or more turns up 47% of the time. The lilac bar is what really happened. Inside the range: passes.

Losses of 15% or more

How often each count turns up if Prometheus is right, worked out exactly

If right: 0 to 253%034%110%22%3<1%4<1%5<1%6+years with such a loss
Expected 0.62, seen 1 (2022). If Prometheus is right, 98 times in 100 the count lands between 0 and 2, and 1 or more turns up 47% of the time. The lilac bar is what really happened. Inside the range: passes.

The 15% chart works the same way with a stricter line. Its chances are smaller, about 2% a year up to 2008 and 3% to 4% after, and they add up to 0.62. Only 2022 lost 15% or more. 2008 lost 14.3%, just short, so it counts as a 10% loss but not a 15% one; the test scores calendar years, and 2008’s worst twelve months, which ran into the next year, lost more. Why is each range lopsided around the expected count? A count can’t go below zero: with 1.56 expected, zero is common and four is possible.

A standard model

A standard bell curve passes too

The same years, judged by a standard bell curve: one for each asset, combined by how the assets moved together, fitted on every year since 1970. It is the comparison the test was built for.

PrometheusStandard bell curveWorstMiddleBest
  • Prometheus
  • Standard bell curve
CheckPrometheusStandard bell curve
Evenness: how often a right model looks this uneven74.4% · pass9.5% · pass
Dots in the worst tenth2 (if right 0 to 4)3 (if right 0 to 4)
Losses of 10% or more: expected · seen1.56 · 2 (if right 0 to 4) · inside0.91 · 2 (if right 0 to 3) · inside
Losses of 15% or more: expected · seen0.62 · 1 (if right 0 to 2) · inside0.36 · 1 (if right 0 to 2) · inside

Both pass every check, so this test can’t rank them: it asks whether a model is honest about how often bad years come, and both are, as far as twenty years can tell. The difference that shows: in 20 of 20 years the bell curve’s dot sits lower than Prometheus’s, because it put more of its futures above what really happened. It was a little rosy, and expected fewer big losses (0.91) than turned up (2). With twenty years that is not enough to fail it. That is why the claim is “Prometheus sees losses about as often as they happen”, and never “and a standard model does not”.

Fair test

What made it fair

  • Rules first

    Written down and frozen on 7 October 2026, before this test was run.

  • Every year

    Every December from 2001 to 2021 was tried. None picked, none dropped once a result was seen.

  • Only the past

    Each rebuild learned only from history up to its December: American markets and the economic history of 18 countries.

  • Checked twice

    Automated reviews of the code before any score was read, and every number recomputed separately after.

  • Nothing hidden

    All 20 years are on this page, good and bad.

  • Bonds earn interest

    In every future and every real year, bonds earn a year’s interest as well as their price change. Added on 9 October, after the first scoring; the claim held before and after.

Limits

What it does not show

  • That Prometheus is exactly right

    Passing means no big mistake showed up in twenty years. It does not prove Prometheus exactly right: a smaller mistake can hide in so few years.

  • That Prometheus beats a standard model

    Two standard bell curves passed the same test: one fitted on every year since 1970, and the crisis test’s own. A test both pass cannot rank them, so this shows Prometheus sees losses about as often as they happen, not that it sees them better.

  • Years it had seen before

    Prometheus was designed in 2026, and changes to it were checked on 19 of these 20 years, 2008 among them, with other portfolios. Fixing this test’s rules first cannot undo that hindsight: it is a bit like a student who has seen most of the exam questions beforehand.

  • Every country and every era

    One country, the United States, over 20 years, with only two losses of 10% or more.

  • 2010

    December 2009 could not be rebuilt, so 2010 was not tested.

  • Crashes that span two years

    Each year runs January to December, so a crash that spans a year end is split across two: 2008’s worst twelve months ran into 2009 and lost more than calendar 2008 did. The crisis test scores each crash on its own twelve months instead.

The maths

Every formula behind the pictures

Landing spot, and why it should be even
landing spot u =(futures at or below the real year)/ (all futures)If the real year R is drawn from themodel’s own distribution F, thenu = F(R), and for any a from 0 to 1:P(u ≤ a) = P(F(R) ≤ a)= P(R ≤ F⁻¹(a))= F(F⁻¹(a)) = awhich is the definition of an even(uniform) spread.

This is the probability integral transform. It needs nothing about the shape of F: it works for fat tails, skews and crashes alike, which is why it is the standard check of a probability forecast.

The evenness check (Kolmogorov–Smirnov)
sort the n landing spots:u(1) ≤ u(2) ≤ … ≤ u(n)staircase:F̂(x) = (number of u ≤ x) / ndiagonal:xD = the biggest |F̂(x) − x|over every x= max over i ofmax(i/n − u(i), u(i) − (i−1)/n)p = the chance that n dots spreadevenly (uniformly) give a gapof D or morepass if p ≥ 5%;for n = 20 that means D below 0.294

The p-value is computed exactly for twenty dots (the Kolmogorov distribution). It is not the chance that Prometheus is right; it is how surprising this much unevenness would be if it were.

Counting dots in part of the line
K = how many dots land in a partof the line covering a share qK ~ Binomial(n, q)P(K = k) = C(n, k) · q^k· (1 − q)^(n − k)“if right” range: fromthe first k with P(K ≤ k) ≥ 5%to the first k with P(K ≤ k) ≥ 95%
Losses expected and the “if right” range
p(t) = the share of the model’sfutures for year t thatlose 10% or moreexpected = p(1) + p(2) + … + p(n)(linearity of expectation)K = how many years bring such aloss, each year with chance p(t),independently of the othersits chances, multiplied out:start with [1]; for each year,multiply by ((1 − p(t)) + p(t)·x);P(K = k) is the coefficient of x^k(the Poisson-binomial)“if right” range: fromthe first k with P(K ≤ k) ≥ 5%to the first k with P(K ≤ k) ≥ 95%

The expected count needs no assumption at all. The range treats the years as independent, which is how the test was written before it ran.

Rebuilt using only the data available each December. Hypothetical, not investment performance.

Each “if right” range runs from the count at which a right model’s chances first add up to 5% to the count at which they reach 95%, so a right model’s count lands inside it at least nine times in ten.