Two rules had already worked on their own: sit out the northern summer and hold shares only from November to April, and separately, hold less whenever markets have recently been turbulent. This combines them. The result was a portfolio identical to the first rule alone — the testing framework keeps total exposure fixed, so it cannot see the second rule at all.
Withdrawn. The combination is mathematically identical to one of its two ingredients — maximum weight difference 1.1e-16, return correlation 1.000 — because the framework holds total exposure fixed. The idea is untested rather than disproved.
Two rules that each passed on their own — holding shares only from November to April, and reducing exposure when markets have recently been turbulent — should work better combined than either does alone.
The harness normalises gross exposure to 1.0 by construction, so scaling an entire row of weights by a scalar is mathematically erased. Measured against W3a_HALLOWEEN_MULTI the maximum absolute weight difference is 1.1e-16, zero rows differ, and the net return streams correlate at 1.000. The hypothesis is not refuted — it is not expressible in this harness.
Two strategies had cleared the bar independently. The first is the old market saying 'sell in May and go away': hold a basket of stock markets from November through April, then sit in cash for the northern summer. The second ignores the calendar entirely and instead holds less whenever the market has recently been turbulent, on the grounds that turbulence persists while returns do not.
They read completely different information — one a date, the other recent volatility — so combining them is the obvious next question rather than a new idea, and it was pre-registered as a single additional trial.
It was initially scored as VALIDATED, beating its own base by 0.679 to 0.674. That margin is implausibly small for a change that should either help or hurt materially, which is what prompted the check.
Measured against the Halloween basket under the canonical harness, the maximum absolute weight difference is 1.1e-16, no row differs at all, and the two net return streams correlate at 1.000.
The cause is structural. The harness normalises every row to a gross exposure of 1.0 so that leverage is explicit and comparable across strategies. Multiplying an entire row of weights by a scalar — which is exactly what volatility targeting does — is removed by that normalisation. The earlier 0.679 came from an evaluation path that did not apply it.
The hypothesis is not refuted. Volatility targeting may well improve the seasonal basket; this harness simply cannot express the difference, because it deliberately holds gross exposure constant. Testing it properly requires an evaluator that permits varying leverage and charges for it.
A composition that cannot change gross exposure cannot express volatility targeting at all.
| Measure | Value |
|---|---|
| Annualised return | 5.4% |
| Annualised volatility | 9.3% |
| Return for risk taken | 0.58 |
| 95% range | 0.32 to 0.85 |
| t-statistic | 4.32 |
| Maximum drawdown | 36.1% |
| Hit rate | 63.6% |
| Annual turnover | 2.07× |
| Return skew | (0.09) |
| Return kurtosis | 10.62 |
| Diagnostic | Value | Reads as | Bad | Okay | Yay | Hmm |
|---|---|---|---|---|---|---|
| Design half (pre-2015) | 0.64 | return for risk | ≤ 0 | 0 – 0.40 | ≥ 0.40 | |
| Holdout half (2015+) | 0.36 | return for risk | ≤ 0 | 0 – 0.40 | ≥ 0.40 | |
| Top-5-month share of profit | 18.7% | how much rode on a few months — lower is better | ≥ 40% | 25 – 40% | < 25% | |
| Top-2-year share of profit | 17.7% | concentration | — | — | — | |
| Distinct trading episodes | 55 | how many independent runs this really is | < 50 | 50 – 100 | ≥ 100 | |
| Chance the edge is real | 100.0% | non-normality corrected | — | — | — | |
| Chance it beats the whole search | 25.6% | chance of beating the whole search by luck | — | — | — |
| Year | Net return |
|---|---|
| 1971 | 7.9% |
| 1972 | 10.7% |
| 1973 | (20.6%) |
| 1974 | (15.5%) |
| 1975 | 25.1% |
| 1976 | 16.0% |
| 1977 | (6.1%) |
| 1978 | 4.6% |
| 1979 | 11.2% |
| 1980 | 4.5% |
| 1981 | (2.0%) |
| 1982 | (0.4%) |
| 1983 | 16.2% |
| 1984 | (2.6%) |
| 1985 | 17.7% |
| 1986 | 9.8% |
| 1987 | 15.3% |
| 1988 | 4.9% |
| 1989 | 14.3% |
| 1990 | 1.4% |
| 1991 | 18.6% |
| 1992 | 3.2% |
| 1993 | 0.4% |
| 1994 | (8.6%) |
| 1995 | 13.1% |
| 1996 | 12.8% |
| 1997 | 7.3% |
| 1998 | 20.1% |
| 1999 | 27.1% |
| 2000 | (4.9%) |
| 2001 | 6.8% |
| 2002 | (1.1%) |
| 2003 | 5.5% |
| 2004 | 7.9% |
| 2005 | 0.5% |
| 2006 | 15.2% |
| 2007 | 4.4% |
| 2008 | (12.4%) |
| 2009 | 10.1% |
| 2010 | 3.4% |
| 2011 | (1.1%) |
| 2012 | 10.5% |
| 2013 | 5.4% |
| 2014 | 2.2% |
| 2015 | 4.1% |
| 2016 | 2.7% |
| 2017 | 10.0% |
| 2018 | 0.0% |
| 2019 | 11.9% |
| 2020 | (3.8%) |
| 2021 | 2.6% |
| 2022 | (3.5%) |
| 2023 | 10.8% |
| 2024 | 2.7% |
| 2025 | (0.7%) |
| 2026 | 0.5% |
| Measure | Value |
|---|---|
| Best years | 1999:+27%, 1975:+25%, 1998:+20% |
| Worst year | 1973:-21% |
| Rev | Stage | Status | Change |
|---|---|---|---|
| 01 | Draft | — | Logged as a single additional trial, control W3a |
| 02 | Tested | Validated | Reported 0.679 against the base strategy's 0.674 |
| 03 | Reviewed | Voided | Identical to the strategy it was built on — the framework holds total exposure fixed, so the change is invisible |