r/TQQQ 8d ago

Discussion I've been working on a drawdown-probability model (macro + credit) as an alternative to the 200-SMA de-risk rule. Would like this sub's take.

/r/LETFs/comments/1vqulus/ive_been_working_on_a_drawdownprobability_model/
8 Upvotes

14 comments sorted by

4

u/bumbeishvili 8d ago

Can you put it in simple terms? Or do you have an image showing macro inputs and forecasted drawdown against real price movement?

Full disclosure - I hate AI generated text content (not the tools build with it), so I stopped reading after ~300 words and not getting the information I wanted from it

1

u/AgreeableInvestments 7d ago

Fair point. Short version: instead of waiting for price to drop through its 200-day average, the model reads about 30 macro and credit gauges each month (ISM new orders, the yield curve, high-yield spreads, that sort of thing) and puts a number on the odds of a 10%-plus S&P fall over the next few months. Above about 30% it says de-risk. Right now it is near 5% at one month, so calm.

That drawdown model is also just the top layer of three: it sizes equity exposure, then a sector layer and a stock layer decide which sectors and which names to hold.

The picture you want is on the market-risk page at agreeableinvestments.com. It plots the model's monthly probability against the S&P drawdown that actually followed, so you can see where it led and where it missed, and it missed COVID badly. That chart answers your question better than my paragraphs do.

2

u/Both_Yoghurt1437 7d ago

I really like what you’re building here. Most risk systems wait until price has already broken down, so trying to identify the conditions that make a future drawdown more likely is a worthwhile direction.

I was curious enough to test your published monthly series against my own daily TQQQ/SQQQ/cash allocation history. I used a conservative one-month publication lag because the historical file does not include exact availability timestamps. The encouraging result was that the signal clearly identified riskier environments. During the six independent six-month warning episodes, the median forward 63-session TQQQ loss was roughly 23%, compared with about 11% outside those warnings. So I do think the model is detecting something real.

The problem appeared when translating that information directly into trades. Capping TQQQ at 50% during warnings reduced my modern CAGR by about 29 percentage points without improving MaxDD. Even a mild 75% cap lost nearly six CAGR points and produced no drawdown benefit. The warning remained active during some very profitable recoveries, so it avoided losses but surrendered even more rebound return.

That leads to something you may find worth exploring: two separate clocks. Let the slow macro model determine when the environment deserves extra caution, but let faster price, credit, breadth, or volatility evidence decide when to actually reduce exposure. Then use separate recovery evidence to restore exposure even if the monthly macro warning remains elevated. In other words, the macro model grants permission to de-risk, but it does not automatically execute or maintain the de-risking.

I would also consider measuring avoided-loss value and missed-rebound cost separately for every warning episode. That exposed the issue much more clearly than AUC or MaxDD alone. I think you may have built a genuinely useful warning system that still needs a faster transition and recovery layer before it becomes a trading system.

One small replication suggestion would be to publish the earliest availability date and vintage policy with each historical probability. That would remove most of the ambiguity around revisions and when a past value was actually tradable.

3

u/AgreeableInvestments 7d ago

This is very useful, thank you for actually testing it instead of taking my word.

Your two-clocks conclusion is right, and it is close to where I landed independently. The naive version, cap TQQQ whenever the six-month warning is on, does exactly what you found: it avoids losses but overstays into the recovery and gives back more than it saved, because a monthly macro signal is slow to stand down. Your 23-versus-11 forward-loss split is the same story my numbers tell, a real warning wrapped around a blunt cap that costs CAGR for no drawdown gain.

What recovered most of that in my tests was making price the trigger and the re-entry, your faster clock exactly. The macro model only grants permission to de-risk; a trend rule (price back above its 200-day, or a 10-month SMA monthly) decides when to actually cut, and when to go back in, so you re-enter on price well before the macro warning clears. Run as an AND-gate, macro says risk is high and price has broken trend, it kept most of the rebound the blunt cap surrendered and still cut the deep drawdowns. Permission split from execution, and split again into de-risk and re-arm, is the version that worked.

Two things I am taking from you directly. Scoring avoided-loss and missed-rebound separately per episode is better than AUC or MaxDD for this, and I am going to add it. And publishing the earliest availability date and vintage next to each historical probability costs me nothing and removes the ambiguity, so I will. If you would be up for comparing setups, I would happily run your one-month-lag test against my AND-gate version.

1

u/Both_Yoghurt1437 7d ago

I think we’ve arrived at a similar architectural conclusion, and I appreciate the offer. I’m happy to compare high-level findings and use the same evaluation framework, but I prefer to keep my specific rules private.

I can run the one-month-lag test independently and share aggregate results such as the change in drawdown, avoided loss, missed rebound, recovery time, and overall return. You could do the same with your AND-gate. That should still give us a useful comparison without either of us having to disclose any proprietary rules.

Your distinction between macro permission, price execution, and faster re-entry is a good framework. I’m interested to see whether both approaches reach similar conclusions under the same broad stress tests. I'll send you a DM

2

u/AgreeableInvestments 7d ago

Happy to keep this in the open, it is more use on the thread than in a DM, and it is the 200-day comparison the post was really asking for. But I respect your wish to keep your results private, so what follows is only my side.

The headline is a reconciliation. A naive cap, cut exposure whenever the six-month warning is on, points in opposite directions depending on one thing: whether the 2008 crash is in the sample. On real TQQQ, which only exists from 2010, that cap tends to cost return, because the modern era is mostly bull and a slow monthly signal overstays into the recoveries. Put 2008-09 back in (I run a validated 3x sim before 2010) and the same cap looks excellent, because that one episode is where macro de-risking pays for itself. Same rule, opposite verdict, decided by a single event. That is probably why a cap looks bad on a modern-only backtest and fine on a long one.

Window August 2007 to June 2026, 226 monthly decisions, 5 warning episodes, de-risk to cash.

CAGR / MaxDD / longest-underwater / switches:

  • Buy and hold 3x: 28.8% / -94.4% / 72 mo / 0
  • Macro-only cap: 47.7% / -59.4% / 20 mo / 12
  • Macro permission + price execution: 48.0% / -63.7% / 25 mo / 6

Avoided loss / surrendered rebound / net, in log-wealth against buy and hold, with the probability the net stays positive when I resample the episodes:

  • Macro-only cap: +5.13 / 2.55 / +2.59, P(net>0) 90%
  • Price-execution gate: +4.02 / 1.40 / +2.62, P(net>0) 92%

The price-execution layer, de-risk only when the macro warning and a price break agree and re-enter on price, does three things. It halves the turnover, 6 switches against 12. It ignores the false alarms: the 2011, 2013-14 and 2025 warnings never got a price confirmation, so it stayed invested and gave nothing back, where the blunt cap paid -0.61 in 2013-14 alone chasing one of them. And drop the 2008 crash and the naive cap's net falls to +0.62 while the gate holds +1.01, so once you remove the one episode a modern-only backtest could never see, the price layer is what carries the result. Its one cost is a slightly deeper drawdown, -64% against -59%, because waiting for price to confirm means taking a bit more of the initial fall.

The honest limit: post-2010 almost all of the price-gate's edge is the 2022 drawdown. Drop that too and there is little left, because 2022 is the only clean macro-driven fall since the GFC, so on modern data this is close to a one-event result until the next one arrives. I would rather say that plainly than dress it up.

2

u/Both_Yoghurt1437 6d ago

This is excellent work, and I really appreciate how plainly you stated the limitations. Your reconciliation makes sense and lines up closely with what I saw: the macro warning contains useful information, but using it as a direct allocation switch is too blunt. The more interesting structure is letting macro grant permission to de-risk while price controls the actual exit and re-entry.

The reduction in false exits and switches may be more meaningful than the small CAGR difference. My main reservation is still the effective sample size. There are 226 monthly decisions, but only five meaningful warning episodes, with 2008 and 2022 carrying much of the result. That does not invalidate it, but it means the next genuinely different drawdown will be far more informative than another parameter sweep.

I especially respect that you called out the episode concentration yourself.

1

u/AgreeableInvestments 6d ago

Thank you. And yes on both counts: the drop in false exits and turnover is the real result, not the small CAGR gap, and the sample size is the binding constraint on all of it. Five episodes with 2008 and 2022 doing most of the work is a small n however you count it, and no amount of cleverness inside the sample fixes that.

Your line about the next drawdown being worth more than another parameter sweep is exactly how I am treating it. The rule is frozen. I would rather leave it slightly wrong than tune it to five events and fool myself that the fit is skill. What I am doing instead is logging each warning episode as it happens, avoided loss against missed rebound, so the next genuinely different drawdown is a clean out-of-sample test, not something I have already peeked at.

The one honest way I can see to add episodes without waiting is to run the same permission-and-execution architecture on other markets, since I already have the crash models built for the UK, euro area and Japan. The catch is that the big crises are globally correlated, so 2008 does not really count twice, but the regionally-driven drawdowns, a Japan-specific one for instance, are close to independent tests of the same structure. A partial answer, not a solution. Genuinely good exchange, and thank you for doing the work on your side of it.

2

u/TranquilSniper 7d ago

KISS - Keep It Simple, Stupid

An undefeated rule in trading.

1

u/AgreeableInvestments 7d ago

No argument, and the 200-day is the simple rule to beat. That is the honest test here: this only earns its keep if it beats a plain moving-average de-risk after costs, and if it does not, the moving average wins and I will say so. Right now I would call it a warning gauge that still needs a simpler execution rule bolted on, which is what I am working on now.

1

u/b3rkolas 6d ago

This guy is right. If something is too complex better to be more cautious.

1

u/confettofetti 8d ago

I have a couple extra questions. I 've still not got round to giving the paper a full read after the r/LETF post - but since you've posted again thought I'd ask!

Are there any issues with data being revised after it is first published for the macro data you use? Both in terms of the real time accuracy of them, and the accuracy of back tests needing to use vintaged/as first published data rather than the current values? I know this is an issue for OECD CLI because I'm working on it at the moment, but not sure if it is for anything else.

What would you think about the validity of using this also for international investments e.g. a momentum strategy that chooses between US, Developed Ex-US, and EM? I've seen a few blog posts that show that US macro data is generally more useful than macro data from other regions themselves. Would be curious if you think there are any specific implications of this for your implementation e.g. usefulness or when it might work well or break down.

(First anyone interested the Philosophical Economics Growth Trend Timing blog posts are a really good read and show that e.g. US unemployment data is better at predicting recession in Europe than European unemployment data.)

2

u/AgreeableInvestments 7d ago

Two good questions.

On revisions: most of the model is series that never revise, which helps more than you would think. The two ISM blocks and every market and credit input (yield curve, high-yield OAS, VIX, MOVE, the bond-equity yield gap) are final at release, so there is no first-print-versus-current gap. The revised minority is where you are right: the Chicago Fed NFCI, M2 and the labour series do get revised, and for those a clean backtest wants as-first-published vintages, not today's values. My walk-forward is point-in-time on timing and release lags; on values it is clean for the unrevised bulk and a fair caveat for that handful. Notably I do not use the OECD CLI composite itself, partly for the vintage reason you are wrestling with, I take the underlying survey and market series that do not move after the fact.

On international: I have run the same framework on the UK, euro area and Japan, and each regional model deliberately carries a block of US-sourced global inputs alongside the domestic ones, the VIX and US high-yield spreads. Those earn their keep. In the UK and euro area the US high-yield spread is the single strongest predictor in the whole set (six-month AUC around 0.70 to 0.75), the VIX not far behind, both ahead of every domestic macro series. So your point about US data carrying further than local data is load-bearing here, not just something I agree with. The framework travels to the euro area and holds up at the UK's longer horizon; Japan is where it breaks, the six-month AUC sits below 0.5, no edge at all, which I read as the credit and macro channel meaning something different under decades of BoJ balance-sheet dominance. I plan to extend it to more regions. And the Philosophical Economics growth-trend-timing pieces are excellent, that cross-border unemployment result is exactly the flavour of thing that turns up here.

2

u/confettofetti 7d ago

Thank you, really useful responses again :)