Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Friday, 3 July 2026

Jumping back in the pool(ing): testing pooling by asset class and portfolio weight distance

This is post #10 in my 2026 series on portfolio optimisation. Time for a quick recap. I'm not going to revisit every post but instead summarise what I now think one should be doing when optimising forecast weights before costs (I haven't yet incorporated costs, nor thought about instrument weights).

(I also confirmed in my very first post it was better to estimate forecast then instrument weights, rather than doing them jointly).

That doesn't seem like much value for the thousands of words I've written, and it's also not a million miles from what I would have down without all this research. A few things haven't worked out: random based methods (bayesian and monte carlo) which don't account for the reduced predictability of real returns compared to synthetic data; formal structural breaks on estimates; grouping and pooling instruments according to forecast SR; shorter EWM windows for SR estimates; and shrinking weights rather than inputs. 

I should probably move on now to looking at costs, and instrument returns, but like a dog with a particularly tasty bone or a cat pulling on an especially interesting piece of string; I can't quite let go of the idea that we should be able to improve on pooling everything.


Prior art

Let's run through the options we potentially have for pooling:

  1. We could cluster things together that have similar characteristics, such as by asset class. 
  2. We could do it on an estimate by estimate basis. We could compare the distribution of returns for say carry10 on US10 year bonds, and on US2 year bonds; and say "Well these distributions aren't significantly different. Let's pool the returns together". 
  3. Or we could look at the estimates of SR across different rules. You could have a vector of the SR for carry10, carry20 and so on. And you'd look at that vector of estimates, and calculate <some measure of> distance between them, and if the distance is low enough, then you'd pool the returns for all the rules for those two instruments.  I covered this in my previous post in this series. 
  4. Or we could do it on portfolio weights. We could for example fit the weights for 2 year bonds, and for 10 year bonds, and then see if they were significantly different. We could then pool the returns if they were not that different. I have also looked at this before.
  5. We don't pool at all, and fit each instrument individually. That sounds terrifying, but remember we're shrinking with our fitting.
  6. We pool everything. So far that seems to the best option, and the one I've used in the past. 

Note that we also have the option of:

  • A pooling the returns before estimating the statistics and then the weights
  • B not pooling the returns, and then pooling the weights
I'm not keen on B because it produces 'over robustness' when combined with a shrinkage methodology. Basically we throw away too much information and end up too close to equal weights.

So returning to the numbered list:

  1. By asset classes - is untried, although it resembles what we used to at AHL when we were organised into asset class teams, each of which fitted their own strategies.
  2. Grouping per estimate: I have objections to in terms of computational time and statistical unpleastness, discussed in the previous post.
  3. Grouping per vector of estimates: I tried this in the previous post. It wasn't effective, and also produced weird undesirable groups.
  4. Grouping by weights: I have tried before in a limited test with some success.

So that leaves us with 1 and 4 as candidates, along with the standard options of #6 full pooling and #5 no pooling at all - fitting each instrument's forecast weights purely on it's own data:

  • Unpooled
  • All instruments pooled
  • Asset class pooled
  • Grouping by portfolio weights


By asset classes - method

This is pretty trivial; the eight asset classes in my system are:

  • Stock indices (58 instruments in my dataset including duplicates and expired instruments like Eurodollar to avoid survivorship bias)
  • Sector stocks (eg 'EU oil companies') 36 instruments
  • Vol 4
  • FX 43
  • Bonds and STIR 39
  • Energies 20
  • Agricultural 39
  • Metals 21 (includes two crypto futures)

So for a given portfolio we fit those instruments in the same asset class together.


Grouping by weights

Well this is easy, as I already did this here and under the heading "get instrument groupings" it tells us we can use k-means clustering and there is even some code there for me to copy and paste. An important difference between this and the grouping by SR vector is that correlations will also be taken into account, at least in an implicit way.

One open question that remains is whether the grouping is done on portfolio weights that have been derived using a shrinkage method, or on weights that haven't (just using naive mean variance). I felt it was better to use 'purer' weights which hadn't used shrinkage so we don't end up discarding useful differences.

As I did in my prior post in this series, let's run the grouping exercise on my entire portfolio. Partly for laughs, and partly to see if the grouping makes sense. How many groups/clusters should we use? Well there are 7 substantive asset classes, excluding vol:

Cluster 0, length 16

BTP3, CANOLA, EU-FOOD, FED, GBPCHF, GOLD_micro, HIGHYIELD, LIVECOW, OMX, R1000, SILVER, SP500_micro, US-INDUSTRY, WHEAT, YENEUR, ZAR

Ags: 3, Metals: 2, Equity: 3, FX: 3, Sector: 2, Bond: 3


Cluster 1, length 29

BRENT_W, CNH, COAL-GEORDIE, COCOA, COFFEE, COTTON, ETHER-micro, EU-AUTO, EU-DJ-TELECOM, EU-DJ-UTIL, EU-MEDIA, EU-REALESTATE, EU-TECH, FTSETAIWAN, GBP, HEATOIL, KOSPI_mini, MILK, MSCIEAFA, MSCISING, OATIES, OJ, RUBBER, SEK, SMI, SONIA3, US-DISCRETE, US2, VNKI

OilGas: 3, Ags: 7, Metals: 1, Equity: 5, FX: 3, Sector: 7, Vol: 1, Bond: 2


Cluster 2, length 20

CAD10, CH10, CHF, CHFJPY, COPPER-micro, CZK, FTSECHINAA, FTSEINDO, IRON, JGB, JGB-SGX-mini, JP-REALESTATE, MUMMY, NIKKEI, SGX, SOYBEAN_mini, SOYOIL, TOPIX, US-ENERGY, US-HEALTH

Ags: 2, Metals: 2, Equity: 6, FX: 3, Sector: 3, Bond: 4


Cluster 3, length 12

AUD_micro, EU-INSURE, FTSE100, FTSECHINAH, GASOIL, HANG_mini, HOUSE-US, NASDAQ_micro, NOK, NZD, US-STAPLES, US-TECH

Sector: 4, Equity: 4, FX: 3, OilGas: 1


Cluster 4, length 16

ALUMINIUM, AUDJPY, BITCOIN, BOBL, BONO, BUND, BUXL, CORN, DOW, GBPJPY, MILKDRY, OAT, RICE, ROBUSTA, SHATZ, STEEL

Ags: 4, Metals: 3, Equity: 1, FX: 2, Bond: 6


Cluster 5, length 79

BB3M, BBCOMM, BRE, BTP, BUTTER, CAD, CHEESE, CHINAA-CON, COAL, COCOA_LDN, COPPER_LME, COTTON2, CRUDE_ICE, CRUDE_W_micro, DJSTX-SMALL, DX, EU-BANKS, EU-CHEM, EU-CONSTRUCTION, EU-MID, EU-OIL, EU-TRAVEL, EURCAD, EURCHF, EURIBOR-ICE, EUROSTX, EUROSTX-SMALL, EUR_micro, FANG, FEEDCOW, FTSE250, GAS-PEN, GASOILINE, GAS_US_mini, GICS, GILT, HANGENT_mini, IBEX_mini, IG, INR, IRS, JPY, KOSDAQ, KR10, KR3, LEAD_LME, LEANHOG, LUMBER-new, MIB, MILKWET, MILLWHEAT, MSCIEMASIA, MSCITAIWAN, MSCIWORLD, MXP, NICKEL_LME, PALLAD, PLAT, REDWHEAT, SARONA, SGD, SMI-MID, SOFR, SOYMEAL, SP400, SPI200, SUGAR11, SUGAR16, SUGAR_WHITE, TIN_LME, TWD, US10, US20, US5, V2X, VIX_mini, WHEAT_ICE, WHEY, ZINC_LME

OilGas: 6, Ags: 18, Metals: 7, Equity: 16, FX: 12, Sector: 6, Vol: 2, Bond: 12


Cluster 6, length 32

AEX, BOVESPA, BRENT-LAST, CAC, CAD2, CAD5, CLP, DAX, EU-BASIC, EU-DIV30, EU-HEALTH, EU-HOUSE, EU-RETAIL, EUA, EURAUD, EURO600, FTSEVIET, GBPEUR, HANGTECH, KRWUSD_mini, MSCIASIA, PLN, RUSSELL, SWISSLEAD, US-FINANCE, US-MATERIAL, US-PROPERTY, US-REALESTATE, US-UTILS, US10U, US3, US30

OilGas: 2, Equity: 11, FX: 5, Sector: 9, Bond: 5

There doesn't seem much congruency with asset classes there. Is this "the data speaking to us", or are we just data mining with a very sharp spade? Let's find out.


Testing

In my older post, here, I did a rather simplistic 'one shot' test on a subset of my available instruments and forecast rules (albeit on a rolling out of sample basis). But I have a rather more exhaustive way of doing things I've been using in this series. 

I cycle through different lengths of in sample (5 years, 10 years, 20 years) and out of sample (1 year and 5 years) lengths of time. For shorter time periods that will allow me to subsample different historic periods. For speed and to get some alternative paths I'm not going to consider all the instruments. Instead I will randomly subsample 50 instruments randomly out of the 214 available. 

Then for a given set of returns I will eithier use fully pooled, asset class pooled, portfolio weight pooled, or unpooled returns. Then I will optimise for each instrument based on the relevant returns, using the shrinkage method with SR shrinkage of 0.5 and correlation of 0.75. Finally I will take the equally weighted across instruments portfolio SR for the 50 instruments, out of sample. 

In previous posts I've discussed a more honest way of backtesting, where we include the opposite of a given trading rule to avoid implicit fitting; and then only bring positive SR rules into the optimisation. All the results here will use that methodology exclusively.


5 years in sample, 1 year out of sample

You should hopefully recognise this format from before. Each row is a fitting option. The first column shows the median SR across the many, many runs of random resampling. The second column shows the t-test p-value from comparing the best option with the others. NaN means this is the best option. A low number in this column, say below 0.01 or 0.05, indicates that the best option is statistically significantly better than the other option.

                           SR  pvalue

unpooled                 0.046     0.0

all pooled               0.425     0.0

asset class pooled       0.557     NaN

weight distanced pooled  0.184     0.0

That is ... pleasing. The least robust method is worse. More robust methods do better. And we get a significant improvement from pooling within asset classes. OK the portfolio weight distancing isn't so good, but we haven't got huge amounts of data to form our portfolio weights with so maybe they are a little unstable.


5 years in sample, 5 years out of sample

                            SR  pvalue
unpooled                 0.418     0.0
all pooled               0.391     0.0
asset class pooled       0.731     NaN
weight distanced pooled  0.376     0.0

Unpooled does a little better here, but asset classes are still the way to go.

10 years in sample, 1 year out of sample

                            SR      pvalue
unpooled                 0.664     NaN
all pooled               0.417   0.000
asset class pooled       0.635   0.127
weight distanced pooled  0.415   0.000

OK interestingly unpooled is making a comeback, but it still isn't significantly better than asset class pooled.

10 years in sample, 5 years out of sample

                            SR  pvalue
unpooled                 0.855     0.0
all pooled               0.575     0.0
asset class pooled       1.028     NaN
weight distanced pooled  0.559     0.0

Asset class is again asserting it's dominance with unpooled a close second.

20 years in sample, 1 year out of sample

                            SR  pvalue
unpooled                 0.571     0.0
all pooled               0.469     0.0
asset class pooled       1.127     NaN
weight distanced pooled  0.410     0.0

OK this is getting a bit silly. I feel like the dad whose kid at sports day is winning everything, proud but also getting a little embarrassed. "Now come on jonny, let one of the other kids win the next one". 

Interestingly it does seem with more data that unpooled is the way to go for a second option.


20 years in sample, 5 years out of sample

                            SR  pvalue
unpooled                 0.251     0.0
all pooled               0.569     NaN
asset class pooled       0.551     0.0
weight distanced pooled  0.553     0.0

"Well done Jonny. Everyone knows you could have won it if you wanted but it's good to show good sportmanship"

So all pooled finally gets it's day in the sun albeit with a slim advantage over the other two pooled methods. Bear in mind only 65 instruments have sufficient history here; with only 18 having two distinct blocks of 25 years so there won't be much genuine variation if we choose 50. So this could be a fluke. Jonny will tell you that it is.


But Rob, What about averaging?

At this point, given the choice between the complexity of weight distancing, and the simplicity and efficiency of asset class pooling; I'm inclined to go with the latter. And it's what we were doing at AHL all those years ago (not because of empirical evidence but because it suited the organisational structure...).

However there is another option which I talked about in the original asset class pooling post, using a blend. Here we take an average of the portfolio weights selected with different methodologies. So that would be an average of:

  • Unpooled
  • All pooled
  • Asset class pooled
Blending weights in this way is a way to improve robustness. It's arguably the correct thing to do, since otherwise we'd be making an in sample choice of methodology - 'meta implicit fitting' if you will. One of my favourite research shops, Resolve asset management, are very keen on doing this. One potential downside is it might be producing 'over robustness' given we're using weights that have already had shrinkage. But let's find out.


5 years in sample, 1 year out of sample

Note these numbers won't be exactly the same as those above, since they're a different set of random experiments. They would eventually converge but it would take millions of runs.

And also, just for fun, I've added an extra column. I started off this series of posts talking about the importance of considering other points of the distribution but I've quietly dropped that and only been quoting the median. For balance then, I've added the 25% SR point as well as the median. The pvalue is as before.

                    SR median  SR 25%  pvalue
unpooled                0.025  -0.569     0.0
all pooled              0.475  -0.147     0.0
asset class pooled      0.620   0.045     0.0
average                 0.650  -0.067     NaN

Averaging is the winner - just - but asset class is better at the more conservative point.

5 years in sample, 5 years out of sample

                    SR median  SR 25%  pvalue
unpooled                0.428   0.164     0.0
all pooled              0.430   0.243     0.0
asset class pooled      0.771   0.537     NaN
average                 0.664   0.430     0.0

A clear win for asset class pooled here. Averaging suffers from it's association with the less performative unpooled / all pooled.

10 years in sample, 1 year out of sample

                    SR median  SR 25%  pvalue
unpooled                0.681   0.034     0.0
all pooled              0.452   0.067     0.0
asset class pooled      0.664   0.167     0.0
average                 0.935   0.256     NaN

This time averaging takes the win, helped by the good performance of unpooled.

10 years in sample, 5 years out of sample

                    SR median  SR 25%  pvalue
unpooled                0.887   0.641     0.0
all pooled              0.570   0.406     0.0
asset class pooled      1.066   0.833     NaN
average                 1.005   0.780     0.0

Asset class pooled is still the winner, but averaging gives a good job.

20 years in sample, 1 year out of sample

                    SR median  SR 25%  pvalue
unpooled                0.538  -0.209     0.0
all pooled              0.515   0.154     0.0
asset class pooled      1.167   0.539     NaN
average                 0.809   0.252     0.0

Asset class pooled by more of a margin now.

20 years in sample, 5 years out of sample

                   SR median  SR 25%  pvalue
unpooled                0.243   0.084     0.0
all pooled              0.531   0.414     NaN
asset class pooled      0.513   0.391     0.0
average                 0.461   0.355     0.0

As before 'all pooled' is the winner, whilst average is dragged down by the poor performance of unpooled. But as I said above, with these longer periods it's hard to know if it's just down to flukey instrument selection.

What to do...

There is enough evidence above to justify asset class pooling as the dominant choice. But equally, I don't think there is enough to discard averaging. And there is something so neat about averaging. We combine three quite disparate source of data together, so we're protected if one of them doesn't work out. It's robustness writ large! It can be justified without any in sample fitting - whereas one could argue that the selection of asset class pooling is an implicit in sample 'meta parameter' choice.

I think we're now (finally) ready to fit our forecast weights, and with costs. This is exciting for me, as whatever comes out I will be using as my new weights. This will be more of a 'literature review' since I've talked about optimising with costs in some detail and at some length before.

Monday, 29 June 2026

Rolling, rolling, rolling.... updating statistical estimates yes or no

 The mega blog post series on portfolio optimisation continues!

A couple of posts ago, here, I looked at using the idea of formal testing for structural breaks in parameter estimates. Important parameters like Sharpe Ratio (SR). Because stuff like this happens:


This is the pre-cost performance of the momentum4 rule on CORN. The formal test found a structural break in 1989.

It's fair to say the structural break stuff didn't work that well. But there may be a much easier way of dealing with the non stationarity of these estimates, and that's to use rolling estimates. For example, if you were to use a 10 year rolling estimate of SR then by the mid 1990s we would conclude that this was a money losing rule. We could also use a rolling estimate for correlation, though as these are stable enough over periods of five years or more this wouldn't affect things much.

Of course I wouldn't be so crude as to use a mere rolling window, instead I'd use an exponential window. As usual I'm going to specify this using the span parameter of the pandas ewm function. A 10 year span has a 3.5 year half life; i.e. the 10 year EWM is roughly equivalent (same halflife) to a 7 year simple moving average.


The test

Regular readers will know exactly what to expect here, but for those that aren't regular here is how I test this procedure to be as sure as possible there is no luck involved.

  • Select 10,20,30 or 40 years of in sample data (shorter periods won't make sense to apply an exponentially weighted [EW] estimation)
  • Select 1 or 5 years of out of sample data
  • Pick a random instrument, ensuring there is enough history available (between 11 and 45 years). We will only choose from instruments with sufficient history for the time required. 
  • Randomly pick N=9 forecasting rules from those available (the same as in previous posts)

Then for each of those sumsamples:

  • Cycle through using an exponentially weighted [EW] span of 5,10,20,30 years; and no span (use all available history). For shorter in sample periods the EW results using longer spans will be very similar to those without EW estimation.
  • Estimate SR using the EW span.
  • Estimate correlation using all the in sample data (we could use an EW span here, but correlations are sufficiently stable that it won't unduly affect results).
  • Use fixed shrinkage levels (estimated here): SR shrinkage 0.5, correlation 0.75 (since we'll always have at least five years of in sample data we don't need to worry about the higher levels of shrinkage required when we have insufficent data). The results won't be much different with any vaguely similar shrinkage; you could argue we'd need more shrinkage with shorter EW spans but I am not going to test this.
  • Run in sample optimisation and out of sample optimisation on all the options above

Finally once we have all our subsamples:

  • Get the median SR from the distribution of subsamples
  • Find the optimal EWM span with the higest median SR
  • Test to see if that optimum is significantly higher than the others

As I also did in my last post I'm going to see if the figures are different without any implicit fitting. To achieve this I include the opposite of a given trading rule as a candidate; and then when I come to do optimisation I pick the version that has a positive SR (there are no costs, so the SR will be identical with a negative sign).


10 years in sample, one year out of sample

We only have five optionts to consider so we can do this in a simple table.

         SR  pvalue
5 0.026 0.065
10 0.021 0.026
20 0.033 NaN
30 0.030 0.042
999999 0.017 0.131

Each row is a different EW span. '99999' means the entire in sample period was used. The next column is the out of sample Sharpe Ratio for each option. In the second column is the p-value for a test of the optimal option against the relevant option. NaN is the optimal option, and lower values (say below 0.05) mean the optimal option is significantly better than the alternatives. We can see that a 20 year span is the optimal, and it's a little better than the other alternatives but not significantly better than the entire in sample period.

Do the results differ when we don't preselect only the 'correct' rules?


SR pvalue
5 0.076 0.188
10 0.093 NaN
20 0.085 0.186
30 0.087 0.169
999999 0.083 0.236

Nothing is really significant there.


10 years in sample, five years out of sample


SR pvalue
5 0.201 0.009
10 0.215 NaN
20 0.214 0.078
30 0.211 0.277
999999 0.199 0.271

SR pvalue 5 0.093 0.000 10 0.112 NaN 20 0.104 0.339 30 0.104 0.466 999999 0.110 0.533

Here we do get better performance with anything more than 5 years.

20 years in sample, one year out of sample

          SR  pvalue
5 -0.154 0.892
10 -0.170 0.687
20 -0.163 0.610
30 -0.164 0.647
999999 -0.138 NaN


          SR  pvalue
5 -0.202 0.006
10 -0.135 0.011
20 -0.110 0.029
30 -0.097 0.059
999999 -0.055 NaN

Longer estimates are better.

20 years in sample, five years out of sample


SR pvalue
5 0.093 0.000
10 0.119 0.000
20 0.134 0.026
30 0.132 0.043
999999 0.149 NaN
          SR  pvalue
5 0.061 0.000
10 0.111 0.003
20 0.129 0.070
30 0.131 0.071
999999 0.140 NaN

Yes, longer estimates are better.


30 years in sample, one year out of sample


SR pvalue
5 -0.073 0.0
10 -0.078 0.0
20 -0.031 0.0
30 -0.005 0.0
999999 0.042 NaN

SR pvalue 5 -0.132 0.007 10 -0.152 0.000 20 -0.156 0.000 30 -0.104 0.000 999999 0.000 NaN


Same story, slightly different numbers.

30 years in sample, five years out of sample

          SR  pvalue
5 -0.038 0.0
10 -0.031 0.0
20 -0.017 0.0
30 -0.006 0.0
999999 0.008 NaN


SR pvalue
5 -0.050 0.001
10 -0.046 0.003
20 -0.038 0.003
30 -0.029 0.003
999999 -0.017 NaN


40 years in sample, one year out of sample


SR pvalue
5 -0.135 0.145
10 -0.113 0.023
20 -0.110 0.001
30 -0.090 NaN
999999 -0.150 0.947

Again, we basically want a very long estimate.
          SR  pvalue
5 -0.159 NaN
10 -0.224 0.044
20 -0.236 0.028
30 -0.238 0.072
999999 -0.254 0.061

That was a little unexpected.

40 years in sample, five years out of sample


SR pvalue
5 -0.049 0.0
10 -0.040 0.0
20 -0.029 0.0
30 -0.021 0.0

999999 -0.015     NaN    


SR pvalue
5 -0.097 0.000
10 -0.064 0.000
20 -0.016 0.006
30 -0.006 NaN
999999 -0.029 0.703


Conclusion

I'm a big believer in publishing (well blogging) research even if it doesn't result in a positive result. And certainly it looks like you don't really gain anything from using exponentially weighted estimates of Sharpe Ratios for optimisation, versus the simpler alternative of using all the data. Still there is that nagging feeling that we should at least have the option of dropping something that hasn't worked for a while which implies a very slow EWM. A 30 year EWM span has a 10.3 year halflife, the same as a 20 year or so SMA; whilst a 40 year EWM span is equivalent to a 28 year SMA. 



Tuesday, 23 June 2026

Breaking Badly: finding the structural breaks in parameter estimates

 Here's a nice picture from a lovely book written by a top bloke:

It shows the cumulative p&l from different speeds of momentum over time (for portfolios containing 102 instruments) over 50 years of data. Notice how the two fastest speeds (2&4) get worse in the second half of the sample. I've called the line #2 here the 'second most famous hockey stick graph in history'. It certainly looks like something changed in 1990. 

This is important. If we're optimising portfolios of such things we only want to consider data that is relevant, but we also want as much data as possible for statistical significance. Now if I were a simpleton I'd do this by looking at graphs like that and going 'aha i only need to use data after 1990'. As a simpleton I don't use capital letters. But I am a big fan of not doing in sample fitting, even of meta parameters like this; and I am an even bigger fan of doing things automatically which means not wading through thousands of graphs like that (since there are thousands of SR estimates in my forecast p&l space, plus a good chunk of correlations).

So we need an automatic way of identifying such breaks. Fortunately this is not a new problem as you will know if, like me, you did undergraduate econometrics. Finding structural breaks is an entire industry. We need two things: a test for how likely it is that a break has occured between two sub-samples A and B. And an algorithim for going through all the options of A and B

And in case you haven't realised this is the seventh post in my summer 2026 series on portfolio optimisation.


What parameters

The first question to think about is what parameters we're going to apply this process to. I do two kinds of optimisation:

  • Forecast weights
  • Instrument weights
And in both cases I have estimates of SR (one per asset) and correlations ([N^2-N]/2 for N assets). I haven't really looked at instrument weight optimisation yet in this series, and there are some wrinkles there so I'm going to park that for now. That just leaves the SR for a forecast (which remember is a pairing of a trading rule and an instrument), and the correlation of such forecasts within an instrument.

Now I am going to ignore correlations in this post. As I discussed in an earlier post, although correlations are relatively unstable in the short term, they are unlikely to have secular trends like SR. And it's quite easy to deal with this by just using a relatively long lookback to estimate them, probably with an ewma on the correlation estimate.


What test

This is quite an easy one, compared to the world of econometrics and linear regression where we have to do such nonsense as a Chow test. Given two sub-samples A and B, to find out if they have different sample means we just do an independent t-test: https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.ttest_ind.html

An important characteristic of such tests is they are only likely to be significant if there is sufficient data. If sample A or sample B is too small in size, it's unlikely we'll find a signficant difference in the means.
Typical values for critical values in t-tests are 1% and 5%; I will also check out 10% later.

Since all the p&l streams in a risk targeted trading system will have the same expected standard deviation in the long run, I can legitmately treat this as a test to see if the SR are different if I adjust the samples so they have an identical standard deviation. 

(If any ex-students of mine are reading this, you should remember this from week 3 of the course!)

What algo / procedure

OK let's make this concrete. Suppose we have 50 years of data as in the original graph above, and we're currently wondering if there were one or more structural breaks. We don't want to have less than five years data to do our estimation. Note these numbers are appropriate for my context, but not in a domain with faster trading, higher SR, and faster alpha decay. We proceed as follows:
  • We compare year 1 with years 2 to 50. Since year 1 is very small it's unlikely we'll find a break; but if we do then we do a split (see below).
  • If no split has occured, we compare years 1-2 with years 3-50. Again if we find a break, we split.
  • If no split has occured, we compare years 1-3 with years 4-50...
  • ...
  • If no split has occured, we compare years 1-45 with years 46-50. If we still don't find a break, then we use the entire period for estimation (since the period 47-50 will only have four years, we terminate here).
Now what if a split occurs at some point? Then we restart the process, but this time without the pre-split data. Suppose for example a split had occured in year 20 (which is 1992 in the original graph). Then:
  • We compare year 21 with years 22 to 50. If we find a break, then we split again.
  • If no split has occured, we compare years 21-22 with years 23-50. Again if we find a break, we split.
  • If no split has occured, we compare years 21-23 with years 24-50...
  • ...
  • If no split has occured, we compare years 21-45 with years 46-50. If we still don't find a break, then we use the entire period from years 21-50 for estimation.
You get the idea. This procedure is quite quick and easy to run; and notice we're identifying any break that exceeds a certain threshold rather than finding the most likely break as we would do with a QLR type test. 

The main downside of it is that it won't identify breaks that reverse. For example if the world is in regime A, then regime B, then regime A again then ideally we'd estimate our parameters using both regime A's. But the above test will eithier use only the final regime A, or the last two regimes, or possible all three regimes depending on whether there is a significant difference at the appropriate point(s). There are fancy things we could do to deal with this, but I feel these are corner cases and life is too short.


An example


Let's use a concrete example. This is the performance of the momentum4 rule on CORN. I've chosen it because we already know momentum4 has a structural break, and CORN has plenty of history. Just for fun, before reading on, see if you can identify where the algo finds the break here (there is exactly one break - this isn't a trick question).

Some code

This is probably the first post in this series where it's been practical to actually include the code, since there isn't much of it, nor are there are any dependencies apart from getting the returns:


from copy import copy
from typing import List, Callable

import numpy as np
import pandas as pd
from scipy.stats import ttest_ind

BUS_DAYS_IN_YEAR =
256
import matplotlib.pyplot as plt

MIN_NUMBER_OF_YEARS =
5

def identify_and_plot_breaks(all_returns: pd.Series, CV: float =0.01):
breaks_as_dict = identify_all_breaks(all_returns)
breaks_as_df = pd.DataFrame(breaks_as_dict)
breaks_as_df = breaks_as_df.bfill(axis=1)
breaks_as_df.cumsum().plot()
plt.show(
block=True)


def identify_all_breaks(all_returns: pd.Series, CV: float):
## returns a dict, turn into a dataframe and you can plot
returns_to_consider= copy(all_returns)
returns_to_consider = returns_to_consider.dropna()
broken_list = identify_all_breaks_recursively(returns_to_consider=returns_to_consider, list_of_returns_broken_off=[],
CV=CV)
broken_list.reverse()
broken_dict = dict([
(
idx, value) for idx, value in enumerate(broken_list)
])

return broken_dict

def identify_all_breaks_recursively(returns_to_consider: pd.Series, list_of_returns_broken_off: List, CV: float) -> List:
years_in_returns = how_many_years_approx(returns_to_consider)

for i in range(years_in_returns):
year_idx=i+1
first_sample, second_sample = split_sample_after_n_years(returns_to_consider, year_idx)
if len(second_sample)<(MIN_NUMBER_OF_YEARS*BUS_DAYS_IN_YEAR):
break
is_broken_here = test_a_break(first_sample, second_sample, CV=CV)
if is_broken_here:
list_of_returns_broken_off.append(first_sample)
return identify_all_breaks_recursively(
second_sample, list_of_returns_broken_off=list_of_returns_broken_off,
CV=CV
)
else:
continue

## No breaks identified or sample size too short
list_of_returns_broken_off.append(returns_to_consider)
return list_of_returns_broken_off

def how_many_years_approx(returns: pd.Series):
return int(np.floor(len(returns)/BUS_DAYS_IN_YEAR))

def split_sample_after_n_years(all_returns: pd.Series, n_years: int):
idx = n_years*BUS_DAYS_IN_YEAR
return all_returns[:idx], all_returns[idx:]

def test_a_break(first_sample: pd.Series, second_sample: pd.Series, CV: float):
## Normalise by standard deviation before considering means
norm_first_sample =first_sample/first_sample.std()
norm_second_sample=second_sample/second_sample.std()
return ttest_ind(norm_first_sample, norm_second_sample).pvalue<CV
On Corn this produces the following:

The break occurs on the 26th August 1982. So to calculate that particular SR estimate we'd only use data from that date onwards.

Is that what you would have guessed? Personally, I would probably have gone for a later break point if identifying it by eye. As humans we are drawn to the sharp upward move in 1989 and would probably have gone for a break just after that. It's possible a search for the likeliest break would have found that point, but remember we are looking for the first break that exceeds the threshold; and once that break happens no further breaks are identified.

A summary of results

Here is a summary of the results for each instrument/forecast pairing with the default 1% critical value:


You can see that breaks are quite rare with only 13% or so of instrument/rules having at least one break. This also suggests that the Sharpe Ratios for trading rule performance are actually quite stable over time; or at least stable enough that they won't fail any statistical tests at a 1% critical value. 

Multiple breaks are even rarer. Just 1.8% have two breaks; 0.4% or 39 instruments have three breaks, ten have four breaks and only two have five breaks. They are:

skewrv365 forecasting EURIBOR-ICE    (yes, there is still a EURIBOR future!)
normmom2 forecasting FTSE250         

Here is Euribor, relative value skew with a 365 day window:


Although five is pushing it, there are certainly three regimes there (pre 2000, 2000 - 2010, and 2010 onwards), and using post 2010 data seems to make some kind of sense.


And FTSE 250 momentum8 (this is pre-cost):


There are certainly at least two regimes there and I wouldn't argue with the automated decision to use only data after 2006 or so, when it looks like; in the words of Pulp in one of my favourite songs, "Something changed".

Of course we will get a different picture with a slacker test. Here is the picture with a 10% critical value:

Now just over half the pairings have at least one break in them.

The decision as to use 0% (equivalent to no breaks at all), 1%, 5% or 10% CV is one we will now address.

An optimisation test

Now the big question is does this actually improve performance? On a pure out of sample test? And is this changed much by using a different critical value?

I follow the same procedure roughly as in previous posts:

  • Select 10,20,30 or 40 years of in sample data (I need at least 10 years because with a minimum of five years required for estimation I certainly won't find any breaks, or I will risk finding a break and not having five years of data leftover)
  • Select 1 or 5 years of out of sample data
  • Pick a random instrument, ensuring there is enough history available (between 11 and 45 years). We will only choose from instruments with sufficient history for the time required.
  • Randomly pick N=9 forecasting rules from those available (the same number as in posts #2 and #3)
Then for each of those sumsamples:
  • Cycle through using no breaks (0% CV), 1% CV, 5% CV and 10% CV
  • Estimate SR on the insample data using eithier all the data (0% CV), or the data after the last break given some critical value. 
  • Estimate correlation using all the in sample data
  • Use fixed shrinkage levels (estimated here): SR shrinkage 0.5, correlation 0.75 (since we'll always have at least five years of in sample data we don't need to worry about the higher levels of shrinkage required when we have insufficent data). The results won't be much different with any vaguely similar shrinkage.
  • Run in sample optimisation and out of sample optimisation on all the options above
Finally once we have all our subsamples:
  • Get the median SR from the distribution of subsamples
  • Find the optimal CV with the highest SR
  • Test to see if that median is significantly higher than the others

10 years in sample, one year out of sample

We only have four options to consider so no need for the huge tables and fancy heatmaps of previous posts:
         SR  pvalue all  pvalue distinct
0.00 -0.021       0.247            0.247
0.01 -0.018       0.295            0.295
0.05 -0.019       0.204            0.204
0.10 -0.014         NaN              NaN
Each row is a different critical value used for breakpoint finding. Zero means the entire in sample period was used. The next column is the out of sample Sharpe Ratio for each option. In the second column is the p-value for a test of the optimal option against the relevant option. NaN is the optimal option, and lower values (say below 0.05) mean the optimal option is significantly better than the alternatives. In the final column I've rerun the staistical t-test but this time I have excluded instances where no breaks were found (so the test is only done comparing the out of sample SR when breaks were found, versus when they were not). This shouldn't affect the p-values, but it's nice to check it doesn't

You can see here that it it looks like a very loose breakpoint policy is the best, but it's not significantly better.

10 years in sample, five years out of sample

         SR  pvalue all  pvalue distinct
0.00 0.166 0.168 0.168
0.01 0.160 0.003 0.003
0.05 0.167 0.007 0.007
0.10 0.170 NaN NaN

Again the loosest breakpoint works best; but it's hardly logical since it's indistinguishable from no breakpoints at all.

20 years in sample, one year out of sample

         SR  pvalue all  pvalue distinct
0.00 -0.160 0.542 0.542
0.01 -0.160 0.323 0.323
0.05 -0.160 0.050 0.050
0.10 -0.155 NaN NaN
Loose is better.

20 years in sample, five years out of sample

         SR  pvalue all  pvalue distinct
0.00 0.123 NaN NaN
0.01 0.116 0.0 0.0
0.05 0.105 0.0 0.0
0.10 0.098 0.0 0.0
No breakpoints are the best. There isn't much consistency here.

30 years in sample, one year out of sample

         SR  pvalue all  pvalue distinct
0.00 0.092 NaN NaN
0.01 -0.019 0.0 0.0
0.05 -0.078 0.0 0.0
0.10 -0.066 0.0 0.0
Another massive win for not using breaks.

30 years in sample, five years out of sample

        SR  pvalue all  pvalue distinct
0.00 -0.005 0.001 0.001
0.01 -0.011 0.000 0.000
0.05 -0.008 0.014 0.014
0.10 -0.003 NaN NaN
Or perhaps we should go for the loosest break....

40 years in sample, one year out of sample

         SR  pvalue all  pvalue distinct
0.00 -0.142 NaN NaN
0.01 -0.249 0.000 0.000
0.05 -0.189 0.000 0.000
0.10 -0.178 0.002 0.002
or no breaks...

40 years in sample, five years out of sample

         SR  pvalue all  pvalue distinct
0.00 -0.060 0.177 0.177
0.01 -0.051 NaN NaN
0.05 -0.055 0.000 0.000
0.10 -0.055 0.000 0.000
A clear vote for strict breaks, but no breaks at all are also good.

Conclusion

Well that was as clear as a bowl full of mud that has been made even more unclear by painting the bowl a very dark colour and then adding some darker mud. Not very clear at all, in other words.

You think up a nice neat simple way of finding structural breaks, and then it doesn't actually work when used for optimisation. In seven out of the cases we examined not using breaks is eithier optimal, or statistically insignificant from the optimal. Only in one case was it inferior. The loosest possible break (CV 10%) was optimal in four cases. In most cases apart from 30 years/1 year the difference in performance was small between CV alternatives. This partly reflects the fact that breaks aren't that common, especially with a 1% CV.

Of course findings like this are very context dependent. I'm using rules that should probably work over long time periods; indeed they have been in sample selected for such a purpose. For the 30 year and 40 year periods there just aren't that many instruments with that much history; so even repeated random sampling is likely to turn up the same suspects repeatedly.

One issue here might be that we are considering the breaks at a forecast/instrument level. We might get different results if we pool estimates for trading rule performance across instruments. And indeed that is the subject of the next post. So I will return to this topic when I've looked at pooling.



Monday, 15 June 2026

FIFA* World Cup (*Fitting and Forecasting Actual data) Portfolio Optimisation competition with real returns

This is my fourth post in my summer 2026 mini series on portfolio optimisation. 

It will very much follow the format of (also with a sports alluding title) blog post number two, so it might be worth rereading that. A reminder if you can't be bothered, I used random data to compare some optimisation methods:

  •  monte carlo (random, parameteric)
  •  bootstrapping (random, non parametric)
  • double shrinkage (shrinking SR towards average SR, and correlations to zero). This encompasses some other methods including:
    • NMV naive mean variance (no shrinkage on anything)
    • EW equal weights (both full shrinkage)
    • MD maximum diversification (no shrinkage correlation, full shrinkage on SR)
    •  EPO (we just shrink the correlation matrix to some degree)

I found that MC/Bootstrap were the best, and didn't require any pesky estimation of the shrinkage meta-parameter. But they are SLOW. I worked out you'd need quite a few iterations to get the weights to converge, so each optimisation took quite a while. Should you wish to estimate that meta-parameter I found that for random data with a nice stable distribution that you didn't need much shrinkage. A little bit on the Sharpe Ratio was the most optimal; a little more wouldn't harm things much, but a lot was bad.  

However as we know from post three, real data is not as nice as random data, and is much harder to forecast. It has a habit of doing annoying things, like changing it's distribution when you're not looking. So we're expecting that we will need, for example, more shrinkage to reflect this.

The real data we will be using will many different runs, each consisting of 9 randomly selected trading rules, chosen for a single randomly chosed instrument. Because we know from post one that fitting within instruments is the way to go. Although I currently have 40 trading rules in my actual portofolio, I am sticking with nine now for speed and intuition. Plus the results shouldn't be too different with more components - that is something I will be looking at later in the series. I'm sampling with replacement so it's feasible - but very unlikely- I'll get the same instrument/rule set more than once.

As per my previous posts I'm also going to compare the results for different lengths of data. In the random data post I could generate as much data as I want; that's tricky here when the absolute longest history I have for any instrument is just over 50 years and many are much less than that. So I'm going to use in sample lengths of 1 year, 5 years and 10 years; and out of sample lengths of 1 year and 5 years. If an instrument doesn't have sufficient data for a given pairing I won't use it; eg for 10 years/5 years I would need 15 years which will be tricky for many instruemnts whilst for 1 year/1 year I would just need 2 years obviously. If it has more data than required, then on a given random run I'll randomly select the required 2 to 15 year long period.

First some speed statistics. We already know that shrinkage will be darn quick, but as I'm using different data lengths from the prior post it's probably worth repeating the stats for montecarlo and bootstrap:

              1 year in sample       5 years in sample      10 years in sample

BS          9.2                     20.6                       33.3

MC          5.1                      6.6                        8.0

Remember from the previous post that convergence is quicker with Monte Carlo than with Bootstrap, hence the substantially longer time taken to do BS which needs twice as many iterations; as well as the slight difference in implementation per iteration which explains the even worse performance of BS at longer iterations.

Results

One year in sample, One year out of sample

Let's begin with the median results. For the moment I'm going to present two data frames. The first is just Sharpe Ratios. Here is the one for an insample and out of sample period of just one year:

      0.00   0.20   0.40   0.60   0.70   0.75   0.80   0.90   1.00
0     0.056  0.057  0.039  0.054  0.047  0.049  0.046  0.055  0.032
0.25  0.061  0.045  0.054  0.063  0.057  0.049  0.044  0.042  0.037
0.5   0.059  0.057  0.048  0.047  0.044  0.046  0.041  0.046  0.044
0.75  0.049  0.041  0.062  0.058  0.041  0.026  0.029  0.047  0.033
0.8   0.030  0.041  0.061  0.054  0.041  0.025  0.026  0.026  0.032
0.85  0.016  0.038  0.050  0.035  0.030  0.029  0.023  0.025  0.036
0.9  -0.002  0.022  0.043  0.030  0.041  0.038  0.029  0.024  0.052
0.95  0.015  0.022  0.045  0.043  0.049  0.043  0.049  0.034  0.056
1.0   0.014  0.003  0.038  0.060  0.056  0.058  0.043  0.032  0.004
MC   -0.000 -0.000 -0.000 -0.000 -0.000 -0.000 -0.000 -0.000 -0.000
BS    0.018  0.018  0.018  0.018  0.018  0.018  0.018  0.018  0.018

This will look very familiar if you looked at the previous post on random data, but there are a couple of extra rows. From the top then each column shows a different degree of correlation shrinkage. On the left 0.0 is no shrinkage where we used the estimated data. 1.0 is full shrinkage, where all correlations are set to zero. Apart from the diagonals. Obviously. Each row then is a different degree of SR shrinkage, from the top row where we use no shrinkage, down to the row labelled 1.0 where we fully shrink all SR to the average SR across assets. 

The bottom two rows are the results for Monte Carlo and Bootstrapping. There is no shrinkage here, so for consistency I've just copied the single value for each across all columns. 

Some elements of interest in the main part of the table, the top left corner (0.0, 0.0) is naive mean variance with no shrinkage, the top right (0.0, 1.0) is full correlation shrinkage, the bottom left (1.0, 0.0) is full SR shrinkage, and the bottom right (1.0, 1.0) is full shrinkage on both which leads to equal weights. The EPO empirical optimal is (0, 0.75).

The optimum value here has some shrinkage: 0.25 on SR and 0.60 on correlations. 

Compare and contrast that with the results for random data. The optimal shrinkage was barely nothing: 0.25 SR, 0 correlations or thereabouts. It isn't surprising we need more shrinkage in general. Remember from the previous post in this series on random data:

Essentially random data sets a lower bound on robustness calibration. For example, suppose we determine that the correct shrinkage for the vector of expected SR on a Bayesian portfolio optimisation using random data is 0.1 (which means we average using 90% of the estimated SR, and 10% of the prior SR). Then it's likely the correct shrinkage on real data will be higher than 0.1.

However the amount of optimal SR versus correlation shrinkage might seem surprising. Quoting now from post three in this series, on forecasting statistical parameters with real data:

In simple terms, we are a little bit worse than forecasting Sharpe Ratios in real data one year ahead than we would be with random data, but a LOT worse with correlations. Partly this is because we are pretty terrible at forecasting SR one year ahead anyway even with a stable underlying distribution; we don't do much worse with real data. However it does seem that correlations are far more unstable in reality than in randomly generated data.... If we recall from the prior post that the optimal shrinkage is zero on correlations with random data; we can now see why with actual data we'd probably want to opt for some correlation shrinkage; purely because the sampling error is much larger in practice. That is the empirical finding of the EPO paper. It does feel a bit weird since up to now my gut feeling has been that we have to shrink means a lot because they are much harder to forecast and because they have an outsized effect on portfolio weights compared to differences in correlation. Whilst the latter is still true it seems the former is not.

There are two different effects here remember: predicability of each estimate compared to random data (where correlation is worse), and more about their outright predictability (where SR is worse), and the different effects each has on MV optimisation (small differences in SR affect the outcome more).

Another surprise might be the relatively poor performance of MC and BS. Remember that the only difference between them is the assumption of joint Gaussian returns in one case and not in the other.  In the random data round each method was the best performing. Both however are making an implicit assumption that there is a stable distribution (parameteric in one case, not in the other), and that any variance in outcome over the out of sample period will be the same as would be expected from the sampling distribution of each parameter. Which is exactly what happens with random data. But we know from post three that the parameter estimates we're making have a wider distribution with real data; and this is especially true for correlations. Hence, the MC/BS methods are too optimistic about predictability and their weights are suboptimal compared to those produced by high shrinkage optimisations.

Note: I have ideas to fix that, which may or may not in a subsequent blog post. Briefly they involve playing with the MC parameter inputs to reflect the higher RMSE of real versus random data.

Now let's run a paired t-test comparision of that optimum median value against all other values. Here are the p=values from doing those tests:


0.00 0.20 0.40 0.60 0.70 0.75 0.80 0.90 1.00 0 0.91 0.64 0.39 0.63 0.86 0.83 0.90 0.56 0.22 0.25 0.84 0.83 0.99 NaN 0.27 0.24 0.64 0.75 0.37 0.5 0.62 0.37 0.34 0.20 0.07 0.19 0.17 0.54 0.79 0.75 0.76 0.81 0.70 0.70 0.79 0.93 0.85 0.99 0.93 0.8 0.81 0.95 0.81 0.84 0.95 1.00 0.99 0.67 0.97 0.85 0.99 0.97 0.72 0.67 0.76 0.69 0.84 0.60 0.99 0.9 0.92 0.68 0.84 0.85 0.66 0.64 0.74 0.71 0.91 0.95 0.73 0.64 0.91 0.73 0.72 0.60 0.56 0.66 0.94 1.0 0.71 0.68 0.87 0.46 0.41 0.42 0.34 0.42 0.47 MC 0.05 0.05 0.05 0.05 0.05 0.05 0.05 0.05 0.05 BS 0.06 0.06 0.06 0.06 0.06 0.06 0.06 0.06 0.06

We can see that the optimum value itself is NaN since the p-value is undefined. We can also see that statistically, there isn't much difference in how shrinkage is used. MC and BS are definitely worse 

Now, how do things change if we are more pessimistic? As before I'm going to look at the 5% distributional point of outcomes from my multiple random results. If I do this, the optimal shrinkage is 0.8 on correlations, but a massive 1.0 on SR. At the 25% point it's 0.75 on SR but 1.0 on correlations. We want more shrinkage for sure!

Now let's think about a nice graphical way of showing these values. I'll start with a heatmap of the median SR:



Now I'm going to do something familar to the students on my course. I'm going to replace every value that is statistically insigificant from the optimal median with the optimal media value. Here I will use a 90% critical value:
Here the result looks like a really shit piece of modern art. Since almost all shrinkage values are not significantly different from the optimal; except weirdly correlation shrinkage 0.7 and SR shrinkage 0.5 (which is adjacent to the optimum); it's just a sea of blue. But we can see the MC/BS methods are inferior.


One year in sample, Five years out of sample


It isn't obvious but I used the same procedure for this plot which shows SR, with all values that can't be distinguished from the optimum in the same colour as that optimum. But every single value other than the optimum, which is full shrinkage or equal weights, is inferior to that optimum.


Five years in sample, One years out of sample

A very interesting picture here. There's clearly a shrinkage area that doesn't work. Note that the results overall are quite poor.

Five years in sample, Five years out of sample

A little clearer here. Modest shrinkage would work well, but then so would random data. Just don't shrink the SR too much.


Ten years in sample, One years out of sample

Importantly here the critical value is 80%, not 90%. With 90% the whole plot goes one colour. Pretty much any amount of shrinkage works. Again the SR results are very poor.

Ten years in sample, Five years out of sample



Summary of results

Well that was messy. I'd conclude that shrinkage of SR 0.5 and correlation 0.75 (the EPO value) is in the optimum region in almost all time periods. That's a reversal of what my original intuition suggested and I've used before, with more shrinkage on the SR. I've explained at length why my intuition was wrong. The random methods (MC/BS) are also inferior in many cases, as well as being slow.

The exception is one year / five years where you need full shrinkage (equal weights). Using 0.5/0.75 isn't so bad however. Although it's significantly worse, the actual loss in SR is small. Still it does seem logical to use more shrinkage with more data; and we can see from the one year/five year plot that we're better off shrinking SR more. So here is my heuristic rule of thumb:

Five or more years of data: SR shrinkage 0.5, correlation 0.75

Four to five years of data: SR shrinkage 0.6, correlation 0.75

Three to four years of data: SR shrinkage 0.7, correlation 0.80

Two to three years of data: SR shrinkage 0.8, correlation 0.85

One to two years of data: SR shrinkage 0.9, correlation 0.90

One or less than one year of data: SR shrinkage 1.0, correlation 1.0 (equal weights)

These results are very domain specific. In particular, I'm mostly dealing with holding periods in the weeks and months. A faster trading system would be able to compress the periods above. But the main lesson is that it's very hard to state categoricially what the exact amount of shrinkage should be. The surface is mostly too noisy. So don't sweat it. Use a vaguely okay value and you'll do vaguely ok.