Showing posts with label clustering. Show all posts
Showing posts with label clustering. Show all posts

Thursday, 18 June 2026

To Cluster Or Not To Cluster That is the Question...

This is the sixth (!) post in a series I'm writing on portfolio optimisation. A quick reminder of the story so far:

  • In the first post I showed that if you are optimising across forecasts from different trading rules and instruments, that the rules within an instrument cluster naturally together, suggesting you should first fit within; and then across, instruments. Luckily, this is what I've always done.
  • In my second post I ran some experiments with optimising with random data. The results showed a supreme indifference between joint winners: monte carlo and bootstrapping, and a shrinkage methodology with a tiny bit of SR shrinkage. Using a more conservative viewpoint didn't affect the results much.
  • ... before moving into the world of real data for post three where I showed that the predictability of sampling distribution of parameter estimates was much worse with real data than with random data.
  • Then in post, number four, I reran my experiments this time with real data. Unsurprisingly I found that the previous winners did badly. Instead a middle ground of shrinkage was the winner across most time periods; unless the in sample period was too short.
  • Post number five attempted a spin on #4 by averaging out the weights produced by different shrinkage levels. It was an abject failure. We shall never speak of it again.
Importantly, all these experiments were done with a relatively small number of assets: nine. Bear in mind that when I fit forecast weights within instruments I have 40 assets; and when fitting instrument weights I have over 200.

Now you will remember that my previous favourite fitting method involves hierarchical clustering. I gave this the grand name of handcrafting, and in it's simplest form it doesn't use SR at all and just clusters by correlations or by some user assigned labels (eg asset class). 

There is plenty of empirical and theoretical evidence elsewhere as to why this makes sense for larger portfolios, but as I have the code and framework to do so, I thought why not battle test for myself; as part of this (long) exercise in checking that all my portfolio optimisation assumptions are correct. The question we want to answer then is what is the optimal number of clusters as a proportion of the total portfolio size? And if it's 1, then we don't cluster.
 
As before I can also vary the amount of available in sample data, and data used for out of sample evaluation. And this should be a joint test with the amount of shrinkage required. With clusters, we might expect less shrinkage to be required, since that is building in some robustness already. However in the interests of time, I won't be using Monte Carlo or Bootstrapping (doing that for each cluster in turn would be .... very..... slow..... indeed).

As before I will be doing this exercise using forecast weights (for which I generally have 40 assets). I can do some random selection, alternating between 20, 30 and 40 assets; randomly choosing instruments to do this with. 

Listing all the permutations:
  • Select 1,5 or 10 years of in sample data
  • Select 1 or 5 years of out of sample data
  • Pick a random instrument, ensuring there is enough history available (between 2 and 15 years). We will only choose from instruments with sufficient history for the time required.
  • Randomly pick N=20, 30 or 40 assets from those available (if 40, just use them all)
Then for a given dataset of in and out of sample:
  • Select correlation (from this list of options: 0, 0.7, 0.75, 0.8) and SR shrinkage (from this list: 0.0, 0.25, 0.5, 0.75, 1.0)
  • Select optimal number of clusters from 1 (no clustering), 2,3,4,6,8,10. 
  • Run in sample optimisation and out of sample optimisation on all the options above.
  • Repeat a few hundred times (it's quite slow!).
As I have done before I will also be checking speed - obviously adding clusters will increase optimisation time; firstly because the clustering and aggregation process itself adds time, but mainly because we will end up doing more optimisations. For example, for N=20 doing ten clusters would require 11 optimisations, one for each cluster, and one across clusters. Whilst the individual optimisations might be slightly faster than for smaller portfolios due to faster convergence, this is still probably going to be longer than a single optimisation. 

The speed of the optimisations varied between 28 seconds up to 294 seconds. Longer in sample time means slightly slower estimation time, more assets means slower convergence, less shrinkage also means slower convergence and as discussed just now smaller clusters also slows things down.

Since we're trying to establish here how much clustering to do, the only thing to note is that clustering imposes a speed penalty.

TLDR: This is a very long post due to the exhaustive number of combinations. There are numerous pretty plots to look at. But if you get bored and skip to the end to see the results, I won't judge you. Honest.

Results - forecast weights

20 assets - one year in sample, one year out of sample

In the previous couple of posts I showed the results as a matrix with different levels of shrinkage, for different time periods. Here it's a bit trickier, since we're considering the performance on 3 axis: two shrinkage, plus cluster size (where 1 is no clustering). But my screen only has two axis. Hence - in the words of many peoples relationship status on facebook circa 2010 - it's complicated.




This is a heatmap and the right hand rectangle is the key. The left hand is the data. The x-axis is the amount of correlation shrinkage. Apologies for the bunched up numbers: the values are 0,.7,.75,.8. Zero is there for fun; the other values are approximately. On the y-axis are the SR shrinkage and the number of clusters. So 1.0,8 is full shrinkage and 8 clusters. The colours are the median SR for each out of sample optimisation except that I've done what I did before in post four. I found the optimum value (zero correlation shrinkage, 0.75 SR shrinkage, cluster size of 5). I then did a t-test to compare that optimum SR value against all the other values. Where that test failed at a critical value of 0.1; in other words when the relevant value isn't significantly different from the optimum; I coloured in the square with the same colour SR value as the optimum.

If that explanation doesn't make sense you should probably reread post#4

The interpretation here is that with a few exceptions pretty much any value is fine as almost everything is one colour.

30 assets - one year in sample, one year out of sample

This is similar to 20 assets - nothing significant.

40 assets - one year in sample, one year out of sample

Here anything works, except:
  • Too much SR shrinkage (the coloured area at the bottom of the plot)
  • Not clustering

20 assets - one year in sample, five years out of sample

Again it's easier to say what doesn't work:
  • No correlation shrinkage
  • Too much SR shrinkage
The amount of clustering doesn't really influence the results.

30 assets - one year in sample, five years out of sample


Similar; perhaps a hint that smaller clusters underperform but just that.

40 assets - one year in sample, five years out of sample


Whilst there is more significance here, and a clear dislike for zero or full SR shrinkage, plus zero correlation shrinkage; again it does look like applying any degree of clustering is equally valid.

20 assets - five years in sample, one year out of sample

At this stage some people will be regretting their decision to print out my blogpost. In colour. For their sakes, I will be only reporting results where there is significance.

And for this combo, nothing is really significantly bad. 

30 assets - five years in sample, one year out of sample

Nothing is really significantly bad. 

40 assets - five years in sample, one year out of sample


 Certainly a preference here for lower amounts of SR shrinkage; with weaker evidence that middling amounts of clustering work well.

20 assets - five years in sample, five years out of sample

Focusing purely on clusters; it looks like 2 clusters is the one to avoid here.

30 assets - five years in sample, five years out of sample

Nothing is really significantly bad. 

40 assets - five years in sample, five years out of sample



20 assets - ten years in sample, one year out of sample

The whole plot is the same colour. Literally nothing to see here.

30 assets - ten years in sample, one year out of sample




40 assets - ten years in sample, one year out of sample

Again it looks like a modest amount of clustering; with N between 4 and 6 might be best.

20 assets - ten years in sample, five years out of sample

Looks like a case for larger cluster size...I think?

30 assets - ten years in sample, five years out of sample

Shrink the correlation and do some kind of clustering and you'll be fine mate.

40 assets - ten years in sample, five years out of sample

The final plot - and the one with the most significance. Shrinkage of about 0.75 on both and big clusters is the way to go here. More generally again it looks like middling cluster sizes are about right.

Summary

I think it is fair to say there aren't many definitive conclusions one can draw from that... experience. I would say however that there is some weak evidence that some level of clustering is better than none at all. And I would say there is even weaker evidence that you don't want your clusters to be too small, i.e. have a larger number of clusters. 

As in the previous post I could test the effect of combining portfolio weights derived with different cluster sizes. But given that we have struggled to find much statistical significance here it seems unlikely we'd get much satisifaction.

In the face of choosing a parameter in the face of no evidence there are two things I like - heuristic rules and powers of 2. Therefore let's say you should use 6 clusters when you're doing your thing. That's roughly equivalent to the number of distinct trading rules I have, and the number of asset classes when optimising for instrument weights. That isn't a power of 2, but 4 clusters seems a bit on the low size, and 8 a little high. In the face of choosing a parameter in the face of no evidence there are three things I like - heuristic rules, powers of 2, and taking an average of potential values.



Friday, 5 June 2026

The crossword puzzle of fitting - why across and then down?

This will be the first in a series of posts about portfolio optimisation. Main reason being I'm planning to write a book about backtesting, and that will include a big chunk of material on optimisation. Yes, I know, my latest book isn't out yet (it's out in December - in time for Christmas). But this backtesting book is going to be quite deep (and probably long!) so I need to start researching now if it's going to be written any time soon. Today's post is not that deep, and is quite short. It has literally been written whilst waiting for the rather extensive testing of the second post to finish. Anyone, let's begin. 

One of the issues when fitting is to decide how to structure and order the process. In very abstract terms, a component of a trading strategy will consist of a forecast to predict the price of an instrument. A forecast might be something like momentum16,64 - that's the exponentially weighted moving average crossover with spans of 16 and 64 days to you my good sir or madam. An instrument is something like the US 10 year bond future. We can represent all these options in a grid like so:

             momentum16,64                momentum4,16                 carry10

US10             X                            X                           X

SP500            X                            X                           X       

US5              X                            X                           X      


.... where each 'X' is a place on the grid. If those were white squares, and if you can imagine that if there were some forecasts missing from certain instruments which were black squares, then we'd have a crossword grid. Yes that's all I've got. Quite a weak link. Apologies.You think it's easy coming up with catchy blog titles?

And note that this is a tiny subset of the full grid. Instead of these 9 possibilities my full trading system currently has 10,373 options. That's 40 trading rules across 260 different instruments. Some of those instruments are duplicated (eg SP500 mini and micro), some are no longer traded; but that still leaves 204 instruments and over 8,100 options.

Anyone it should be obvious that in doing our fitting we have a few options:

  • A joint fit where we fit everything in one go. 8,100 options. In one go. Let that sink in.
  • A natural clustering where we cluster together things that are correlated.
  • A down and then across structured clustering where we first fit within rules - so for example working out what the best blend of US10, SP500 and US5 is within the carry10 rule - and then across rules - so estimating the best blend of carry10, momentum16,64 and momentum4,16.
  • An across and then down  structured clustering where we first fit within instruments - so for example working out what the best blend of carry10, momentum16,64 and momentum4,16 is within US10 - and then across instruments - so estimating the best blend of US10, SP500 and US5.
Now I've generally done the final one of these four options: across and down. And it's quite a natural way of doing things; because we seperate out the idea of predicting the price of a given instrument, and then put together a portfolio of trading substrategies one per instrument. But I've never actually tested the assumption that this is the right thing to do. In particular, if we were to do a natural clustering, would it come out quite like across and down; or would it be something weird? Or would, for example, all the momentum rules be more correlated with other irrespective of instrument, in which case down and across would make more sense?

Programming note: I've done this type of clustering before; here for underlying instrument returns, and here for forecasts

Incidentally, I've got three different approaches in the post; as I iteratively updated trying different things each time. TLDR: one of these approaches was inconclusive, the other two were firmly in favour of sticking to 'across then down'.


FIRST ATTEMPT: ONE GIANT CLUSTER

(This was the original blogpost)

Before I begin, I did decide to limit the analysis to the last 10 years. It's kind of slow just calculating a correlation matrix from 52 years of data and many instruments dont't have data except for the last 10 years anyway. I also resampled the returns to a weekly frequency. This is what I do when optimising anyway. Unless your returns are very quick, this won't affect correlations much. That leaves us with 520 weeks (I know that isn't exactly 10 years. Let me check to see if I give a toss. No, I don't) or rows in a dataframe, with over 8,100 columns remember. That's about the limit as to what my laptop can calculate a correlation for; and it's quite a painful process to cluster these bad boys as well.

Anyway let's begin. I've come up with quite a fun way to visualise these clusters which you can see here for the first two cluster plot:

OK as you can see each cluster has a subplot. Each splot has two stacked bars. The lefthand side shows the composition by asset class. The righthand side shows the composition by trading rule. You can just about make out from the legend what the various colours mean. Note these aren't portfolio weights, and just reflect the number of instrument/rule combos in each category. 

If your eyes are very good you will see the tops of the left hand bars don't quite reach 1. This is because I've removed anything with less than 2% of the total from the plots for clarity. Mainly this is a few odd instruments in a few very small asset classes (like volatility). It doesn't affect trading rules, since even the mrinasset rule has 1/40 = 2.5% of the total count.

The important thing here is we aren't yet seeing any evidence of clustering eithier by instrument or by trading rule. If we were, there would be a preponderance of colours on one side or the other. For example, if correlations were higher amoungst trading rules even from different instruments, then one cluster would have a lot of blue and purple in the right hand bar of one cluster (colours I assign to more divergent rules like momentum), and yellow and orange in the other right hand bar (colours that are reserved for convergent rules).

At this stage then there is no evidence that rows or columns makes more sense. 
Here is the three cluster plot. All three clusters look very similar for both instruments and rules, again suggesting there isn't much going on here yet on eithier axis.

I'm going to skip through the next few plots, since none of them show anything interesting. Let's check in at N=16:

Still not much going on here! We aren't seeing clumps of colours developing for eithier bar.

N=36:

49:

Never have so many coloured pixels mean wasted for so little result!

Let's close out with N=64, a nice power of 2 to finish on, as we're reaching the point where the plots are getting so small they're impossible to see:


There you have it folks. No clear evidence for eithier across-> down, or for down-> across. 

Now, there are a few alternative conclusions we could draw here. One is that there is some weird deep correlation pattern that the simple analysis by asset class and trading rule doesn't pick up. I don't buy that. If for example we look at the final cluster from N=64, it looks like this:

[assettrend4 forecasting AEX, breakout80 forecasting AEX, momentum32 forecasting ALUMINIUM, breakout160 forecasting AUD_micro, assettrend16 forecasting BITCOIN, normmom64 forecasting BITCOIN, accel64 forecasting BONO, carry125 forecasting BOVESPA, momentum64 forecasting BOVESPA, breakout80 forecasting BRE, normmom4 forecasting BRENT_W, accel32 forecasting BTP, momentum8 forecasting BTP3, accel16 forecasting BUTTER, assettrend16 forecasting BUTTER, carry125 forecasting CAD, assettrend32 forecasting CAD10, relcarry forecasting CAD2, breakout20 forecasting COAL, breakout40 forecasting COPPER_LME, relmomentum40 forecasting COPPER_LME, relmomentum20 forecasting COTTON, breakout320 forecasting DOW, assettrend8 forecasting EU-AUTO, relmomentum40 forecasting EU-CHEM, assettrend2 forecasting EU-CHEM, normmom2 forecasting EU-CHEM, relmomentum80 forecasting EU-DIV30, relmomentum10 forecasting EU-DJ-TELECOM, relmomentum20 forecasting EU-HOUSE, relmomentum20 forecasting EU-INSURE, carry10 forecasting EU-MEDIA, accel32 forecasting EU-OIL, assettrend2 forecasting EU-TECH, carry10 forecasting EU-TECH, breakout20 forecasting EURAUD, momentum8 forecasting EURCAD, carry30 forecasting EURCHF, carry60 forecasting EURCHF, normmom8 forecasting EURIBOR-ICE, mrinasset1000 forecasting EUROSTX, momentum4 forecasting EUROSTX-SMALL, normmom16 forecasting EUR_micro, assettrend32 forecasting FANG, mrinasset1000 forecasting FED, assettrend4 forecasting FED, normmom32 forecasting FEEDCOW, carry60 forecasting FEEDCOW, accel16 forecasting FTSECHINAH, breakout20 forecasting FTSEINDO, relmomentum20 forecasting FTSEINDO, breakout10 forecasting FTSETAIWAN, normmom8 forecasting FTSEVIET, breakout80 forecasting FTSEVIET, breakout40 forecasting GASOIL, relmomentum10 forecasting GAS_US_mini, momentum8 forecasting GAS_US_mini, mrinasset1000 forecasting GICS, accel32 forecasting HANGENT_mini, skewabs365 forecasting HANGENT_mini, carry30 forecasting HOUSE-US, momentum8 forecasting HOUSE-US, breakout320 forecasting HOUSE-US, assettrend2 forecasting IBEX_mini, relmomentum20 forecasting INR, normmom2 forecasting INR, normmom64 forecasting IRON, accel64 forecasting JGB, carry10 forecasting JGB-SGX-mini, normmom16 forecasting JGB-SGX-mini, assettrend4 forecasting KOSPI_mini, mrinasset1000 forecasting KR10, normmom32 forecasting KR3, momentum4 forecasting KR3, relmomentum40 forecasting KRWUSD_mini, normmom64 forecasting LEAD_LME, carry125 forecasting LEAD_LME, breakout10 forecasting LEANHOG, relmomentum40 forecasting LIVECOW, relcarry forecasting LIVECOW, breakout160 forecasting MILKWET, assettrend2 forecasting MSCIEAFA, skewabs180 forecasting MSCIEMASIA, relmomentum20 forecasting MSCIEMASIA, skewrv365 forecasting MSCIEMASIA, breakout160 forecasting MUMMY, normmom16 forecasting OAT, breakout80 forecasting OJ, skewrv180 forecasting PALLAD, normmom2 forecasting PALLAD, normmom64 forecasting PLAT, normmom16 forecasting RUSSELL, normmom2 forecasting SEK, skewrv365 forecasting SHATZ, normmom64 forecasting SHATZ, accel16 forecasting SHATZ, carry30 forecasting SHATZ, breakout10 forecasting SILVER, assettrend8 forecasting SMI, breakout40 forecasting SMI-MID, momentum16 forecasting SMI-MID, momentum64 forecasting SOFR, carry30 forecasting SOYBEAN_mini, normmom2 forecasting SOYBEAN_mini, skewabs365 forecasting SOYBEAN_mini, skewrv365 forecasting SP500_micro, skewabs365 forecasting STEEL, skewrv365 forecasting SUGAR16, relmomentum80 forecasting TIN_LME, carry10 forecasting TIN_LME, relmomentum40 forecasting TWD, normmom8 forecasting US-ENERGY, carry30 forecasting US-MATERIAL, mrinasset1000 forecasting US10U, carry10 forecasting US2, momentum32 forecasting US30, breakout40 forecasting US30, relcarry forecasting US30, skewrv180 forecasting US5, momentum64 forecasting V2X, carry125 forecasting VIX_mini, breakout10 forecasting VIX_mini, skewabs180 forecasting VNKI, normmom4 forecasting WHEAT, relmomentum10 forecasting YENEUR, assettrend4 forecasting ZAR, accel64 forecasting ZAR]

I look forward to anyone who can give me a coherent story as to why those things are lumped together. 

Personally I'm taking the absence of any contradictory evidence as evidence that I should continue to do what I've done before: fit across and down. Doing some kind of all group clustering or all in one fit, or doing down and then across; none of these seem to offer clear advantages. So why not stick with a simple thing that works?

Note - there is no reason why in theory you could not do an 'all in one' fit or weird clustering, and still use the procedure of generating a combined forecast for an instrument, then a subsystem position for an instrument, and then forming a portfolio. You just take the weights for each rule/instrument pairing from your all in one or weird cluster fit, and then take them across each instrument for forecast weights, and then add up the weights for a given instrument to find instrument weights.

Arguably this has been a waste of time, but the good news is I can recycle this code to visualise forecast weights across a strategy so that's something....


SECOND ATTEMPT: USE AVERAGE CORRELATIONS

As I was posting this up I thought of a much simpler test: if I measure the average correlation of forecasts within instruments (rule/instrument pairings for the same instrument) then it is 0.29. However the average correlation of forecasts within rules (eg rule/instrument pairings with the same rule) is much lower: 0.05. That reinforces that across and down is indeed the way to go.


THIRD ATTEMPT: TRY WITH SMALLER UNIVERSE

Arguably we might not see the pattern we required until we had hundreds of clusters; equal to the number of instruments in the universe. 

So I decided to do a 'small sample' approach. If I were to pick say 5 arbitrary instruments, and 5 arbitrary trading rules, what is the likelihood that these 25 components would cluster into instrument groups rather than rule groups; or nothing at all?

To make this a proper test I need to repeat the random choice of instruments and trading rules many times over. That means I can't just eyeball the charts each time, looking for a preponderance of colours on one side or the other. This is both timeconsuming and subjective.

Let's come up with a systematic rule. Seems on brand for this blog, doesn' it? Given 5 clusters, we consider them to have formed into instrument groups if  50% or more of the weight in a cluster is for a single instrument, and this has happened in 3 or more of the clusters. Alternatively, we consider them to be rule groups if 50% or more of the weight in a  cluster is for a single rule, and this has happened in 3 or more of the groups. Otherwise we consider them to be ungrouped (and that would be the case for all the charts above). Also any cluster with only one or two components is excluded from the calculation.

I do want to show you an example, since I don't want to waste my nice pictures. Here is what might be a particularly extreme example since it includes only US bonds (that are somewhat correlated): US2, US5, US10, US20, US30; and a smattering of trading rules (less likely to be correlated):

Cluster#1: [breakout40 forecasting US2, breakout40 forecasting US20, breakout40 forecasting US30, breakout40 forecasting US10, breakout40 forecasting US5] 

100% in single rule: meets rule grouping criteria

Cluster#2: [momentum64 forecasting US30, momentum64 forecasting US5, momentum64 forecasting US20, momentum64 forecasting US10, momentum64 forecasting US2], 

100% in single rule: meets rule grouping criteria

Cluster#3: [skewrv365 forecasting US2, skewrv365 forecasting US5, skewrv365 forecasting US20, skewrv365 forecasting US10, skewrv365 forecasting US30]

100% in single rule: meets rule grouping criteria

Cluster#4: [carry10 forecasting US10, carry10 forecasting US20, relcarry forecasting US20]

66.6% in single rule: meets rule grouping criteria; 66.6% in single instrument: meets instrument grouping criteria

Cluster#5: [carry10 forecasting US2, relcarry forecasting US2, relcarry forecasting US5, relcarry forecasting US30, relcarry forecasting US10, carry10 forecasting US30, carry10 forecasting US5]

57% in single rule: meets rule grouping criteria

Since 5/5 meet the rule group criteria (more than half), and only 1/5 meets the instrument group criteria, this is a rule group clustering.

Here is a more random selection:

['HEATOIL', 'SMI-MID', 'RUSSELL', 'BOBL', 'EUA']
['skewrv365', 'momentum64', 'carry10', 'relcarry', 'breakout40']

Which clusters as follows:

Cluster#1 [carry10 forecasting RUSSELL, relcarry forecasting RUSSELL, carry10 forecasting SMI-MID, breakout40 forecasting SMI-MID, relcarry forecasting SMI-MID, momentum64 forecasting SMI-MID, breakout40 forecasting RUSSELL, momentum64 forecasting RUSSELL] - this is actually two instrument clusters with exactly 50% in each

Cluster#2 [skewrv365 forecasting SMI-MID, carry10 forecasting EUA, relcarry forecasting EUA, skewrv365 forecasting EUA] - meets both criteria

Cluster#3 [carry10 forecasting HEATOIL, breakout40 forecasting HEATOIL, relcarry forecasting HEATOIL, momentum64 forecasting HEATOIL] - Heating oil cluster

Cluster#4 [breakout40 forecasting EUA, momentum64 forecasting EUA] - ignored, only 2 components.

Cluster#5 [carry10 forecasting BOBL, breakout40 forecasting BOBL, momentum64 forecasting BOBL, skewrv365 forecasting BOBL, relcarry forecasting BOBL, skewrv365 forecasting HEATOIL, skewrv365 forecasting RUSSELL] mostly BOBL


Since 3/4 valid clusters meet the 50% instrument threshold, and only one meets the 50% rule threshold, this would be a case where we would cluster by instruments most logically.

Anyway I repeated this exercise a few thousand times, and here are the results as a proportion of the total:

Meets neithier criteria: 3.8%
Meets both criteria: 0.81%
Meets rule grouping criteria: 0.11%
Meets instrument grouping criteria: 95.3%

That seems pretty conclusive