Showing posts with label asset allocation. Show all posts
Showing posts with label asset allocation. Show all posts

Friday, 3 July 2026

Jumping back in the pool(ing): testing pooling by asset class and portfolio weight distance

This is post #10 in my 2026 series on portfolio optimisation. Time for a quick recap. I'm not going to revisit every post but instead summarise what I now think one should be doing when optimising forecast weights before costs (I haven't yet incorporated costs, nor thought about instrument weights).

(I also confirmed in my very first post it was better to estimate forecast then instrument weights, rather than doing them jointly).

That doesn't seem like much value for the thousands of words I've written, and it's also not a million miles from what I would have down without all this research. A few things haven't worked out: random based methods (bayesian and monte carlo) which don't account for the reduced predictability of real returns compared to synthetic data; formal structural breaks on estimates; grouping and pooling instruments according to forecast SR; shorter EWM windows for SR estimates; and shrinking weights rather than inputs. 

I should probably move on now to looking at costs, and instrument returns, but like a dog with a particularly tasty bone or a cat pulling on an especially interesting piece of string; I can't quite let go of the idea that we should be able to improve on pooling everything.


Prior art

Let's run through the options we potentially have for pooling:

  1. We could cluster things together that have similar characteristics, such as by asset class. 
  2. We could do it on an estimate by estimate basis. We could compare the distribution of returns for say carry10 on US10 year bonds, and on US2 year bonds; and say "Well these distributions aren't significantly different. Let's pool the returns together". 
  3. Or we could look at the estimates of SR across different rules. You could have a vector of the SR for carry10, carry20 and so on. And you'd look at that vector of estimates, and calculate <some measure of> distance between them, and if the distance is low enough, then you'd pool the returns for all the rules for those two instruments.  I covered this in my previous post in this series. 
  4. Or we could do it on portfolio weights. We could for example fit the weights for 2 year bonds, and for 10 year bonds, and then see if they were significantly different. We could then pool the returns if they were not that different. I have also looked at this before.
  5. We don't pool at all, and fit each instrument individually. That sounds terrifying, but remember we're shrinking with our fitting.
  6. We pool everything. So far that seems to the best option, and the one I've used in the past. 

Note that we also have the option of:

  • A pooling the returns before estimating the statistics and then the weights
  • B not pooling the returns, and then pooling the weights
I'm not keen on B because it produces 'over robustness' when combined with a shrinkage methodology. Basically we throw away too much information and end up too close to equal weights.

So returning to the numbered list:

  1. By asset classes - is untried, although it resembles what we used to at AHL when we were organised into asset class teams, each of which fitted their own strategies.
  2. Grouping per estimate: I have objections to in terms of computational time and statistical unpleastness, discussed in the previous post.
  3. Grouping per vector of estimates: I tried this in the previous post. It wasn't effective, and also produced weird undesirable groups.
  4. Grouping by weights: I have tried before in a limited test with some success.

So that leaves us with 1 and 4 as candidates, along with the standard options of #6 full pooling and #5 no pooling at all - fitting each instrument's forecast weights purely on it's own data:

  • Unpooled
  • All instruments pooled
  • Asset class pooled
  • Grouping by portfolio weights


By asset classes - method

This is pretty trivial; the eight asset classes in my system are:

  • Stock indices (58 instruments in my dataset including duplicates and expired instruments like Eurodollar to avoid survivorship bias)
  • Sector stocks (eg 'EU oil companies') 36 instruments
  • Vol 4
  • FX 43
  • Bonds and STIR 39
  • Energies 20
  • Agricultural 39
  • Metals 21 (includes two crypto futures)

So for a given portfolio we fit those instruments in the same asset class together.


Grouping by weights

Well this is easy, as I already did this here and under the heading "get instrument groupings" it tells us we can use k-means clustering and there is even some code there for me to copy and paste. An important difference between this and the grouping by SR vector is that correlations will also be taken into account, at least in an implicit way.

One open question that remains is whether the grouping is done on portfolio weights that have been derived using a shrinkage method, or on weights that haven't (just using naive mean variance). I felt it was better to use 'purer' weights which hadn't used shrinkage so we don't end up discarding useful differences.

As I did in my prior post in this series, let's run the grouping exercise on my entire portfolio. Partly for laughs, and partly to see if the grouping makes sense. How many groups/clusters should we use? Well there are 7 substantive asset classes, excluding vol:

Cluster 0, length 16

BTP3, CANOLA, EU-FOOD, FED, GBPCHF, GOLD_micro, HIGHYIELD, LIVECOW, OMX, R1000, SILVER, SP500_micro, US-INDUSTRY, WHEAT, YENEUR, ZAR

Ags: 3, Metals: 2, Equity: 3, FX: 3, Sector: 2, Bond: 3


Cluster 1, length 29

BRENT_W, CNH, COAL-GEORDIE, COCOA, COFFEE, COTTON, ETHER-micro, EU-AUTO, EU-DJ-TELECOM, EU-DJ-UTIL, EU-MEDIA, EU-REALESTATE, EU-TECH, FTSETAIWAN, GBP, HEATOIL, KOSPI_mini, MILK, MSCIEAFA, MSCISING, OATIES, OJ, RUBBER, SEK, SMI, SONIA3, US-DISCRETE, US2, VNKI

OilGas: 3, Ags: 7, Metals: 1, Equity: 5, FX: 3, Sector: 7, Vol: 1, Bond: 2


Cluster 2, length 20

CAD10, CH10, CHF, CHFJPY, COPPER-micro, CZK, FTSECHINAA, FTSEINDO, IRON, JGB, JGB-SGX-mini, JP-REALESTATE, MUMMY, NIKKEI, SGX, SOYBEAN_mini, SOYOIL, TOPIX, US-ENERGY, US-HEALTH

Ags: 2, Metals: 2, Equity: 6, FX: 3, Sector: 3, Bond: 4


Cluster 3, length 12

AUD_micro, EU-INSURE, FTSE100, FTSECHINAH, GASOIL, HANG_mini, HOUSE-US, NASDAQ_micro, NOK, NZD, US-STAPLES, US-TECH

Sector: 4, Equity: 4, FX: 3, OilGas: 1


Cluster 4, length 16

ALUMINIUM, AUDJPY, BITCOIN, BOBL, BONO, BUND, BUXL, CORN, DOW, GBPJPY, MILKDRY, OAT, RICE, ROBUSTA, SHATZ, STEEL

Ags: 4, Metals: 3, Equity: 1, FX: 2, Bond: 6


Cluster 5, length 79

BB3M, BBCOMM, BRE, BTP, BUTTER, CAD, CHEESE, CHINAA-CON, COAL, COCOA_LDN, COPPER_LME, COTTON2, CRUDE_ICE, CRUDE_W_micro, DJSTX-SMALL, DX, EU-BANKS, EU-CHEM, EU-CONSTRUCTION, EU-MID, EU-OIL, EU-TRAVEL, EURCAD, EURCHF, EURIBOR-ICE, EUROSTX, EUROSTX-SMALL, EUR_micro, FANG, FEEDCOW, FTSE250, GAS-PEN, GASOILINE, GAS_US_mini, GICS, GILT, HANGENT_mini, IBEX_mini, IG, INR, IRS, JPY, KOSDAQ, KR10, KR3, LEAD_LME, LEANHOG, LUMBER-new, MIB, MILKWET, MILLWHEAT, MSCIEMASIA, MSCITAIWAN, MSCIWORLD, MXP, NICKEL_LME, PALLAD, PLAT, REDWHEAT, SARONA, SGD, SMI-MID, SOFR, SOYMEAL, SP400, SPI200, SUGAR11, SUGAR16, SUGAR_WHITE, TIN_LME, TWD, US10, US20, US5, V2X, VIX_mini, WHEAT_ICE, WHEY, ZINC_LME

OilGas: 6, Ags: 18, Metals: 7, Equity: 16, FX: 12, Sector: 6, Vol: 2, Bond: 12


Cluster 6, length 32

AEX, BOVESPA, BRENT-LAST, CAC, CAD2, CAD5, CLP, DAX, EU-BASIC, EU-DIV30, EU-HEALTH, EU-HOUSE, EU-RETAIL, EUA, EURAUD, EURO600, FTSEVIET, GBPEUR, HANGTECH, KRWUSD_mini, MSCIASIA, PLN, RUSSELL, SWISSLEAD, US-FINANCE, US-MATERIAL, US-PROPERTY, US-REALESTATE, US-UTILS, US10U, US3, US30

OilGas: 2, Equity: 11, FX: 5, Sector: 9, Bond: 5

There doesn't seem much congruency with asset classes there. Is this "the data speaking to us", or are we just data mining with a very sharp spade? Let's find out.


Testing

In my older post, here, I did a rather simplistic 'one shot' test on a subset of my available instruments and forecast rules (albeit on a rolling out of sample basis). But I have a rather more exhaustive way of doing things I've been using in this series. 

I cycle through different lengths of in sample (5 years, 10 years, 20 years) and out of sample (1 year and 5 years) lengths of time. For shorter time periods that will allow me to subsample different historic periods. For speed and to get some alternative paths I'm not going to consider all the instruments. Instead I will randomly subsample 50 instruments randomly out of the 214 available. 

Then for a given set of returns I will eithier use fully pooled, asset class pooled, portfolio weight pooled, or unpooled returns. Then I will optimise for each instrument based on the relevant returns, using the shrinkage method with SR shrinkage of 0.5 and correlation of 0.75. Finally I will take the equally weighted across instruments portfolio SR for the 50 instruments, out of sample. 

In previous posts I've discussed a more honest way of backtesting, where we include the opposite of a given trading rule to avoid implicit fitting; and then only bring positive SR rules into the optimisation. All the results here will use that methodology exclusively.


5 years in sample, 1 year out of sample

You should hopefully recognise this format from before. Each row is a fitting option. The first column shows the median SR across the many, many runs of random resampling. The second column shows the t-test p-value from comparing the best option with the others. NaN means this is the best option. A low number in this column, say below 0.01 or 0.05, indicates that the best option is statistically significantly better than the other option.

                           SR  pvalue

unpooled                 0.046     0.0

all pooled               0.425     0.0

asset class pooled       0.557     NaN

weight distanced pooled  0.184     0.0

That is ... pleasing. The least robust method is worse. More robust methods do better. And we get a significant improvement from pooling within asset classes. OK the portfolio weight distancing isn't so good, but we haven't got huge amounts of data to form our portfolio weights with so maybe they are a little unstable.


5 years in sample, 5 years out of sample

                            SR  pvalue
unpooled                 0.418     0.0
all pooled               0.391     0.0
asset class pooled       0.731     NaN
weight distanced pooled  0.376     0.0

Unpooled does a little better here, but asset classes are still the way to go.

10 years in sample, 1 year out of sample

                            SR      pvalue
unpooled                 0.664     NaN
all pooled               0.417   0.000
asset class pooled       0.635   0.127
weight distanced pooled  0.415   0.000

OK interestingly unpooled is making a comeback, but it still isn't significantly better than asset class pooled.

10 years in sample, 5 years out of sample

                            SR  pvalue
unpooled                 0.855     0.0
all pooled               0.575     0.0
asset class pooled       1.028     NaN
weight distanced pooled  0.559     0.0

Asset class is again asserting it's dominance with unpooled a close second.

20 years in sample, 1 year out of sample

                            SR  pvalue
unpooled                 0.571     0.0
all pooled               0.469     0.0
asset class pooled       1.127     NaN
weight distanced pooled  0.410     0.0

OK this is getting a bit silly. I feel like the dad whose kid at sports day is winning everything, proud but also getting a little embarrassed. "Now come on jonny, let one of the other kids win the next one". 

Interestingly it does seem with more data that unpooled is the way to go for a second option.


20 years in sample, 5 years out of sample

                            SR  pvalue
unpooled                 0.251     0.0
all pooled               0.569     NaN
asset class pooled       0.551     0.0
weight distanced pooled  0.553     0.0

"Well done Jonny. Everyone knows you could have won it if you wanted but it's good to show good sportmanship"

So all pooled finally gets it's day in the sun albeit with a slim advantage over the other two pooled methods. Bear in mind only 65 instruments have sufficient history here; with only 18 having two distinct blocks of 25 years so there won't be much genuine variation if we choose 50. So this could be a fluke. Jonny will tell you that it is.


But Rob, What about averaging?

At this point, given the choice between the complexity of weight distancing, and the simplicity and efficiency of asset class pooling; I'm inclined to go with the latter. And it's what we were doing at AHL all those years ago (not because of empirical evidence but because it suited the organisational structure...).

However there is another option which I talked about in the original asset class pooling post, using a blend. Here we take an average of the portfolio weights selected with different methodologies. So that would be an average of:

  • Unpooled
  • All pooled
  • Asset class pooled
Blending weights in this way is a way to improve robustness. It's arguably the correct thing to do, since otherwise we'd be making an in sample choice of methodology - 'meta implicit fitting' if you will. One of my favourite research shops, Resolve asset management, are very keen on doing this. One potential downside is it might be producing 'over robustness' given we're using weights that have already had shrinkage. But let's find out.


5 years in sample, 1 year out of sample

Note these numbers won't be exactly the same as those above, since they're a different set of random experiments. They would eventually converge but it would take millions of runs.

And also, just for fun, I've added an extra column. I started off this series of posts talking about the importance of considering other points of the distribution but I've quietly dropped that and only been quoting the median. For balance then, I've added the 25% SR point as well as the median. The pvalue is as before.

                    SR median  SR 25%  pvalue
unpooled                0.025  -0.569     0.0
all pooled              0.475  -0.147     0.0
asset class pooled      0.620   0.045     0.0
average                 0.650  -0.067     NaN

Averaging is the winner - just - but asset class is better at the more conservative point.

5 years in sample, 5 years out of sample

                    SR median  SR 25%  pvalue
unpooled                0.428   0.164     0.0
all pooled              0.430   0.243     0.0
asset class pooled      0.771   0.537     NaN
average                 0.664   0.430     0.0

A clear win for asset class pooled here. Averaging suffers from it's association with the less performative unpooled / all pooled.

10 years in sample, 1 year out of sample

                    SR median  SR 25%  pvalue
unpooled                0.681   0.034     0.0
all pooled              0.452   0.067     0.0
asset class pooled      0.664   0.167     0.0
average                 0.935   0.256     NaN

This time averaging takes the win, helped by the good performance of unpooled.

10 years in sample, 5 years out of sample

                    SR median  SR 25%  pvalue
unpooled                0.887   0.641     0.0
all pooled              0.570   0.406     0.0
asset class pooled      1.066   0.833     NaN
average                 1.005   0.780     0.0

Asset class pooled is still the winner, but averaging gives a good job.

20 years in sample, 1 year out of sample

                    SR median  SR 25%  pvalue
unpooled                0.538  -0.209     0.0
all pooled              0.515   0.154     0.0
asset class pooled      1.167   0.539     NaN
average                 0.809   0.252     0.0

Asset class pooled by more of a margin now.

20 years in sample, 5 years out of sample

                   SR median  SR 25%  pvalue
unpooled                0.243   0.084     0.0
all pooled              0.531   0.414     NaN
asset class pooled      0.513   0.391     0.0
average                 0.461   0.355     0.0

As before 'all pooled' is the winner, whilst average is dragged down by the poor performance of unpooled. But as I said above, with these longer periods it's hard to know if it's just down to flukey instrument selection.

What to do...

There is enough evidence above to justify asset class pooling as the dominant choice. But equally, I don't think there is enough to discard averaging. And there is something so neat about averaging. We combine three quite disparate source of data together, so we're protected if one of them doesn't work out. It's robustness writ large! It can be justified without any in sample fitting - whereas one could argue that the selection of asset class pooling is an implicit in sample 'meta parameter' choice.

I think we're now (finally) ready to fit our forecast weights, and with costs. This is exciting for me, as whatever comes out I will be using as my new weights. This will be more of a 'literature review' since I've talked about optimising with costs in some detail and at some length before.

Wednesday, 17 June 2026

Honey I shrunk the weights (instead of the inputs!)

TLDR: This is a post about something that doesn't work. So don't read if you only care about cherry picked delightful backtests.

This is my fifth post in a rapid fire intense series on portfolio optimisation. In my last post I looked at the optimal amount of shrinkage to use with real data, when running a bayesian methodology for mean variance optimisation. I found two things. Firstly, the optimal shrinkage was different for different sizes of in and out of sample periods. Secondly, that there was mostly a great deal of uncertainty about what the optimium was, with fairly flat surfaces and insigificant t-statistics abounding. I also found that random based methods (monte carlo and bootstrapping) don't work as well as the best shrinkage methods (and in some cases, do worse than the poorest methods). That's three things, but the latter point isn't relevant to this post.

Hence, shrinkage of 0.5 on SR and 0.75 for correlations seemed reasonable; but the truth is we don't really know for sure.

Now I am a big fan of the work of Resolve asset managment. And one thing they are fond of doing is if two or more things seem to work equally well, just taking an average of them (for example they do this here with CTA replication). And I also know intuitively that taking an average of portfolio weights is better than taking an average of inputs. Therefore might we not do better by taking an average of the weights produced by different shrinkage methods?

For example, if we averaged the weights produced by naive mean variance (NMV - zero shrinkage) and equal weights (full shrinkage on both inputs), then we're basically shrinking the weights.

This leaves us with two open questions (apart from the obvious question, which is how long I will continue flogging this subject to death):

  • What are we averaging?
  • What averaging weights should we use?
For the second part I'm going to keep things simple and just use equal weights. For the first part, consider this grid of shrinkage options. This is a subset of what we have seen before:

      0.00  0.50  1.00
0     A       B     C
0.5   D       E     F
1.0   G       H     J


Each row is a different SR shrinkage. Each column is another level of correlation shrinkage. There are 9 options of shrinkage. Some have special names. A is no shrinkage; naive mean variance. B is closest to the optimal shrinkage from the EPO paper I have referenced before. E is not that different from the empirical option I selected in the previous post. J is full shrinkage; equal weights. 

Now if I said to me "Rob, you can only choose two options from this list", I would select:
  • A and J
  • or perhaps, C and G
If allowed three options, I'd throw in E, so:

  • A, E,J
  • C, E, G
With four options I would hit the corners:
  • A,C,G,J
Finally with five options I would hit the corners and the centre:
  • A,C,G,J, E
With everything equally weighted in all of the above. This gives me six different permutations. These in turn can be compared to each of the individual shrinkage options (since we have to calculate them anyway...), so we're comparing 15 possibles.

I'm going to use exactly the same set up as the previous post; randomly chosen portfolios of nine trading rules for a random instrument; varying the size of the in sample and out of sample periods.

Note: yes the title is an allusion to this paper.


One year in sample, one year out of sample

       SR median  SR 0.05  T statistic
A 0.029 -1.531 0.104
B 0.014 -1.549 0.745
C 0.008 -1.598 0.025
D 0.031 -1.527 0.507
E 0.038 -1.573 NaN
F 0.014 -1.572 0.028
G 0.008 -1.729 0.004
H 0.023 -1.704 0.036
J 0.012 -1.694 0.019
AJ 0.013 -1.584 0.031
CG 0.003 -1.668 0.001
AEJ 0.006 -1.603 0.020
CEG 0.002 -1.591 0.001
ACGJ 0.001 -1.610 0.001
ACEGJ 0.005 -1.595 0.001


      0.00  0.50  1.00
0     A       B    C 
0.5   D       E    F 
1.0   G       H    J


Hopefully the format of this table makes sense. The first two columns are median SR across all the random portfolios, and the 5% point of the distribution of random portfolios. You can see that option E is best at the median point (0.5 shrinkage on both), but D is slightly better at the 5% point. The final column is the result of a paired t-statistic comparing the optimal choice (which has NaN in this column) and the choice on the appropriate line. A number below 0.05 means the optimum is significantly better at a 5% critical value, i.e. there is a 95% or more chance it isn't just pure luck.  One benefit of doing this optimisation is that there are fewer options, plus no random methods; so it's very quick. Hence I can get more reasonable t-statistics here (I have 4,000 values in my sample). 

But you can still see that E isn't significantly better than D, and nor is A or B. C, F, H and J are insignificant at a 5% level, but not a 1% level. All the values in the top left quadrant are fine.

You will remember from the last post that the optimum was shrinkage of 0.25 SR, 0.6 correlations but also that there almost no statistical difference between sensible shrinkage results. Option E is closest to that previous optimum. 

Sadly none of the new 'combo' options are any good.

One year in sample, five years out of sample

       SR median  SR 0.05  T statistic
A 0.148 -0.628 0.0
B 0.157 -0.603 0.0
C 0.164 -0.578 0.0
D 0.140 -0.623 0.0
E 0.156 -0.600 0.0
F 0.161 -0.582 0.0
G 0.155 -0.594 0.0
H 0.170 -0.567 0.0
J 0.193 -0.491 NaN
AJ 0.173 -0.549 0.0
CG 0.176 -0.539 0.0
AEJ 0.170 -0.571 0.0
CEG 0.175 -0.561 0.0
ACGJ 0.178 -0.541 0.0
ACEGJ 0.179 -0.542 0.0
      0.00  0.50  1.00
0     A       B    C 
0.5   D       E    F 
1.0   G       H    J

Some amazing significance there - basically equal weights is better than everything by some margin. This is exactly the result from before. And the combos don't perform as well.
 

Five years in sample, one year out of sample


       SR median  SR 0.05  T statistic
A 0.051 -1.742 0.276
B 0.057 -1.769 NaN
C 0.052 -1.763 0.147
D 0.048 -1.738 0.825
E 0.049 -1.757 0.883
F 0.043 -1.754 0.062
G 0.010 -1.842 0.005
H 0.026 -1.767 0.011
J 0.004 -1.803 0.008
AJ 0.025 -1.759 0.031
CG 0.019 -1.789 0.013
AEJ 0.042 -1.757 0.355
CEG 0.034 -1.749 0.043
ACGJ 0.016 -1.747 0.081
ACEGJ 0.037 -1.745 0.047
      0.00  0.50  1.00
0     A       B    C 
0.5   D       E    F 
1.0   G       H    J

Again, the newer combo methods aren't much cop although AEJ is a little better than J.

Five years in sample, five years out of sample


       SR median  SR 0.05  T statistic
A 0.159 -0.666 0.000
B 0.166 -0.648 0.365
C 0.165 -0.642 0.233
D 0.157 -0.666 0.000
E 0.173 -0.660 NaN
F 0.168 -0.647 0.259
G 0.128 -0.676 0.000
H 0.141 -0.667 0.000
J 0.144 -0.660 0.000
AJ 0.162 -0.659 0.051
CG 0.157 -0.653 0.001
AEJ 0.172 -0.667 0.086
CEG 0.162 -0.654 0.013
ACGJ 0.161 -0.655 0.030
ACEGJ 0.163 -0.658 0.054
      0.00  0.50  1.00
0     A       B    C 
0.5   D       E    F 
1.0   G       H    J

Again the middle ground of E is the best; we're also seeing more extreme shrinkage (the bottom row) do very badly as does mean variance. None of the combos do as well.

Ten years in sample, one year out of sample

       SR median  SR 0.05  T statistic
A -0.016 -1.850 0.166
B -0.025 -1.827 0.269
C -0.015 -1.775 0.223
D -0.007 -1.853 NaN
E -0.028 -1.815 0.991
F -0.017 -1.795 0.654
G -0.076 -1.894 0.000
H -0.082 -1.883 0.268
J -0.097 -1.856 0.000
AJ -0.074 -1.846 0.095
CG -0.069 -1.858 0.000
AEJ -0.057 -1.843 0.622
CEG -0.058 -1.844 0.001
ACGJ -0.069 -1.857 0.006
ACEGJ -0.058 -1.840 0.004
      0.00  0.50  1.00
0     A       B    C 
0.5   D       E    F 
1.0   G       H    J

I struggled to get statistical significance for this set before; I have some now, but basically again somewhere in the region of D and E is best. Combo methods do not win though again AEJ isn't significantly worse.

Ten years in sample, five years out of sample

       SR median  SR 0.05  T statistic
A 0.089 -0.825 0.962
B 0.095 -0.865 0.578
C 0.099 -0.848 0.083
D 0.086 -0.842 0.581
E 0.095 -0.852 0.431
F 0.101 -0.848 NaN
G 0.072 -0.923 0.000
H 0.076 -0.930 0.000
J 0.089 -0.939 0.000
AJ 0.082 -0.889 0.000
CG 0.082 -0.908 0.000
AEJ 0.089 -0.874 0.000
CEG 0.091 -0.907 0.000
ACGJ 0.085 -0.896 0.000
ACEGJ 0.088 -0.892 0.000
      0.00  0.50  1.00
0     A       B    C 
0.5   D       E    F 
1.0   G       H    J

'Somewhere in the middle row' isn't a song from Wizard of OverFitting; but roughly where you want to be once again. The combo results are a dismal failure.

Summary

Someone once told me "I love your blog and your books because you talk about failures as well as successes". Well whoever that was - you'll have loved this one! 

Tuesday, 9 January 2024

Skew preferences for crypto degens


An old friend asking for help... how can I resist? Here is the perplexing paper:

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4042239

And here is the (not that senstional) abstract:


Bitcoin (BTC) returns exhibit pronounced positive skewness with a third central moment of approximately 150% per year. They are well characterized by a mixture of Normals distribution with one “normal” regime and a small probability of a “bliss” regime where the price appreciation is more than 100 times at the annual horizon. The large right-tail skew induces investors with preferences for positive skewness to add significant BTC holdings to equity-bond portfolios. Even when BTC is forecast to lose half of its value in the normal regime, investors with power utility optimally add 3% allocations to BTC when the probability of the bliss regime is around 1%. Cumulative Prospect Theory investors are even more sensitive to positive skewness and hold BTC allocations of around 3% when the probability of the bliss regime is 0.0006 and the mean of BTC in the normal regime corresponds to a loss of 90%.


3% in BTC doesn't sound too crazy to me, but what has been really setting the internet on fire is this out of context quote from later in the paper:


Starting with a 60-40 equity-bond portfolio, which is produced with a risk aversion of 𝛾 = 1.50, the optimal BTC allocation is a large 84.9%! The remainder of the portfolio, 15.1% is split 60-40 between equities and bonds. Although BTC has an extremely large volatility of 1.322 (see Exhibit 1), the pronounced positive skewness leads to large allocations and dominates in the utility function (see equation (9)). The certainty equivalent compensation required to not invest in BTC is close to 200%. [my emphasis]


Here's my English translation of this:

- Bitcoin has pronounced positive skew 

- Some people really like positive skew (people  with 'power utility' and 'cumulative prospect theory' preferences)

- This justifies a higher allocation to Bitcoin than they would otherwise have, since it has lots of positive skew (both on an outright basis, and as part of a 60:40 portfolio).

- There is a 'Bliss' regime when Bitcoin does really well ('goes to the moon') but which isn't very likely

- Even if there is a tiny probaility of this happening, and if things are generally terrible in the non bliss regime, then people who like positive skew should have more Bitcoin. Some of them should have a lot!

Now, I could just as easily write this:

- Lottery tickets have (very!) pronounced positive skew 

- Some people really like positive skew 

- This justifies a higher allocation to lottery tickets than they would otherwise have

- There is a 'Bliss' regime when lottery tickets do very well ('winning the jackpot') but which isn't very likely

- Even if there is a tiny probaility of this happening, and if things are generally terrible in the non bliss regime, then people who like positive skew should have more lottery tickets. Some of them should have a lot!

I see nothing here that I can argue with (sorry Ben)! And it certainly doesn't require an academic to make the argument that people who like lottery ticket type payoffs, and think that there is a chance that Bitcoin will go up a lot, should buy more Bitcoin. But I think there is a blogpost to be written about the interaction of skew prefences and allocations; and hopefully one that is perhaps easier to interpret. Two key questions for me are:

- to what extent does the expectation of return distributions affect allocations?

- just how far from 'skew neutral' does ones prefence have to be before we allocate significant amounts to Bitcoin

Luckily, I already have an intuitive framework for analysing these problems, which I used in a fairly complete way in my previous post - bootstrapping the return distribution. 


Setup

The goal then, is to understand the asset allocation that comes out of (a) a set of return distributions and (b) a preference for skew.

For the return distributions we have two broad approaches we can use. Firstly, we can use actual data. Secondly, we can use made up return distributions fitted to the actual data. This is what the paper does, mostly "We use monthly frequency data at the annual horizon from July 2010 to December 2021 for BTC and from January 1973 to December 2021 for stocks and bonds. The univariate moments for each asset are computed using the longest available sample, and the correlation estimates are computed with the common sample across the assets."

The paper also uses a third approach, which is to see what happens if they mess with the return distributions once fitted by changing the probability of 'Bliss'.

I'm going to use the first approach, which is to use real return data at least initially. Other slight differences, I will use returns from July 2010 to November 2023 for all three assets, I will use excess rather than total returns (which given the low interest rates in the period makes almost no difference) with futures prices for S&P 500 (equity proxy) and US 10 year bonds (bond proxy), with Bitcoin total return deflated by US 3 month treasury yields, and I'm going to use daily rather than monthly data to improve my sample size.

The next consideration is the utility preference of the investor. I am going to assume that the investor wants to maximise the Nth percentile point of the distribution of geometric returns. This is the approach I have used before which requires no assumptions about utility function and allows an intuitive measure of risk preference to be used by modifying N. 

As I have noted at length, someone with N=50 is a Kelly optimiser. That is the absolute maximum you should bet, irrespective of your appetite for skew or risk. Thus the Kelly bettor must have the maximum possible appetite for skew. Someone with N<50 would be very nervous about the downside and much more worried about small losses than the potential for large gains; and hence they would have less of a preference for positive skewed assets.

I personally think this is a much more intuitive way to proceed than randomly choosing utility functions and risk aversion parameters, and choosing from a menu of theoretical distributions. The downside is that isn't possible to decompose skew and risk preferences, since both have been replaced with a different measure - the 'appetite for uncertainty'.

An important point is that maximising CAGR will naturally lead to a higher allocation to crypto than you would get from the more classical method of maximising mean subject to some standard deviation constraint or risk aversion penalty. 

The method I will use then is:

- sample the returns data repeatedly to create multiple new sets of data.The new set of data would be the same length as the original, and we'd be sampling with replacement (or we'd just get the new data in a different order). 

- from this new set of data and a given set of possible portfolio allocations, estimate the geometric return

- for a given set of allocations, take the Nth percentile of the distribution of geometric means

- plot the Nth percentile for each allocation to work out roughly where the optimal might be

I say 'roughly', because as readers of previous posts on this subject are aware, we never know exactly where the optimal is when bootstrapping, which is a much better reflection of reality than the precise analytical calculations done by the original authors. Still, we can get a feel for how the optimal changes as we vary N (skew preference).

Note: As a fan of Red Dwarf, the use of the term 'Bliss' in this context is very confusing!


The data


Since we're pretending to be proper academics, here are the summary statistics of the real data:

Annualised mean:
equity 0.256
bonds 0.000
bitcoin 1.280

Annualised standard deviation:
equity 0.160
bonds 0.064
bitcoin 0.816

Correlation:
equity bonds bitcoin
equity 1.000 -0.245 0.069
bonds -0.245 1.000 -0.014
bitcoin 0.069 -0.014 1.000
Sharpe ratio:
equity 0.783
bonds 0.166
bitcoin 1.612

Skew:
equity -0.431
bonds 0.294
bitcoin 0.970

Note that if anything the statistics here are more favourable to Bitcoin than in the original paper. Importantly, we are assuming that as in the past Bitcoin will more than double every year on average (the figure in bold), and that it will have a Sharpe Ratio well north of 1.0. Given these raw statistics, it isn't then very suprising regardless of skew preferences that we would potentially dump a large part of our portfolio into Bitcoin. And indeed, if I run these numbers through my optimisation the optimal position is 100% in Bitcoin for a Kelly maximiser.

To add another line to my 'dumb' bullet point translation of the paper earlier:

- if you think Bitcoin will go up a lot like it did in the past, you should only own Bitcoin

To make things more realistic and interesting, I'm doing to massage the data to reflect what I think is a fairly conservative forward looking position: All assets will have the same Sharpe Ratio (which I will set arbitrarily at an annualised 0.5). I achieve this by shifting the mean returns up or down respectively, which means all the other return characteristics remain the same - only Sharpe Ratio and means are affected. Note that this also means that bonds will look better relative to equities.

This still implies that Bitcoin will, on average, go up by 40% a year, which means it will double every two years. Personally I still think this is extremely optimistic, but I'm going to put my own views to one side for this exercise.

Note: even if you are not a Bitcoin skeptic, it seems unlikely that Bitcoin will behave in the same way going forward as it did when it was worth less than $1,000 and had the market cap of a penny stock rather than a decent sized country; both the mean, skewness, and the standard deviation have reduced in the last few years since Bitcoin has become a bigger market.


Results

Right, let's see some pictures. 

The following heatmap shows what happens to the median of the distribution of bootstrapped geometric returns (Kelly maximiser, with maximum appetite for skew) as we allocate to equities (y-axis) and Bitcoin (x-axis). The allocation to bonds will be whatever is leftover. The white area is where we can't allocate, since we are putting more than 100% into the portfolio, and my working assumption is here is that leverage isn't allowed (if it was, we'd have much more bonds, much less equities and Bitcoin, and use leverage to maximse CAGR).




The optimal allocation to Bitcoin is somewhere around 50% with equities taking most of the rest. So even the most gung-ho optimistic skew loving nutjob shouldn't put more than half their wealth in BTC. For context, a 50% equity and BTC portfolio would have a standard deviation of around 42%, nearly 3 times the risk of equities. To be Kelly optimal, that implies the Sharpe Ratio would need to be at least 0.42. This is a much higher risk target than pretty much every hedge fund uses.

Now let's see what happens if we reduce our N to the 25% percentile point. Importantly: this is roughly the N that produces a 60:40 portfolio considering a portfolio with only equities and bonds. So we can think of this as the 'base case' for risk and skew preference. Again with CAGR below 4% washed out to produce a more granular z axis:




You can see the optimal allocation to Bitcoin is lower here, around 30%, with perhaps 50% in equities and the rest in bonds.


What about N=10%?

Again, we are looking at a bit less again in Bitcoin; with something around 20% with perhaps 70% in equities and the rest in bonds. This would give you something with a standard deviation not much higher than equities, at least in theory.


Summary- ignore everything I have said

The original paper has been toted around the internet to say that you should have 85% of your portfolio in Bitcoin ('this is optimal'). But:

- this is a single figure taken out of context from a much more nuanced paper; note again that the abstract does not include such an extreme figure
- it assumes that historic Bitcoin performance is matched going forward, including performance from 2010 back when BTC cost less than $1 and the total 'market cap' was less than $200,000. 
- it assumes particular risk aversion, preferences for skew and utility functions; such that you would hold quite a lot of Bitcoin even if you thought it's performance would generally be bad except in rare 'Bliss' regimes. Basically it says 'if you like lottery tickets, you are going to love Bitcoin!'.

In this post I take a different approach which hopefully is more intuitive for the non economist, and gives a bit more insight into the interplay between return skew and skew preference, which is also useful beyond the narrow problem of allocating to crypto currency. But what you couldn't or shouldn't do is take anything I or anyone else has written, and claim it 'proves' that the 'optimal' allocation to Bitcoin is x%. All it can do is say based on these assumptions and assuming this set of preferences what your allocation should be. That can quite easily come out to 85%, or 100%. It can also quite easily come out to less than 1%, or even zero.  

What's my own personal allocation to Bitcoin, I hear you ask? On a long only basis it is zero, and nothing I have written here will change that. Partly this is because of my long standing and well known aversion to this 'asset class', both in principle* and in practice**.

* to summarize it's a ponzi that wastes energy with the ownership structure of a pyramid scheme, and which will never be useful for anything except the current use cases: 1% illegal money transfer, 99% gambling
** it's a real pain and very expensive to buy Bitcoin 'properly' i.e. owning your own coins and putting them into cold wallet storage 

But it's also because unlike in this example, there are more than three assets in the world! Concretely, I trade well over 100 futures; of which just a couple are crypto coins. Accordingly it also makes no sense to me to put more than a few % of my trading account into crypto - an account where I can go long and short and hence my personal biases are irrelevant.

My allocation to Bitcoin and Ether in my futures trading strategy is a touch under 5%. And those are risk weights; the equivalent cash weight would be lower: as I write this my position in Bitcoin is long 3 micro futures with a notional value of perhaps £12K or around 3% of my trading capital. Of course it could just as easily be zero, or a short position...

Tuesday, 12 December 2023

Portfolio optimisation, uncertainty, bootstrapping, and some pretty plots. Ho, ho, ho.

Optional Christmas themed introduction

Twas the night before Christmas, and all through the house.... OK I can't be bothered. It was quiet, ok? Not a creature was stirring... literally nothing was moving basically. And then a fat guy in a red suit squeezed through the chimney, which is basically breaking and entering, and found a small child waiting for him (I know it sounds dodgy, but let's assume that Santa has been DBS checked*, you would hope so given that he spends the rest of December in close proximity to kids in shopping centres)

* Non british people reading this blog, I could explain this joke to you, but if you care that much you'd probably care enough to google it.

"Ho ho" said the fat dude "Have you been a good boy / girl?"

"Indeed I have" said the child, somewhat precociously if you ask me.

"And what do you want for Christmas? A new bike? A doll? I haven't got any Barbies left, but I do have a Robert Oppenheimer action figure; look if you pull this string in his stomach he says 'Now I am become Death destroyer of worlds', and I'll even throw in a Richard Feynman lego mini-figure complete with his own bongo drums if you want."

"Not for me, thank you. But it has been quite a long time since Rob Carver posted something on his blog. I was hoping you could persuade him to write a new post."

"Er... I've got a copy of his latest book if that helps" said Santa, rummaging around in his sack "Quite a few copies actually. Clearly the publisher was slightly optimistic with the first print run."

"Already got it for my birthday when it came out in April" said the child, rolling their eyes.

"Right OK. Well I will see what I can do. Any particular topic you want him to write about in this blog post?"

"Maybe something about portfolio optimisation and uncertainty? Perhaps some more of that bootstrapping stuff he was big on a while ago. And the Kelly criterion, that would be nice too."

"You don't ask for much, do you" sighed Santa ironically as he wrote down the list of demands.

"There need to be really pretty plots as well." added the child. 

"Pretty... plots. Got it. Right I'll be off then. Er.... I don't suppose your parents told you to leave out some nice whisky and a mince pie?"

"No they didn't. But you can have this carrot for Rudolf and a protein shake for yourself. Frankly you're overweight and you shouldn't be drunk if you're piloting a flying sled."

He spoke not a word, but went straight to his work,And filled all the stockings, then turned with a jerk. And laying his finger aside of his nose, And giving a nod, up the chimney he rose! He sprang to his sleigh, to his team gave a whistle, And away they all flew like the down of a thistle. But I heard him exclaim, ‘ere he drove out of sight,

"Not another flipping protein shake..."

https://pixlr.com/image-generator/ prompt: "Father Christmas as a quant trader"

Brief note on whether it is worth reading this

I've talked about these topics before, but there are some new insights, and I feel it's useful to combine the question of portfolio weights and optimal leverage into a single post / methodology. Basically there is some familar stuff here but now in a coherent story, plus some new stuff.

And there are some very nice plots.

Somewhat messy python code is available here (with some data here or use your own), and it has no dependency on my open source trading system pysystemtrade so everyone can enjoy it.


Bootstrapping

I am a big fan of bootstrapping. Some definitional stuff before I explain why. Let's consider a couple of different ways to estimate something given some data. Firstly we can use a closed form. If for example we want the average monthly arithmetic return for a portfolio, we can use the very simple formula of adding up the returns and dividing by the number of periods. We get a single number. Although the arithmetic mean doesn't need any assumptions, closed form formula often require some assumptions to be correct - like a Gaussian distribution. And the use of a single point estimate ignores the fact that any statistical estimate is uncertain. 

Secondly, we can bootstrap. To do this we sample the data repeatedly to create multiple new sets of data. Assuming we are interested in replicating the original data series, the new set of data would be the same length as the original, and we'd be sampling with replacement (or we'd just get the new data in a different order). So for example, with ten years of daily data (about 2500 observations), we'd choose some random day and get the returns data from that. Then we'd keep doing that, not being bothered about choosing the same day (sampling with replacement), until we had done this 2500 times. 

Then from this new set of data we estimate our mean, or do whatever it is we need to do. We then repeat this process, many times. Now instead of a single point estimate of the mean, we have a distribution of possible means, each drawn from a slightly different data series. This requires no assumptions to be made, and automatically tells us what the uncertainty of the parameter estimate is. We can also get a feel for how sensitive our estimate is to different variations on the same history. As we will see, this will also lead us to produce estimates that are more robust to the future being not exactly like the past.

Note: daily sampling destroys any autocorrelation properties in the data, so it wouldn't be appropriate for example for creating new price series when testing momentum strategies. To do this, we'd have to sample larger chunks of time period to retain the autocorrelation properties. For example we might restrict ourselves to sampling entire years of data. For the purposes of this post we don't mind about autocorrelation, so we can sample daily data.

Bootstrapping is particularly potent in the field of financial data because we only have one set of data: history. We can't run experiments to get more data. Bootstrapping allows us to create 'alternative histories' that have the same basic character as our actual history, but aren't quite the same. Apart from generating completely random data (which itself will still require some assumptions - see the following note), there isn't really much else we can do.

Bootstrapping helps us with the quant finance dilemma: we want the future to be like the past so that we can use models calibrated on the past in the future, but the future will never be exactly like the past. 

Note: that bootstrapping isn't quite the same as monte carlo. With that we estimate some parameters for the data, making an assumption about it's distribution. Then we randomly sample from that distribution. I'm not a fan of this. We have all the problems of making assumptions about distribution, and of uncertainty about the parameter estimates we use for that distribution. 


Portfolio optimisation

With all that in mind, let's turn to the problem of portfolio opimisation. We can think as this as making two decisions:

  • Allocating weights to each asset, where the weights sum to one
  • Deciding on the total leverage for the portfolio
Under certain assumptions we can seperate out these two decisions, and indeed this is the insight of the standard mean variance framework and the 'security market line'. The assumption is that enough leverage is available that we can get to the risk target for the investor. If the investor has a very low risk tolerance, we might not even need leverage, as the optimal portfolio will consist of cash + securities.

So basically we choose the combination of asset weights that maximises our Sharpe Ratio, and then we apply leverage to hit the optimal risk target (since SR is invariant to leverage, that will remain optimal). 

To begin with I will assume we can proceed in this two phase approach; but later in the post I will relax this and look at the effect of jointly allocating weights and leverage.

I'm going to use data for S&P 500 and 10 year Bond futures from 1982 onwards, but which I've tweaked slightly to produce more realistic forward looking estimates for means and standard deviations (in fact I've used figures from this report- their figures are actually for global equities and bonds, but this is all just an illustration). 

My assumptions are:
  • Zero correlation (about what it has been in practice since 1982)
  • 2.5% risk free rate (which as in standard finance I assume I can borrow at)
  • 3.5% bond returns @ 5% vol
  • 5.75% equity returns @ 17% vol
This is quite a nice technique, since it basically allows us to use forward looking estimates for the first two moments (and first co-moment - correlation) of the distribution, whilst using actual data for the higher moments (skew, kurtosis and so on) and co-moments (co-skew, co-kurtosis etc). In a sense it's sort of a blend of a parameterised monte-carlo and a non parameterised bootstrap.


Optimal leverage and Kelly

I'm going to start with the question of optimal leverage. This may seem backwards, but optimal leverage is the simpler of the two questions. Just for illustrative purposes, I'm going to assume that the allocation in this section is fixed at the classic 60% (equity), 40% (bonds). This gives us vol of around 10.4% a year, a mean of 4.85%, and a Sharpe Ratio of 0.226

The closed form solution for optimal leverage which I've written about at some length, is the Kelly Criterion. Kelly will maximise E(log(final wealth)) or median(final wealth), or importantly here it will maximise the geometric mean of your returns.

Under the assumption of i.i.d. Gaussian returns optimal Kelly leverage is achieved by setting your risk target as an annual standard deviation equal to your Sharpe Ratio. With a SR of 0.226 we want to get risk of 22.6% a year, which implies running at leverage of 22.6 / 10.4 = 2.173

That of course is a closed form solution, and it assumes that:
  • Return parameters are Guassian i.i.d. (which financial data famously is not!)
  • The return parameters are fixed
  • That we have no sampling uncertainty of the return parameters
  • We are fine running at fully Kelly, which is a notoriously aggressive amount of leverage
Basically that single figure - 2.173 - tells us nothing about how sensitive we would be to the future being similar to, but not exactly like, the past. For that we need - yes - bootstrapping. 


Bootstrapping optimal leverage 

Here is the bootstrap of my underlying 60/40 portfolio with leverage of 1.




Each point on this histogram represents a single bootstrapped set of data, the same length as the original. The x-axis shows the geometric mean, which is what we are trying to maximise. You can see that the mean of this distribution is about 4.1%. Remember the arithmetic mean of the original data was 4.85%, and if we use an approximation for geometric mean that assumes Gaussian returns then we'd get 4.31%. The difference between 4.1% and 4.31% is because this isn't Guassian. In fact, mainly thanks to the contribution of equities, it's left tailed and also has fat tails. Left fat tails result in lower Geometric returns - and hence also a lower optimal leverage, but we'll get to that in a second.

Notice also that there is a fair bit of distributional range here of the geometric mean. 10% of the returns are below 2%, and 1% are below 0.4%.

Now of course I can do this for any leverage level, here it is for leverage 2:



The mean here is higher, as we'd probably expect since we know the optimal leverage would be just over 2.0 if this was Gaussian. It comes in at 4.8%; versus the 7.2% we'd expect if this was the arithmetic mean, or the 5.04% that we would have for Gaussian returns.

Now we can do something fun. Repeating this exercise for many different levels of leverage, we can take each of the histograms that are producing and pull various distributional points off them. We can take the median of each distribution (50% percentile, which in fact is usually very close to the mean), but also more optimistic points such as the 75% and 90% percentile which would apply if you were a very optimistic person (like SBF, as I discussed in a post about a year ago), and perhaps more usefully the 25% and 10% points. We can then plot these:


How can we use this? Well, first of all we need to decide what our tolerance for uncertainty is. What point on the distribution are you optimising for? Are you the sort of person who worries about the bad thing that will happen 1 in 10 times, or would you instead be happy to go with the outcome that happens half the time (the median)?

This is not the same as your risk tolerance! In fact, I'm assuming that your tolerance for risk is sufficient to take on the optimal amount of leverage implied by this figure. Of course it's likely that someone with a low tolerance for risk in the form of high expected standard deviation would also have a low tolerance for uncertainty. And as we shall see, the lower your tolerance for uncertainty, the lower the standard deviation will be on your portfolio.

(One of the reasons I like this framing of tolerance is that most people cannot articulate what they would consider to be an appropriate standard deviation, but most people can probably articulate what their tolerance for uncertainty is, once you have explained to them what it means)

Next you should focus on the relevant coloured line, and mentally remove the odd bumps that are due to the random nature of bootstrapping (we could smooth them out by using really large bootstrap runs - note they will be worse with higher leverage since we get more dispersion of outcomes based on one or two bad days eithier being absent or repeated in the sample), and then find the optimium leverage.

For the median this is indeed at roughly the 2.1 level that theory predicts (in fact we'd expect it to be a little lower because of the negative skew), but this is not true of all the lines. For inveterate gamblers at 90% it looks like the optimum is over 3, whilst for those who are more averse to bad outcomes at 10% and 25% it's less than 2; in fact at 10% it looks like the optimium could easily be 1 - no leverage. These translate to standard deviations targets of somewhere around 10% for the person with a 10% risk tolerance . 

Technical note: I can of course use corrections to the closed form Kelly criterion for non Gaussian returns, but this doesn't solve the problem of parameter estimation uncertainty - if anything it makes it worse.

The final step, and this is something you cannot do with a closed form solution, is to see how sensitive the shape of the line is to different levels of leverage, thus encouraging us to go for a more robust solution that is less likely to be problematic if the future isn't exactly like the past. Take a slightly conservative 25% quantile person on the red line in the figure. Their optimium could plausibly be at around 1.75 leverage if we had a smoother plot, but you can see that there is almost no loss in geometric mean from using less leverage than this. On the other hand there is a steep fall off in geometric mean once we get above 1.75 (this assymetry is a property of the geometric mean and leverage). This implies that the most robust and conservative solution would be to choose an optimal leverage which is a bit below 1.75. You don't get this kind of intuition with closed form solutions.



Optimal allocation - mean variance

Let's now take a step backwards to the first phase of solving this problem - coming up with the optimal set of weights summing to one. Because we assume we can use any amount of leverage, we want to optimise the Sharpe Ratio. This can be done in the vanilla mean-variance framework. The closed form solution for the data set we have, which assumes Gaussian returns and linear correlation, is a 22% weight in equities and 78% in bonds. That might seem imbalanced, but remember the different levels of risk. Accounting for this, the resulting risk weights are pretty much bang on 50% in each asset. 

As well as the problems we had with Kelly, we know that mean variance has a tendency to produce extreme and not robust outcomes, especially when correlations are high. If for example the correlation between bonds and equities was 0.65 rather than zero, then the optimal allocation would be 100% in bonds and nothing in equities.

(I actually use an optimiser rather than a single equation to calculate the result here, but in principal I could use an equation which would be trivial for two assets - see for example my ex colleague Tom's paper here - and not that hard for multiple assets eg see here).

So let's do the following; boostrap a set of return series with different allocations to equities (bond allocation just 100% - equity allocation), then measure the Sharpe Ratio of each allocation/bootstrapped return series, and then measure the distribution of those Sharpe Ratios for different distributional points.


Again, each of these coloured lines represents a different point on the distribution of Sharpe Ratios. The y-axis is the Sharpe Ratio, and the x-axis is the allocation to equities; zero in equities on the far left, and 100% on the far right. 
Same procedures as before: first work out your tolerance for uncertainty and hence which line you should be on. Secondly, find the allocation point which maximises Sharpe Ratio. Thirdly, examine the consequences of having a lower or higher allocation - basically how robust is your solution.
For example, for the median tolerance (green line) the best allocation comes in somewhere around 18%. That's a little less than the closed form solution; again this is because we haven't got normally distributed assets here. And there is a reasonably symettric shape to the gradient around this point, although that isn't true for lower risk tolerances.
You may be surprised to see that the maximum allocation is fairly invarient to uncertainty tolerance; if anything there seems to be a slightly lower allocation to equities the more optimistic one becomes (although we'd have to run a much more granular backtest plot to confirm this). Of course this wouldn't be the case if we were measuring arithmetic or even geometric return. But on the assumption of a seperable portfolio weighting problem, the most appropriate statistic is the Sharpe Ratio. 
This is good news for Old Skool CAPM enthusiasts! It really doesn't matter what your tolerance for uncertainty is, you should put about 18% of your cash weight - about 43% of your risk weight in equities; at least with the assumption that future returns have the forward looking expectations for means, standard deviations, and correlations I've specified above; and the historic higher moments and co-moments that we've seen for the last 40 years.



 

Joint allocation

Let's abandon the assumption that we can seperate our the problem, and instead jointly optimise the allocation and leverage. Once again the appropriate statistic will be the geometric return. We can't plot these on a single line graph, since we're optimising over two parameters (allocation to equities, and overall leverage), but what we can do is draw heatmaps; one for each point on the return distribution.
Here is the median:



The x-axis is the leverage; lowest on the left, highest on the right. The y-axis is the allocation to equities; 0% on the top, 100% on the bottom. And the heat colour on the z-axis shows the geometric return. Dark blue is very good. Dark red is very bad. The red circle shows the highest dark blue optimum point. It's 30% in equities with 4.5 times leverage: 5.8% geometric return.
But the next question we should be asking is about robustness. An awful lot of this plot is dark blue, so let's start by removing everything below 3% so we can see the optimal region more clearly:



You can now see that there is still quite a big area with a geometric return over 5%. It's also clear from the fact there is variation of colour within adjacent points that the bootstrapped samples are still producing enough randomness to make it unclear exactly where the optimium is; and this also means if we were to do some statistical testing we'd be unable to distinguish between the points that are whiteish or dark blue. 
In any case when we are unsure of the exact set of parameters to use, we should use a blend of them. There is a nice visual way of doing this. First of all, select the region you think the optimal parameters come from. In this case it would be the banana shaped region, with the bottom left tip of the banana somewhere around 2.5x leverage, 50% allocation to equities; and the top right tip around 6.5x leverage, 15% allocation. And then you want to choose a point which is safely within this shape, but further from steep 'drops' to much lower geometric returns which means in this case you'd be drawn to the top edge of the banana. This is analogous to avoiding the steep drop when you apply too much leverage in the 'optimal leverage' problem. 
I would argue that something around the 20% point in equities, leverage 3.0 is probably pretty good. This is pretty close to a 50% risk weight in equities, and the resulting expected standard deviation of 15.75% is a little under equities. In practice if you're going to use leverage you really should adjust your position size according to current risk, or you'd get badly burned if (when) bond vol or equity vol rises.
Let's look at another point on the distribution, just to get some intuition. Here is the 25% percentile point, again with lower returns taken out to better intuition:




The optimal here stands out quite clearly, and in fact it's the point I just chose as the one I'd use with the median! But clearly you can see that the centre of gravity of the 'banana' has moved up and left towards lower leverage and lower equity allocations, as you would expect. Following the process above we'd probably use something like a 20% equity allocation again, but probably with a lower leverage limit - perhaps 2.



Conclusion

Of course the point here isn't to advocate a specific blend of bonds and equities; the results here depend to some extent on the forward looking assumptions that I've made. But I do hope it has given you some insight into how bootstrapping can give us much more robust outcomes plus some great intuition about how uncertainty tolerance can be used as a replacement for the more abstract risk tolerance. 
Now go back to bed before your parents wake up!