Skip to content
Archive
← r/algotrading
1
100%
u/Grand-Fly-6090 2 days ago Strategy

White’s reality check crapped all over 1.4 million models

Finally ran white’s reality check on my big mean reversion parameter sweep. The adjusted reality check p value for the entire run is 0.58. This is very disappointing. I used 0% return for the null. I tried running Hansen’s superior predictive ability. The p value was 0.65. Filtering out low returns and high drawdowns improve p to 0.47. I clearly haven’t found an edge.
12 comments held Reddit says 0 on reddit ↗
  1. u/golden_bear_2016 1 2 days ago
    I prefer black's reality check imo
  2. u/Grand-Fly-6090 OP 1 2 days ago
    What should I be using?
  3. u/Many-Pick5066 1 2 days ago
    the 0.47 is the number to throw out. filtering by return and drawdown and then re-running the check pulls candidates out of the max statistic, and the model you filter out is never the winner. so the null distribution drops and your observed max sits exactly where it was. p can only fall. that direction is known before you run it, which means the 0.47 isnt measuring anything. the SPA landing above the RC is backwards from what its for. hansen recenters so the dead models stop inflating the null, and that normally pulls it under white's. check which variant you computed. the l/c/u versions spread out a long way and the u bound is close to a studentized reality check. also 1.4 million models off a parameter sweep is not 1.4 million bets. neighbouring parameters are almost the same strategy and the bootstrap already sees that correlation, so you paid a much smaller penalty than the count makes it sound. 0.58 against a 0% null, before costs, is just a weak family.
  4. u/Grand-Fly-6090 OP 1 2 days ago
    Thank you!
  5. u/Grand-Fly-6090 OP 1 2 days ago
    |Test|p-value| |:-|:-| |White RC|0.5747| |SPA lower, liberal|0.5419| |SPA consistent, recommended|0.6539| |SPA upper, conservative|0.6969|
  6. u/Grand-Fly-6090 OP 1 2 days ago
    Thank you! On Hansen, I computed three p values: |Test|p-value| |:-|:-| |White RC|0.5747| |SPA lower, liberal|0.5419| |SPA consistent, recommended|0.6539| |SPA upper, conservative|0.6969| Yes, the 0.47 is not right. I was trying to explore removing poor models before analyzing to see the impact on p.
  7. u/Effective_Manager273 1 2 days ago
    honestly this is the best outcome you could have gotten from that sweep, even though it does not feel like it. a p of 0.58 across 1.4 million models is the test doing exactly what it exists to do. one thing worth checking before you write the whole family off. white's reality check compares your best model against the distribution of the best model under the null, so with a sweep that size the benchmark you have to beat is brutal, and a real but small edge inside one corner of the parameter space can get buried by the other 1.39 million junk configs you dragged along. if you have a prior about which region should work, run the test on just that block as a pre-registered subset rather than the whole sweep. that is not p-hacking as long as you decide the block before you look. the other thing is your null. 0% return is a weak null for mean reversion on equities, because a lot of mean reversion strategies inherit a long bias and the underlying drifted up over the sample. try the test against a matched random-entry benchmark with the same holding period and exposure. mine looked considerably worse against that than against zero.
  8. u/Grand-Fly-6090 OP 1 2 days ago
    I thought that I would start easy with a zero return null. I have some parameter sets that I’m running live. I could try those families to see how they fair.
  9. u/kestrel_42 1 2 days ago
    costs re-rank a sweep, they don't shift it down evenly. the gross-best configs are usually the highest-turnover ones, so the argmax moves once you charge per fill. rerun net of costs, take the max from that, then run the check on it. good chance it's a different model than the one you tested. on the 1.39m junk configs dragging the max around, worth checking whether what survives is a plateau or a single cell. neighbouring params are nearly the same strategy, so a plateau is structure and a lone spike is the search finding noise
  10. u/Grand-Fly-6090 OP 1 2 days ago
    Thank you!
  11. u/[deleted] 1 2 days ago

    [removed] — already gone when the archive first saw it

  12. u/[deleted] 1 2 days ago

    [removed] — already gone when the archive first saw it