2 COMMA,,INVESTOR @2commainvestor

NotesA hit rate proves nothing

Note · Opinion, not rated

A hit rate proves nothing.

A stock picker who goes seven for eight looks like a genius and is statistically indistinguishable from a coin. This is not modesty. It is arithmetic, and it applies to us with full force.

OPINION · This note is opinion and general education, not investment advice and not a recommendation about any security. The calculations are standard statistics applied to a hypothetical and to this site's own record. Full disclosures →

The claim, aimed at ourselves

This site keeps a public scoreboard. Every call is dated, benchmarked against the index twelve months later, and left on the record. That is unusual and we think it is the right way to publish. But there is a trap sitting inside it, and honesty requires us to spring it ourselves before a reader does: a good hit rate on a small number of calls is worth almost nothing, and no amount of it, over any span we will produce for years, will prove that anyone here can pick stocks. The scoreboard is a discipline. It is not yet evidence, and pretending otherwise would be the first dishonest thing on it.

What a coin does

Start with the null hypothesis every track record has to beat: a person with no skill at all, calling each position as a coin flip against the benchmark. Give that coin eight calls, which is where our board stands today.

The coin scores six or more out of eight 14% of the time. It scores seven or more 3.5% of the time. It runs the table, eight for eight, once every 256 attempts, or 0.4%. Those are not trivial numbers. They mean that if you handed eight calls each to a room of coin-flippers, a meaningful share would walk out looking like analysts worth paying, having done nothing but get lucky in a small sample. A seven-for-eight record, the kind that would get a newsletter breathless subscribers, is the sort of thing chance produces roughly one time in thirty for a person with no ability whatsoever.

The problem is not that a good short-run record might be luck. The problem is that at eight calls, luck and skill produce records that look identical, and nothing in the numbers can separate them.

How many calls it actually takes

The honest question is not whether a record looks good. It is how many scored calls are required before a genuinely skilled operator would stand clearly apart from the coin. This is a standard power calculation, and the answers are humbling.

To distinguish a true 60% hit rate from a coin, with normal statistical confidence, takes about 153 calls. A true 55% picker, which would still be a valuable edge compounded over a career, needs over six hundred. Even a genuinely excellent 70% operator needs roughly 37 before the record alone rules out luck.

Scored calls required to distinguish a given true skill level from a coin, at standard confidence and power.
True hit rateCalls needed
55%, a real but slim edge616
60%, a strong edge153
70%, exceptional37

At a cadence of one call a week, 153 calls is roughly three years, and that only settles the case if the true skill is a robust 60%. For the slim edges that describe most genuinely good investors, the sample needed runs to a decade or more. This is why a fund's three-year track record tells you so much less than the marketing implies, and why the honest answer to whether we can pick stocks is: ask again in several years, and even then read the confidence interval before the point estimate.

The confidence interval nobody prints

Here is the same fact in the form that should accompany every hit rate ever published and almost never does. Suppose we finish our first eight calls at five correct, a 62.5% hit rate that would look perfectly respectable in a headline. The 95% confidence interval around that number runs from about 29% to 96%. The data is consistent with us being slightly worse than a coin and consistent with us being nearly infallible, at the same time. A number whose error bars span from below random to almost perfect is not a measurement. It is a placeholder.

Only sample size shrinks that interval, and only slowly, because it narrows with the square root of the number of calls. The same 62.5% hit rate measured over fifty calls carries a confidence interval of roughly plus or minus 13 points instead of 34. Better, still not tight. Four times the calls buys half the uncertainty. There is no shortcut around this and no clever adjustment that rescues a small sample.

Why we keep score anyway

If the record cannot prove skill for years, why publish it at all? Three reasons, none of which is statistical.

It removes the exit. A dated, public call cannot be quietly forgotten if it ages badly, and the knowledge that every call will be scored changes how carefully each one is made. It disciplines the writer more than it informs the reader, which is the correct order.

It forces the benchmark. Measuring against the index twelve months out kills the two favorite tricks of performance storytelling: quoting the winners and dropping the timeframe that flatters them. You cannot cherry pick a scoreboard that records everything on a fixed clock.

And it compounds into evidence eventually. Every scored call is one more data point toward the sample size that would actually mean something. The record is worthless as proof today and slightly less worthless every week, which is the most any honest track record can claim in its early years.

What to do with anyone's hit rate, including ours

Treat a short record as a character reference, not a performance measurement. It tells you whether someone is willing to be measured, whether they publish the misses next to the hits, and whether they benchmark honestly. Those are real signals about the person, available immediately, and they are the ones worth weighing. The hit rate itself, until the sample is large, tells you mostly about the coin. When a newsletter leads with a gaudy record over a handful of picks, the record is not the reason to trust it. The willingness to show you the whole thing, error bars included, is.

Follow @2commainvestor