StockLearnGuides › Why a t-statistic below 2.4 means we don't claim it
Methodology

Why a t-statistic below 2.4 means we don't claim it

The conventional threshold assumes you ran one test. We run many. Here's the correction, and what it cost one of our own articles.

By ClusterMicro · Updated 2026-08-02 · 6 min read · Research & education

Every research guide on this site reports a t-statistic alongside its headline number, and several of them end with some version of "this didn't clear our bar, so we're not claiming it." This piece explains what the bar is, why it sits at 2.4 rather than the conventional 2.0, and what happens when one of our own results lands just underneath it.

What a t-statistic is, without the algebra

Suppose we measure the average forward return of stocks showing some condition, and it comes out at −0.4%. The obvious question is whether −0.4% is a real effect or just the kind of number you get from noise. The t-statistic answers that by comparing the size of the effect to the size of its own uncertainty.

Roughly: t = effect ÷ standard error of the effect. A t of 1 means the effect is about the same size as your uncertainty about it — indistinguishable from nothing. A t of 3 means the effect is three times its own error bar, which is much harder to explain away as luck.

Method

Effects are mean forward excess returns at 3, 10 and 20 sessions, market-matched over the same window. The standard error is computed across the full set of qualifying stock-days. Horizons are fixed before the test, not chosen afterwards.

Why 2.0 isn't enough

The textbook threshold is roughly t = 2, which corresponds to about a one-in-twenty chance of seeing a result that large under pure noise. That's fine if you run one test. We don't run one test.

Across our component library we test many conditions, at several horizons each. If you run forty independent tests on pure noise, you should expect about two of them to clear t = 2 by luck alone. Publish those two and you've published nothing, dressed as a finding. This is the multiple comparisons problem, and it's the single most common reason retail-facing "research" is wrong.

Raising the bar to 2.4 is a blunt correction. It doesn't fully solve multiple comparisons — nothing simple does — but it materially cuts the rate at which noise gets promoted to a claim, at the cost of occasionally missing something real. We prefer that trade.

What happens when a result lands underneath

This isn't hypothetical. One of our published research pieces originally described a relationship between relative volume and forward returns as significant at the ten-session horizon. When a reviewer pushed back, we re-ran the decomposition properly.

The headline result held and got stronger: the top relative-volume quintile reverted over three sessions with a t-statistic around −3.6, comfortably clear. But the specific ten-session claim came back with the relative-volume-to-excess-return relationship at t ≈ 2.34 — below our own bar.

So we rewrote it. The article now says the ten-session relationship is around t = 2.3, explicitly under the threshold, and describes the pattern as a gradient rather than the cleaner two-populations story the original draft implied. The headline stayed because the headline was solidly supported. The weaker claim got demoted to what the data actually showed.

The rule we try to follow

If a number clears the bar, we state it. If it doesn't, we either say so explicitly with the t-value attached, or we don't make the claim at all. What we try never to do is report a below-bar result in language that sounds above-bar.

Null results are findings

A test that comes back with t = 0.4 has told you something genuinely useful: whatever folklore surrounds that indicator, it isn't visible in this data over this period. Almost nobody publishes these, because they make poor marketing, which is exactly why they're valuable — the published record is skewed towards positive findings by selection, not by truth.

Several of our guides exist to report a null or a negative. The MACD crossover measurement came back negative. Strong technical setups underperformed over short horizons rather than outperforming. These are the pieces we'd most like readers to take seriously.

What the bar does not do

Clearing t = 2.4 does not mean an effect will persist, that it's large enough to be worth acting on after costs, or that it applies to any individual stock. It means the historical average difference was unlikely to be noise over the period measured. That's a narrow claim, and we try to keep it narrow.

Statistical significance and practical significance are different questions. A reliably measured −0.1% edge is real and almost certainly useless. We report the size alongside the t so you can judge both.

Key takeaways

  • t = effect divided by its own standard error; it asks whether a number is bigger than its uncertainty.
  • The textbook t = 2 assumes a single test — run forty on noise and about two clear it by luck.
  • Our bar is 2.4: a blunt correction that trades some missed real effects for far fewer false ones.
  • A published article of ours was rewritten when a claim came back at t = 2.34, below the bar.
  • Statistical significance is not practical significance — we report effect size alongside the t.

See these ideas on real stocks

StockLearn runs this read on ~2,000 NSE stocks every evening. Nifty 50 is free, no login.

Browse today's scan →

This guide is educational and explains how StockLearn interprets common technical indicators, using illustrative examples. It is not investment advice or a recommendation to buy or sell any security.