Peptide marketing loves citing “a study.” Almost nobody explains what that study could actually prove, given how it was built. You do not need a statistics background to spot the difference between strong evidence and a small, fragile finding. You need four questions, and a willingness to ask them of every study — including the ones that support a conclusion you like.
Question one: how many people?
This is “n,” the number of participants, and it is the fastest gut-check available. The single adult trial behind most sermorelin claims had nineteen people in it [1]. Nineteen, split roughly in half by sex. A finding in a trial that size can be entirely real and still be fragile. One or two unusual participants can swing a small trial’s result in a way that would wash out in a larger one.
Compare that to the trial that put a GHRH analog head-to-head against real growth hormone, in growth-hormone-deficient children: sixty participants, split across three treatment arms [2]. Still not huge, but more than three times the size. It found a clear, statistically decisive result — actual GH produced faster growth than the analog, at either dose tested, p < 0.01. A bigger n does not guarantee a study is right. It does mean a real effect has more room to show up clearly, and a fluke has less room to hide. See what that size difference means for before-and-after claims specifically.
Question two: what was the study actually powered to detect?
“Powered” means this: before the trial started, researchers calculated how many people they would need to reliably detect the effect they were looking for. A trial can be adequately powered for its main question, and badly underpowered for everything else it happens to measure along the way.
The nineteen-person sermorelin trial measured lean body mass, bone density, sleep quality, well-being, and libido, in men and women separately [1]. That is a lot of outcomes for nineteen people. Split those nineteen into male and female subgroups, then test each subgroup on five separate outcomes, and several of those individual comparisons involve well under ten people. A trial like that can be perfectly well run, and still not have enough people in each slice to reliably catch anything but a large effect.
Contrast that with a modern trial of tesamorelin, a related GHRH-analog drug that is still FDA-approved. It enrolled sixty adults and ran a full year, specifically measuring visceral fat, with enough power to detect a treatment effect and report a precise confidence interval around it: −35 cm², 95% CI −58 to −12, p = 0.003 [3]. That is what an adequately powered trial in this drug class looks like. It is not what sermorelin has, which is part of why the honest answer to “does it work” is mixed.
Question three: primary endpoint, or one they happened to also measure?
A study’s primary endpoint is the one outcome it was specifically designed and powered to test, usually declared before the trial started. Everything else measured along the way is a secondary endpoint. Secondary endpoints are far more prone to noise, because the study was not built with enough people to nail them down.
In the nineteen-person sermorelin trial, it is genuinely unclear from the published abstract which single outcome was the pre-specified primary one, versus which were exploratory measures tacked on. That ambiguity is itself a signal. When a study reports five or six outcomes with no clear “this was the one we were testing for,” treat every individual result inside it as closer to a secondary finding than a headline one. That includes the lean-mass result, which is the strongest thing this drug has going for it.
Question four: “no significant difference” does not mean “no effect”
This is the one people get backward most often, in both directions. When the sermorelin trial reported sleep quality “unaffected in both genders,” that is a real, negative finding worth taking seriously. It was directly measured, not just absent from the report. But a null result in a study this small also cannot rule out a real, modest effect the study simply did not have enough people to detect. “Not significant” in a nineteen-person trial means “we did not find enough evidence to be confident.” It does not mean “we proved there is nothing there.”
The reverse matters just as much. A positive finding in a study this small deserves the same caution, not less. The lean-mass result, the one number most marketing repeats, passed a significance threshold in men, in a trial of nineteen people, and has never been tested again in the decades since. “Statistically significant” tells you the result probably was not random noise, at the sample size tested. It does not tell you the effect is large, reliable, or something a different group of nineteen people would reproduce.
Put together: what an unreplicated trial of nineteen actually means
It means real: someone ran a controlled trial, found a genuine signal, and reported it honestly, antibodies and all. It also means nobody has spent the money or time to confirm it since, in a field that has had three decades to do so. Every larger, better-designed trial run in this drug class since then has been run on a different molecule, tesamorelin, not on sermorelin itself. Sellers publishing an actual dose and a real monitoring plan, rather than a headline statistic with none of this context, are listed on the injections board.
That is not proof sermorelin does not work. It is the honest description of how thin the evidence is, underneath a claim that gets repeated as if it were settled. Once you know to ask about n, power, primary endpoints, and what “not significant” actually means, that thinness is visible in about thirty seconds. No statistics course required.