A supplement claim is only as good as the study underneath it, and most of the time you can judge that study in about two minutes without any scientific training. Ask seven things: was it done in people or in animals and cells, how many people took part, did it measure something a person would actually notice or only a number in a blood sample, was the result the one the researchers set out to look for, was the trial registered before it started, who paid for it, and has anyone else found the same thing since. This article shows you how to apply that checklist, with the published evidence for why each question matters, and then applies it to our own claims.

First, read what the study actually tested
Almost every disappointing supplement story starts with a mismatch between the sentence on the box and the sentence in the study. The study gave 3 grams a day for twelve weeks to trained young men; the box implies a benefit at 500 mg for everyone. Or the study used an injected form and the product is a capsule. Or the effect appeared after six months and the marketing suggests a week.
So before you judge the quality of a study, write down four things from it: the population (who), the dose and form (what), the duration (how long), and the outcome measured (what changed). If any of those four differ meaningfully from how the product is sold to you, the study may be perfectly good science and still not support the claim being made with it. That gap is where most supplement marketing lives, and it is not a scientific failure at all. It is a rhetorical one.
Question 1: was it done in humans?
A very large share of the ingredient science quoted in supplement marketing was done in cells or in rodents. That research is legitimate and necessary, but it is a hypothesis generator, not proof of a human benefit.
The clearest illustration comes from a systematic review in the BMJ that compared animal experiments with the human trials of the same treatments. Tirilazad, for example, reduced infarct volume by 29% and improved neurobehavioural scores by 48% in animal models of stroke. In patients it was associated with a worse outcome. Corticosteroids for head injury looked beneficial in animals (pooled odds ratio 0.58 for a poor functional outcome) and showed no benefit in clinical trials. The authors concluded that the discordance may reflect bias in the animal studies or the failure of animal models to mimic human disease closely enough.
Practical rule: if the only evidence for an ingredient is a mouse study or a petri dish, treat the claim as an open question, not a finding. In EU food law that instinct is already codified, because the health claims that may legally appear on a supplement are assessed by EFSA against a published set of scientific principles, and the outcome is a public register of what is and is not authorised. We wrote about the legal side of that separately in which supplement label words are legally defined.
Question 2: how many people took part?

Sample size is the single easiest quality signal to check, and it does more than you would expect. A widely cited analysis in Nature Reviews Neuroscience showed that the average statistical power of studies in a whole field was very low, and spelled out the two consequences: overestimated effect sizes and poor reproducibility. Low power does not only mean a real effect may be missed. It also means that when a small study does report a significant result, that result is more likely to be an exaggeration, or simply wrong.
This is why the same ingredient can look spectacular in a 20 person crossover trial and unremarkable in a 400 person randomised trial. The bigger trial is not being pessimistic. It is being precise.
Rough working guide for a nutrient study: fewer than about 30 participants means treat any effect size as provisional; a few hundred participants with a control group is where numbers start to settle down. And note the direction of the bias, because it matters when you are reading marketing: small studies do not err randomly, they err upward.
Question 3: did it measure something you would notice?
There is a real difference between an outcome that matters to you (you ran further, you slept through the night, you got fewer colds) and a surrogate marker (a blood level rose, an enzyme fell, an antioxidant assay improved).
Surrogates are useful. Raising a low blood level of a nutrient is a legitimate goal in itself. But they systematically flatter interventions. A meta-epidemiological study in the BMJ compared trials in six high impact journals and found that trials reporting surrogate primary outcomes reported larger treatment effects (odds ratio 0.51) than trials reporting outcomes that mattered directly to patients (odds ratio 0.76), a ratio of odds ratios of 1.47 (95% confidence interval 1.07 to 2.01) that survived adjustment. The surrogate trials were also smaller: median sample size 371 against 741.
So when a product is sold on "clinically shown to increase X in the blood", the honest translation is: it changed a number. Whether that number translates into anything you would feel is a separate question that the study may not have asked. Our own hydrogen tablets are a good example of a category where much of the published work is short, small and marker based, which is exactly why we describe them factually rather than promising an outcome. We went through that literature in detail in what the hydrogen water research actually says.
Question 4: was this the result they went looking for?
A trial declares a primary outcome in advance. What gets published does not always match. In a cohort study of 102 trials approved by Danish ethics committees, researchers compared the protocols with the published papers: 62% of trials had at least one primary outcome that was changed, introduced or omitted. Half of the efficacy outcomes and 65% of the harm outcomes per trial were reported too incompletely to be used in a meta-analysis, and statistically significant outcomes were about 2.4 times more likely to be fully reported than non-significant ones (4.7 times for harms). Asked directly, 86% of responding trialists denied that any outcomes had gone unreported.
You cannot audit protocols yourself. What you can do is notice the tell: a headline result that sits oddly with the stated aim of the trial, or a paper whose abstract celebrates an outcome that reads like an afterthought.
Question 5: whole group, or a slice of it?
"It worked in women over 50 with low baseline levels" can be a genuine finding or a fishing expedition. A systematic review in the BMJ assessed every subgroup claim in randomised trials published in core clinical journals: of 207 trials reporting subgroup analyses, 64 made a claim about the primary outcome. Judged against ten credibility criteria, 84% of those claims met four or fewer. Only 41% of the claims had clearly prespecified the hypothesis, only 6% had correctly predicted the direction in advance, and only 9% were supported by a statistically significant test of interaction.
The same instinct applies to how the data were analysed. Text mining across the scientific literature has shown that p-hacking, meaning collecting or selecting data and analyses until a non-significant result becomes significant, is widespread. In fairness, that same analysis concluded the effect of p-hacking looks weak relative to real effect sizes, and probably does not overturn conclusions drawn from good meta-analyses. Treat it as a reason to distrust single striking papers, not as a reason to distrust science.
Question 6: was the trial registered before it started?
Prospective registration is the cheapest integrity check in existence, and you can verify it in a minute. Medical journals following the ICMJE recommendations require registration in a public registry at or before the time of first patient consent for enrolment as a condition of publication, and the ICMJE definition of a clinical trial explicitly includes dietary interventions.
A registered trial has a public record of what it intended to measure before it knew the answer. If a supplement study has a registration number, you can compare intention with result. If it does not, you are trusting the authors' memory of their own plan.
Question 7: who paid, and does it show?

Funding is not disqualifying. Someone has to pay for research, and manufacturers fund a great deal of it. But the effect of funding is measurable, and it is not small. A Cochrane methodology review of 75 studies found that industry sponsored drug and device studies more often had favourable efficacy results (risk ratio 1.27, 95% confidence interval 1.17 to 1.37) and more often favourable conclusions (risk ratio 1.34, 1.19 to 1.51) than studies funded by other sources. The reviewers noted that this bias was not explained by the standard risk of bias assessments, and that in industry sponsored studies the conclusions agreed with the actual results less often.
Nutrition is not exempt. A study of 206 articles on soft drinks, juice and milk found that among interventional studies, the proportion with unfavourable conclusions was 0% when all funding was industry funding, against 37% when there was none. The odds ratio for a favourable versus unfavourable conclusion was 7.61 (95% confidence interval 1.27 to 45.73).
So read the funding statement, and weight the conclusion, not the data. Industry funded trials are often well run. It is the interpretation, more than the measurement, that drifts.
Question 8: is this one study, or a body of evidence?
Single studies are the weakest unit of scientific knowledge, and the literature you can see is not the literature that exists. In an analysis of 74 FDA registered antidepressant trials, 31% were never published. Reading only the published literature, 94% of trials appeared positive; according to the FDA's own analysis, 51% were. Comparing the two datasets, the published effect size was inflated by 32% overall.
That is publication bias in one paragraph, and it is the reason a single positive study, however elegant, is a weak foundation. The theoretical version of this argument was made in a much cited 2005 paper arguing that a research finding is less likely to be true when studies in a field are small, when effect sizes are small, and when there is greater flexibility in designs and analyses. Supplement research has all three properties more often than we would like.
The practical move: look for a systematic review or meta-analysis rather than a study. If the only support for a claim is one trial, the honest word is "promising", not "proven".
Now run the test on us
A checklist you never point at yourself is decoration. So, briefly and honestly:
- Creatine. Our creatine monohydrate carries the single authorised EU claim for creatine: creatine increases physical performance in successive bursts of short-term, high-intensity exercise. That claim exists because it survived exactly the kind of scrutiny described above: repeated human trials, a real performance outcome rather than a blood marker, and an assessment against published EFSA principles. It is also narrow, and we do not stretch it.
- Magnesium. Our magnesium complex states contributions to normal muscle and nervous system function and to the reduction of tiredness and fatigue. Those are authorised nutrient function claims, which is a different and lower bar than "this product will make you feel different". They describe what the nutrient does in the body, not what a specific pill will do to your week.
- Shilajit. There is no authorised EU health claim for shilajit, so we make none. The product page describes what it is and how it is tested. Any brand telling you what shilajit does for your energy or hormones is going beyond what the evidence has been judged to support.
If you apply the seven questions to our marketing and find something that fails, that is a real finding and we would like to hear about it.
Frequently Asked Questions
Does a study in a famous journal automatically mean the finding is solid?
No. The subgroup analysis review and the surrogate outcome study cited above both examined trials in core clinical and high impact journals, and still found large credibility problems. Journal prestige tells you about competition for space, not about whether one particular result will hold up.
How many participants is enough?
There is no universal number, because it depends on how large the true effect is and how variable the outcome is. As a reading habit: below roughly 30 participants, treat the size of any effect as unreliable; the published evidence is clear that low powered studies systematically overestimate effects. A few hundred participants with a control group is where you can start taking the magnitude seriously.
Is industry funded research worthless?
No, and treating it that way would leave you with very little nutrition research at all. The measured pattern is that industry sponsored studies more often report favourable results and, more strongly, favourable conclusions. The sensible response is to read the numbers and discount the discussion section, not to discard the paper.
What is the difference between a health claim and a study result?
A study result is one measurement in one population. An authorised health claim in the EU is a wording that has been assessed centrally and entered in a public register, with conditions of use attached (for example a minimum daily amount). A brand may quote a study at you freely, but it may only use the authorised wording on the product. Checking whether a claim appears in the EU register is often faster than reading the study.
Can a supplement work even if the evidence is weak?
Possibly, and that honesty cuts both ways. Absence of good evidence is not evidence of absence. But you are the one paying, and a weak evidence base means you are funding a hypothesis. That is a legitimate choice if you make it knowingly, and it is the reason we would rather describe an ingredient factually than dress a hypothesis as a promise.
Where can I check whether a claim is allowed at all?
The European Commission maintains a public register of authorised and rejected health claims. Search the nutrient, read the exact permitted wording and the conditions of use, and compare it with the sentence on the product you are looking at.
The Bottom Line
You do not need to read methods sections to protect yourself. Ask whether the study was in humans, how many of them there were, whether it measured something you would notice or only a marker, whether the result was the one they set out to find, whether the trial was registered in advance, who paid, and whether anyone has replicated it. Each of those questions maps onto a documented, quantified failure mode in the published literature. Most supplement claims that fall apart fall apart at question one or question two, long before anything technical is needed. And a claim that survives all seven is worth taking seriously, whoever is selling it.
Sources
- Perel P, et al. Comparison of treatment effects between animal experiments and clinical trials: systematic review. BMJ, 2007.
- Button KS, et al. Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 2013.
- Ciani O, et al. Comparison of treatment effect sizes associated with surrogate and final patient relevant outcomes in randomised controlled trials: meta-epidemiological study. BMJ, 2013.
- Chan AW, et al. Empirical evidence for selective reporting of outcomes in randomized trials: comparison of protocols to published articles. JAMA, 2004.
- Sun X, et al. Credibility of claims of subgroup effects in randomised controlled trials: systematic review. BMJ, 2012.
- Head ML, et al. The extent and consequences of p-hacking in science. PLoS Biology, 2015.
- ICMJE. Clinical Trial Registration, Recommendations for the Conduct, Reporting, Editing and Publication of Scholarly Work in Medical Journals.
- Lundh A, et al. Industry sponsorship and research outcome. Cochrane Database of Systematic Reviews, 2017.
- Lesser LI, et al. Relationship between funding source and conclusion among nutrition-related scientific articles. PLoS Medicine, 2007.
- Turner EH, et al. Selective publication of antidepressant trials and its influence on apparent efficacy. New England Journal of Medicine, 2008.
- Ioannidis JPA. Why most published research findings are false. PLoS Medicine, 2005.
- EFSA NDA Panel. General scientific guidance for stakeholders on health claim applications (Revision 1). EFSA Journal, 2021.
This article is general information about reading research, not medical advice. Food supplements are not a substitute for a varied and balanced diet and a healthy lifestyle.


