Written by the Inclusive Developmental and Therapy Center therapy team · medically reviewed by Dr Muhammad Suffyan, MB BS (GMC 8023727) · Last reviewed July 2026
Evidence-based practice means combining three things — the best available research, the clinician’s own judgement, and your family’s values and circumstances — and a claim is only as strong as the study design behind it. Knowing how to weigh a claim is one of the most protective skills a family can have, because a child with additional needs attracts confident promises: a programme that guarantees a breakthrough, a headline saying a food ‘causes’ or ‘cures’ something, a brochure citing ‘studies’ you cannot see. For a worked example of weighing claims in practice, our guide to food, feeding and nutrition goes through the diet research for autism study by study.
This guide explains what evidence-based practice actually means, how researchers rank the strength of different kinds of study, and the handful of ideas — control groups, reliability, validity, effect size — that separate a trustworthy finding from a flashy one. It is written for an interested parent, but the definitions are precise enough for a student or trainee to rely on, and it sits alongside the rest of our plain-language guides to child development. The aim is not to make you cynical about research. It is to make you a fair, confident judge of it.
Is ‘evidence-based’ just a marketing phrase?
Evidence-based practice is not simply doing whatever the research says. It means bringing three things together: the best available research evidence, the clinician’s own expertise and judgement, and the values, culture and circumstances of the family. Remove any one of the three and the decision is not evidence-based, however many studies are quoted at you.
That definition has real consequences. Research describes averages across a group, never the individual child in front of you — which is what clinical judgement is for. And the best-supported intervention in the world is the wrong choice if it clashes with a family’s routines, language, resources or the child’s own distress. A therapy a family cannot sustain is not evidence-based for that family merely because a journal supports it.
So when you hear the phrase, the fair question is not ‘is there research?’ but ‘is there good research, read by someone with the judgement to apply it, in a way that fits this child and this home?’
Are all studies equally strong? The hierarchy of evidence
No. Researchers picture the relative strength of study designs as a pyramid, usually called the hierarchy of evidence. Designs higher up the pyramid do more to protect against bias and chance, so they carry more weight. Knowing the ladder lets you judge a claim by the kind of study behind it before you read a single result.
Near the bottom sits the case study or case series — a detailed description of one child, or a handful. These are valuable for describing something new and generating questions, but they cannot show that a treatment works, because there is nothing to compare against. A step up are observational studies, such as cohort and case-control designs, which follow or compare groups without assigning who gets what. They can reveal associations across many people, but because the groups differ in other ways, they cannot firmly establish cause.
Higher still is the randomised controlled trial, the strongest single design for testing whether a treatment actually causes an effect. At the top sit the systematic review and the meta-analysis. A systematic review uses an explicit, repeatable method to find and appraise every relevant study on a question; a meta-analysis statistically combines their results into one more precise estimate. Because they pool many studies rather than resting on one, well-conducted reviews sit at the highest tier. One caveat is worth carrying with you: a review is only as good as the studies inside it. Pooling weak studies does not manufacture strong evidence.
Why is a randomised controlled trial the strongest single study?
A randomised controlled trial compares people who receive the treatment with a control group who do not, and decides who goes into which group purely at random. That random allocation spreads the other differences between children — age, severity, family circumstances, motivation — evenly across the groups, so a difference in outcome afterwards can fairly be credited to the treatment.
The control group is doing the harder work than most people realise. Without one you cannot know whether children would have improved anyway, through ordinary maturation, extra adult attention, or simply the passage of time. Randomisation then handles confounding: a confounder is a hidden third factor linked to both the treatment and the outcome, which can create a convincing impression of cause where there is none. Statistical adjustment can only correct for confounders someone thought to measure; randomisation balances the unknown ones too.
Many good trials add blinding, meaning the people involved do not know who received the real treatment. In a single-blind study the participants do not know; in a double-blind study neither the participants nor the people scoring the outcomes know. Blinding matters because expectation is powerful — a parent or an assessor who believes a child received a promising therapy may genuinely perceive more improvement. Blinding stops hope from masquerading as a result.
Reliability or validity: which question are you asking?
Whenever a study measures something — a child’s language, attention or behaviour — ask two separate questions about that measurement. Is it reliable, and is it valid? The words sound interchangeable and are not, and a measure can easily have one without the other. The same two questions apply to the standardised tests used in an assessment and diagnosis, which is why a good report explains what each test can and cannot tell you.
Reliability is consistency. A reliable measure gives the same answer when nothing has truly changed: the same result if the same child is tested twice in a short window (test–retest reliability), or if two examiners score the same performance and agree (inter-rater reliability). An unreliable measure is a bathroom scale that reads differently each time you step on it — no single reading can be trusted.
Validity is accuracy: does the measure capture what it claims to? A test can be perfectly reliable and completely invalid — a scale that always reads six kilograms heavy is beautifully consistent and consistently wrong. In child development this matters constantly. A questionnaire may measure something very steadily, but is that something really ‘anxiety’, or is it shyness, or an unnoticed language difficulty? Reliability is hitting the same spot every time; validity is hitting the actual target. You need both, and reliability alone is never enough.
Why small studies and strong correlations mislead people
Sample size is the first thing to look for — how many children took part. Small studies are wobbly: with a handful of participants, an impressive result can easily be a fluke of who happened to be included. Larger samples give steadier estimates and make chance a less plausible explanation. When a dramatic claim rests on five or ten children, treat it as a lead worth following, not a conclusion to act on.
The second idea is the most important in this guide: correlation is not causation. Two things can rise and fall together without either causing the other. Across a summer, ice-cream sales and drowning rates both climb — but hot weather independently drives both. The same trap is everywhere in child development. If children who receive a particular therapy tend to do better, it may be the therapy, or it may be that the families who can reach that therapy differ in income, time, language or support, and those are the real drivers. Only a design that controls for the other factors can turn an association into a claim about cause.
Does the effect matter, or does it only exist?
Statistical significance and effect size answer different questions. Significance, usually reported as a p-value, only says a result is unlikely to have arisen by chance. Effect size says how big the difference actually is. With a large enough sample, a change far too small to matter in a child’s daily life can still be reported as ‘statistically significant’.
That is the gap marketing lives in. A ‘statistically significant’ result (conventionally a p-value below 0.05) means the finding is probably not a fluke. It says nothing at all about whether the effect is big enough to be worth your time, your money or your child’s patience.
Clinical significance is the question families actually care about: is the change large enough to make a real difference to everyday life — being understood by a grandparent, sitting through a lesson, asking for what he wants? A result can be statistically significant and clinically trivial. The mature reader wants both: confidence that the effect is real, and evidence that it is big enough to be worth doing. A headline that says ‘significant’ almost always means the first, and quietly hopes you will assume the second.
How do you read a bold claim without a science degree?
You do not need statistics to appraise a claim sensibly. Five questions do most of the work. What kind of study is behind it — a single case, or a controlled trial and a review? Was there a control group and random allocation, or only a before-and-after story? How many children took part? Is the effect merely statistically significant, or big enough to matter in real life? And is the source independent of the people selling the product?
The same five questions work whatever is being offered — a speech programme, a sensory and occupational therapy package, or a behaviour programme for attention and ADHD. The subject changes; the standard of proof does not.
Some warning signs deserve a pause: a cure or a breakthrough that is said to work for everyone; testimonials offered in place of controlled evidence; the word ‘studies’ with no way to check them; one treatment sold as effective for a long list of unrelated conditions; and a package of sessions priced and sold before anyone has assessed your child. None of these proves a therapy is worthless. Each is a reason to slow down and ask for better evidence.
One more rule protects you in the other direction. No evidence is not the same as evidence of no effect. Plenty of sensible, ordinary things have never been trialled, and very little research has been done in Urdu, Saraiki or Punjabi at all. The honest phrase is ‘not studied’, not ‘does not work’. What should worry you is not a thin evidence base — it is a confident promise built on top of one.
Why do testimonials for unproven treatments sound so convincing?
Testimonials feel convincing because they leave out everything a trial controls for. A child who improves after an unproven treatment may have been developing anyway; may have been in a difficult patch that was always going to pass; and may be scored more generously by adults who are hoping hard. Only a control group can separate those explanations from the treatment itself.
Two further biases stack on top. Families who saw no change rarely post about it, so you only ever meet the successes — and any treatment started at a child’s worst moment tends to be followed by improvement, because extremes drift back towards the average on their own. Add the cost and effort already spent, and it becomes genuinely hard for anyone to judge their own experience fairly. This is not gullibility. It is how human judgement works, which is exactly why controlled research exists.
Reviews of the research have not found good-quality evidence that the treatments most heavily marketed as cures for autism — stem-cell procedures, restrictive elimination diets, high-dose supplement protocols — change a child’s autism, and some carry real physical risk alongside real cost. If you are considering anything invasive, or anything that removes food groups from a child’s diet, speak to your paediatrician first. A responsible clinician will never promise a cure, and should be able to tell you plainly what their approach cannot do. Our guide to autism and social communication sets out the support that is actually established.
Your own child is a study of one
All of this can feel abstract until you apply it to your own home, where there is no control group and no randomisation — just one child, and a change you either can or cannot see. The way to make that judgement fair is a baseline: a record of where things actually stood before anything started. Countable things work best. Words used without prompting in ten minutes of play; sounds produced clearly; instructions followed first time; minutes of settled attention at a table; times he asked for something instead of pulling your hand.
Without a baseline, ‘he seems a bit better’ can never be tested, and everyone involved has a reason to believe it. With one, progress becomes something that can be shown rather than felt. Agree in advance what would count as progress and when you will both look again — a fixed review point, not an open-ended arrangement. It is also worth asking what usually shifts first, because it is rarely the headline goal: understanding tends to move before talking, attempts and imitation before accurate production, and skills appear in a familiar room before they appear at school.
The same honesty has to work in reverse. If nothing has moved against the baseline after a fair block of work, the answer is to change the method, change the goal, or refer on — not to ask for more of the same. In our sessions in Multan that is also where hearing gets revisited: unclear speech and patchy responding can be an ear problem rather than a speech one, and a hearing test is a reasonable early step. If you are worried about your child’s development, that worry is worth acting on now rather than waiting for certainty — speak to your paediatrician or a qualified therapist, and ask them to show you the measurements behind whatever they recommend. If you would like to organise what you are seeing first, our free child development check is a plain-language place to start.
Where can you look a claim up yourself?
You can check most claims yourself in about ten minutes, in plain English and without a subscription. Cochrane publishes plain-language summaries of its systematic reviews. The NHS, the CDC, NIDCD and ASHA publish accessible summaries of what is established for speech, language and developmental conditions, and NICE publishes clinical guidance. Searching the method’s name together with the words ‘systematic review’ is usually enough to see whether one exists at all.
Two habits make that search far more useful. First, search the name of the method, not the name of the clinic — a named approach can be looked up, whereas ‘our own special technique’ cannot, and that is often the whole point of the vagueness. Our guide to the main therapy approaches and how they are meant to work names the methods you are most likely to be offered. Second, when you find a study, read the abstract for the design and the numbers rather than the conclusion: how many children, over how long, compared with what. ‘Peer-reviewed’ means independent experts checked the methods before publication. It is a useful quality filter, not a guarantee that the study got everything right.
Key takeaways
- Evidence-based practice stands on three legs together — the best research evidence, clinical expertise, and the family’s values and circumstances. Research alone is not enough.
- The hierarchy of evidence runs from case studies (weakest) up through observational studies and randomised controlled trials to systematic reviews and meta-analyses (strongest), because higher designs better protect against bias and chance.
- A randomised controlled trial establishes cause through a control group plus random allocation, which balances confounders the researchers never thought to measure; blinding stops expectation from masquerading as a result.
- Reliability is consistency (the same answer each time); validity is accuracy (measuring what you claim). A measure can be reliable yet invalid, and reliability alone is never enough.
- Correlation is not causation: two things moving together may share a hidden common cause, so only a controlled design can turn an association into a claim about cause.
- A p-value only says a result is probably not chance; effect size and clinical significance say whether it is big enough to matter. Always ask for both.
- No evidence is not the same as evidence of no effect — the honest phrase is ‘not studied’. What should worry you is a confident promise built on a thin evidence base.
- For your own child, a written baseline and a fixed review point do the job a control group does in a study: they turn ‘he seems better’ into something you can actually check.
For students & professionals
A few deeper points worth knowing if you’re studying this area — think of it as a study aid, not a replacement for your course or supervisor.
- Memorise the hierarchy of evidence and why it is ordered as it is: each step up (case study → cohort/case-control → randomised controlled trial → systematic review and meta-analysis) adds protection against bias and chance. A meta-analysis gives a more precise estimate but inherits the weaknesses of the studies it pools.
- Be able to explain randomised controlled trial logic precisely: the control group supplies the comparison, and random allocation balances known and unknown confounders across groups, so a later difference in outcome can be attributed to the intervention.
- Distinguish reliability from validity with the target analogy, and know the subtypes you will be examined on — test–retest and inter-rater reliability — plus the principle that reliability is necessary but not sufficient for validity.
- Define a confounder as a variable associated with both the exposure and the outcome that can produce a spurious association, and be able to say why randomisation, not statistical adjustment, is the stronger defence against unmeasured confounders.
- Keep effect size and the p-value firmly separate. A p-value addresses only whether an effect is likely to be non-chance; effect size quantifies magnitude — and in a large sample a trivial effect can reach statistical significance without being clinically significant.
- Know the difference between a systematic review and a narrative literature review: the systematic review pre-specifies its question, search strategy, inclusion criteria and appraisal method so that another team could repeat it, which is what keeps selection bias in check.
- Learn to phrase null findings correctly. ‘No evidence of effect’ and ‘evidence of no effect’ are different statements, and an underpowered study can produce the first while telling you almost nothing about the second.