Part 2 — The How
Study Protocol and Basic Statistics
The most important decisions in a clinical trial are made before it begins — and written down before anyone knows whether the treatment works. Which outcome will count, how many participants it takes to believe the answer, what result will be called real: the protocol fixes every one of these in advance, precisely so that no one can choose the answer after seeing the data.
That is what this final module asks you to see in the document you work from every day. The protocol is not paperwork around the science; it is the science, pinned down in advance, so the answer deserves belief.
The module reads the protocol in three passes. First the document: what it is and what it must contain. Then the design decisions that protect the answer — endpoints, randomisation, blinding, controls, and who gets enrolled. Then the statistics that size the trial and judge its result. It closes with faithfulness in practice — adherence and amendments — and the red flags to catch when you read a protocol.
Learning Objectives
After this module, you can:
- Read a protocol as what it is — scientific plan, ethical contract and regulatory submission — and check its content against E6(R3) Appendix B
- Judge a trial's endpoints — primary, secondary, surrogate — and recognise what an estimand pins down before the data exists
- Explain how randomisation, blinding, the choice of control and the choice of participants protect a trial's answer from bias — and decide when a placebo arm is ethically defensible
- Interpret p-values, confidence intervals, power and sample size well enough to respect them — and spot when a number is being misread
- Hold a running trial to its protocol: document deviations, route changes through amendments, and catch the red flags when you read one
The Protocol Is Three Promises at Once
Ask three people what the protocol is and you get three honest answers. The statistician calls it the trial's scientific plan. The ethics committee calls it the thing it approved. The authority calls it the thing it authorised. All three are right, and each answer binds your site differently.
As a scientific plan, the protocol commits the trial to one question and one way of answering it — design, endpoints, analysis — before any data can tempt anyone to prefer a different question. Swiss law makes that scientific quality a duty held jointly by the sponsor and the investigator: a research question grounded in the current state of knowledge, an appropriate methodology, and the resources to carry it. KlinV
As an ethical contract, the protocol is the specific document the ethics committee weighed when it said yes — Module 5's ground. The approval attaches to that version, its risks and its safeguards; a trial run differently is a trial the committee never approved.
As a regulatory submission, it is Swiss ground: the protocol sits in the application dossier the ethics committee reviews, and in the dossier Swissmedic authorises for the trials that need its approval. KlinV
One part of the protocol deserves special respect from the start: the analysis. The statistical principles for trials are built on prespecification — the analysis is described before the data are unblinded, and Module 12 taught the lock that enforces the deadline. An analysis invented after the results were visible, presented as if it had been the plan, is not a shortcut; it is choosing the answer. ICH E9
What a Protocol Must Contain
A protocol your site cannot actually run is not a scientific problem waiting to happen — it is a deviation list waiting to happen. E6(R3) Appendix B sets out what a trial protocol and its amendments contain, from general information through background, objectives, design, participant selection, treatment, assessments, statistics, direct access, quality control and ethics. ICH E6(R3)
The full list is the sponsor's writing job. Your reading job is narrower and sharper: before your site commits, check the parts that will govern your daily work — and check them against each other. A protocol whose sections disagree makes every team member choose a version, and every choice is a potential deviation. Consistency and comprehensibility are not editorial niceties; they are what makes adherence possible. The quality-by-design thinking behind a well-built protocol — identify what is critical, strip what is unnecessary — is Module 8's ground.
| What you check | Where it lives |
|---|---|
| One primary endpoint, stated exactly | B.4.1 |
| The design, with a schematic you could explain to a participant | B.4.2 |
| The measures against bias — randomisation, blinding | B.4.3 |
| Eligibility criteria your clinic can apply the same way every time | B.5 |
| What ends a participant's treatment, and what follows for them | B.6 |
| An assessment schedule that agrees with itself across sections | B.8–B.9 |
| The sample size, with its reasoning | B.10.2 |
| Planned interim looks and their stopping rules, if any | B.10.1 |
Vague wording fails this check as surely as contradiction does. An eligibility criterion two investigators read two ways is not flexible — it is unreproducible, and enrolment decisions will not match across sites. Say so before signature, while the authors can still fix it rather than the users work around it.
Endpoints — and the Estimand Behind Them
Every trial makes one promise it cannot take back: what will count as the answer. That promise is the endpoint — the measured outcome on which the trial stands or falls. The protocol states the primary endpoint, the one the trial is sized and judged on, and any secondary endpoints, stated in advance but subordinate. ICH E6(R3) ICH E9
Why only one primary? Because every additional outcome given primary status is another chance for noise to look like an effect — ask enough questions of one dataset and one eventually says yes by accident. Statistics can handle several primary variables honestly, but only with the adjustments planned in advance — which is why a protocol carrying several, with no plan for the multiplicity, is a protocol built to find something. ICH E9
Some endpoints measure the thing itself: survival, stroke, recovery. Others measure a stand-in — a surrogate endpoint, a variable like blood pressure or tumour shrinkage that substitutes for the clinical outcome because it is faster to observe. A surrogate is a bet that improving the measurement improves the patient. The bet can fail: a treatment can move the surrogate beautifully while leaving the real outcome untouched, or worse. ICH E9 A documented case in the statistics section shows how much can ride on it.
There is one more layer, newer than the others, and E6(R3) names it in the protocol's objectives section: the estimand. An endpoint says what is measured. The estimand pins down the precise treatment effect the trial claims to estimate — in which population, on which variable, summarised how, and crucially, what the comparison means once real life intervenes. ICH E6(R3) ICH E9(R1)
What does "pins down" mean in practice? E9(R1) writes the trial's full question as five plain attributes, all fixed before the data exists — here they are for one chronic-pain trial. ICH E9(R1)
| Attribute | This trial |
|---|---|
| Population | adults with chronic low-back pain |
| Treatment vs comparator | drug X versus placebo |
| Variable (endpoint) | pain score at 12 weeks |
| Intercurrent-event handling | counted regardless of rescue-painkiller use |
| Summary measure | difference in mean scores between arms |
The fourth line is where real life bites. After treatment starts, things happen that muddy the comparison: a participant switches treatment, stops early, or reaches for rescue medication — extra painkillers when symptoms break through despite the assigned treatment. These are intercurrent events, and the card fixes how each is counted, rather than leaving it to be decided after the results are seen. ICH E9(R1)
That is one line of five: change any line — a different population, the mean swapped for the share who halved their pain — and the trial answers a different question. Settling all five in advance is the estimand's whole job, so the analysis and missing-data rules serve that question, not one picked after the results are in. ICH E6(R3)
You will not build estimands; that is the statistician's craft, brought in by R3 from the E9(R1) addendum. What you need is recognition: when a protocol names its estimand and its intercurrent-event strategy, it is fixing the question — and a question fixed in advance is one nobody can quietly swap later.
Two trials test the same drug for depression and name the identical primary endpoint: symptom score at week 8. They differ in one line of the estimand. In Trial A, a participant who stops the assigned drug and starts a different antidepressant still has their week-8 score counted as measured; in Trial B, that same participant is counted as a non-responder. A reviewer says: 'Identical primary endpoint, so the two trials estimate the same treatment effect.' Is that right?
Randomisation and Blinding: The Machinery Against Bias
Suppose the trial's two groups were assembled by human judgment — this patient looks robust, put her on the new drug; this one is frail, spare him. Whatever difference the trial then finds, nobody can say whether the treatment caused it or the sorting did. Bias — systematic distortion of the answer, as opposed to random noise — got there first, at assignment.
The most important design techniques against bias are blinding and randomisation, and the protocol must describe both. In the common parallel-group design each participant is randomised to one arm and stays there; in a crossover design each participant receives the treatments in sequence, as their own comparison. Either way, the protection starts at the same place. ICH E9 ICH E6(R3)
Randomisation assigns participants to arms by chance. Chance is not fair by accident — it is fair by construction. It tends to balance the groups on everything, including the confounders nobody has thought to measure: background differences, like age or undiagnosed disease, that could produce an outcome difference on their own. A human sorter can only balance what they know about. Chance balances the unknown unknowns too, and that is the whole trick: it is what lets the end-of-trial difference be attributed to the treatment. How the sequence is built — in blocks that keep arm sizes level, in strata that balance a key factor like site or severity — is craft on top of the principle. ICH E9
The principle survives only if the sequence is obeyed and the next assignment stays unpredictable. The moment anyone steers — holding a frail patient back for the gentler arm, timing an enrolment to catch a particular slot — selection re-enters through the side door, and comparability quietly dies. If a participant should not receive one of the trial's arms, the honest answer is an eligibility answer, decided before randomisation — never a steered allocation after it.
Blinding protects everything that happens next. People who know the assignment behave differently, with complete sincerity: participants report symptoms differently, clinicians look harder for side effects on the active arm, assessors nudge borderline scores. In a single-blind trial participants do not know their assignment; in a double-blind trial neither participants nor the investigator and site staff do. Where full blinding is impossible — surgery against a pill — the trial can still blind the outcome assessor, so the person scoring the result does not know what they are scoring. ICH E9
Blinding is a site discipline as much as a design feature. The protocol says who is blinded; your site's job is to keep them that way — in conversation, in filing, in who sees what. Breaking a blind for a safety emergency has its own controlled procedures; what design cannot survive is the casual leak.
Controls and the Ethics of Placebo
A treatment group alone proves almost nothing, because diseases wax and wane on their own and people improve under attention. The control group — participants who receive the comparator instead of the new treatment — is what turns "they got better" into "they got better because of it". The comparator can be the best existing treatment (an active control) or a placebo, an inert treatment indistinguishable from the real one. ICH E10
The choice of control also sets the trial's objective. Against placebo, the natural question is superiority — is the new treatment better than nothing dressed as something? Against an active control, the question is often non-inferiority: is the new treatment not meaningfully worse than the standard, while offering some other advantage? Non-inferiority is the harder trial to interpret, because "no difference found" must be distinguishable from "a trial too blunt to find one" — a trial's capacity to detect a difference that is really there is its assay sensitivity, and it is exactly what an active-control design cannot take for granted. ICH E9
Placebo is scientifically clean and ethically expensive. A participant on the placebo arm forgoes the new treatment; in a placebo-only design they forgo active treatment altogether, though add-on designs give standard care plus placebo rather than placebo alone. The Declaration of Helsinki — whose ethical framework is Module 6's ground — prices it explicitly.
A sponsor proposes a double-blind, placebo-controlled trial of a new antibiotic for community-acquired pneumonia. The standard treatment, amoxicillin, is effective and inexpensive. Is the placebo control ethically acceptable under the Declaration of Helsinki?
The Numbers You Must Respect
You do not need to do statistics to run a good site. You need something rarer: enough to respect what the numbers protect, and to notice when someone — a colleague, a paper, your own hopeful reading — is bending them. Three steps: the test, the size, the interval.
One idea sits under all three. A trial measures a sample to estimate a parameter — the true value in the whole population, which nobody observes directly. What the sample actually hands you is a statistic: a number computed from the participants you happened to measure, standing in as the estimate of that truth. Run the same trial again and the statistic moves; the parameter it points at does not. Results scatter around the truth in a distribution, so two groups always differ somewhat even when the treatment does nothing. Every tool below answers one question: is this difference more than the scatter alone would produce?
The distribution is not decoration — it decides the method. Blood pressures ranging symmetrically around a centre, skewed hospital stays, event counts, times to an event: each shape carries different assumptions, and each statistical test is valid only where its assumptions hold. That is why the protocol's statistical section does not just promise an analysis but names one. Where the expected shape would break the chosen method's assumptions, the protocol also pre-specifies the fix — a transformation of the variable — with the rationale written down in advance. ICH E9
Hypothesis Testing and p-Values
A trial does not set out to prove its treatment works. It sets out to embarrass the opposite claim: the null hypothesis (H₀), that the treatment has no effect, and asks how implausible the data would be in a world where H₀ holds.
The measure of that implausibility is the p-value: the probability of seeing a result at least as extreme as the one observed, if the treatment truly had no effect. Small p means the no-effect world would rarely produce data like this; at a pre-set threshold — conventionally 5% — the result is called statistically significant. E9 asks for precise values in reports (p = 0.034, not "p < 0.05"), because the number carries information the label discards. ICH E9
Now the misreadings, because you will meet them weekly. The p-value is not the probability that the treatment doesn't work. It is computed assuming the no-effect world is true, so it can never hand back the probability that that world is true — the condition it starts from cannot be the answer it delivers. Reversing a conditional changes the number: the chance a card is red given it is a heart is 100%, but the chance it is a heart given it is red is only 50%. The p-value is the first kind of statement, never the second.
Two more, same root. The p-value is not the probability the result was chance — chance is the assumption it runs on, not a verdict it returns. And p above 0.05 does not show the treatment has no effect; it shows this trial failed to demonstrate one, which may say more about the trial's size than the drug's worth. The p-value answers exactly one question, about data in a hypothetical no-effect world; every grander claim put in its mouth is someone else's wish.
Power, Sample Size and Who Gets Counted
Every trial verdict risks two distinct mistakes, and the design prices each in advance.
| Truth: no effect | Truth: real effect | |
|---|---|---|
| Trial finds an effect | Type I error — false positive (α) | Correct |
| Trial finds none | Correct | Type II error — false negative (β) |
In plain terms: a Type I error is a false alarm — the trial declares an effect that is not real, and an ineffective treatment can reach patients on the strength of it. A Type II error is the reverse miss — the trial finds nothing when the effect is real, and a working treatment is wrongly read as useless.
E9 sets the conventional prices. The probability of a Type I error (α) is set at 5% or less; the probability of a Type II error (β) at 10% to 20%. And α is fixed in the protocol at design, then left alone: it is not a dial anyone turns after seeing the data.
Power is the plain-words counterpart: the trial's probability of detecting an effect that is really there, at the size the trial assumed. At 80% power, if the treatment truly works as assumed, roughly 8 of every 10 such trials find it and 2 miss it — reading as negative when the effect was real all along. In the arithmetic, power is simply 1 − β, so a 10–20% β is 80–90% power. ICH E9
The sample size is where those prices are paid. From the assumed effect size, the outcome's variability, the protocol's α and the desired power, the statistician computes how many participants the question needs. The protocol must state the number and the reasoning behind it, power calculation and clinical justification included. Watch the first input hardest: the calculation is only as honest as the effect size fed into it, and an optimistic assumption quietly buys a smaller, cheaper, weaker trial. ICH E6(R3)
Two more pieces of protective machinery, briefly. A trial may plan interim analyses — looks at the accumulating data before the end, usually to stop early for clear benefit, futility or harm. Planned is the operative word: the protocol must state their timing, purpose and statistical stopping criteria, because every unplanned peek is another chance for noise to cross the significance line. ICH E6(R3) ICH E9
And when the analysis comes, who gets counted? The intention-to-treat (ITT) principle answers: analyse everyone as randomised, in the arm chance gave them, whether or not they complied, switched or dropped out. The moment you exclude the inconvenient, you unpick the comparability randomisation built. The full analysis set applies that principle as closely as practicable; the stricter per-protocol set — only participants who followed the protocol — can support it. But leaning on per-protocol as the main answer favours the treatment in exactly the trials where adherence went wrong. ICH E9
Confidence Intervals
A p-value says whether; it never says how much. Think of an election poll reporting a candidate at 52%, margin of error ±3. The real finding is not 52 — it is the range 49 to 55, the band the data are consistent with. A trial's result is the same kind of object: a best estimate wearing a band of uncertainty.
The confidence interval (CI) is that band — the range of true effects the data are compatible with. A narrow interval means a precise estimate, usually from a large trial; a wide one means an imprecise estimate from a small one. The "95%" names the method's reliability: intervals built this way capture the true effect 95% of the time. ICH E9
Two things to read off it. First, does the interval cross no effect — a zero difference, or a risk ratio of 1? If it does, the result is compatible with nothing. Second, read both ends. A trial reports a 15% relative risk reduction (the drop in event rate), with a 95% CI from 2% to 26%. That range sits as comfortably with a barely-worth-it 2% as with a substantial 26%, and cannot tell them apart. A significant result whose interval still spans trivial-to-large has proven direction, not size; a "negative" one whose interval still holds important effects has shown imprecision, not absence. The interval is the honest sentence; the p-value is its punctuation.
A trial is powered at 80% to detect a 20% relative risk reduction. It reports p = 0.12 and concludes 'no significant effect.' A colleague argues the treatment might still work and the trial was simply too small to show it. Who is right?
The section closes with the machinery running for real — a case where the endpoint decision, fixed in the protocol, was everything.
Newer Designs, in One Breath
The machinery you just learned is the classical, or frequentist, approach: it judges the data against a fixed threshold, the 5% we met. It is still the backbone of confirmatory trials — the late-phase trials built to prove a hypothesis and support approval, as opposed to the exploratory ones that generate the hypotheses.
Protocols now name designs beyond it, and three trial shapes are worth recognising. A platform trial runs one master protocol in which treatments enter and leave over time against a shared control. An umbrella trial studies one disease through parallel sub-studies split by molecular subtype. A basket trial tests one targeted treatment across several diseases that share its target. E6(R3)'s design section lists these alongside adaptive designs, decentralised elements — trial activities moved to the participant through home visits, telemedicine and local labs — and the familiar double-blind parallel design. ICH E6(R3)
Two terms are worth recognising. A Bayesian design starts from prior evidence, updates it as trial data accumulates, and declares success when the probability of benefit — the posterior probability — crosses a pre-set threshold; the protocol states either a significance level or that Bayesian threshold. An adaptive design pre-plans its own changes — dropping a dose, re-estimating the sample size, stopping arms — with each permitted adaptation and its decision rule in the protocol from the start. ICH E6(R3)
An adaptive trial does not improvise; its changes were pre-specified. The principle just moved up one level, from "no changes" to "no unplanned changes". Recognition is all this course asks; the methods belong to the statisticians.
Designing for the Right Participants
A trial can be randomised, blinded, controlled and powered — and still answer its question for the wrong people: efficacy shown in fit fifty-year-old men stretches thin over the eighty-year-olds in Monday's clinic.
E8(R1) — the ICH guideline on general study design that Module 8 placed in the family — pushes on both fronts. Trial design is best informed by input from a broad range of stakeholders, patients and healthcare providers included. Early engagement catches burdens a sponsor cannot see, and demands that fit real lives recruit and retain better. And later-phase studies should involve participants representative of the diverse populations that will receive the intervention in clinical practice. ICH E8(R1)
Protocol Adherence and Amendments
The protocol you signed in the first section now has to survive contact with a running trial — with sick participants, staff turnover and eighteen months of Fridays. Two disciplines keep the document and the trial pointing at the same thing: adherence when reality slips, and amendment when the document itself must change.
GCP states the adherence half plainly: the investigator should comply with the protocol, and should document all protocol deviations. That covers the ones your site finds and the relevant ones the sponsor's monitoring follow-up communicates back. A deviation is unplanned noncompliance, handled after the fact: documented, reviewed, and for the important ones, escalated and corrected so the same slip stops recurring. ICH E6(R3)
An amendment is the opposite creature: the controlled, prospective way the protocol itself changes. Appendix B governs the protocol and its amendments — an amended protocol is a new version of the same binding document, approved through the same machinery. The approval route runs through the ethics committee before implementation, with the participant-protection exception — Module 9 taught that process. What this module adds is the document discipline at your site. The amended version is re-signed, re-distributed, and in the hands of every team member — because a nurse working from version 3 in a version-4 trial generates deviations with every visit. ICH E6(R3)
The analysis has its own guarded door — a standing confession rule in the protocol's statistical section. Any deviation from the statistical analysis plan (SAP), the pre-specified description of how the data will be analysed, must be described and justified in the clinical trial report. Changes before unblinding are amendments; changes after are disclosures. There is no third, silent category — choosing or redefining the hypothesis or analysis after the results are in and presenting it as the original plan is HARKing (Hypothesising After Results are Known), the exact move the confession rule is built to expose. ICH E6(R3)
Red Flags When Reading a Protocol
Everything this module taught compresses into a reading skill: wherever you meet a protocol, these patterns deserve a raised hand.
- Several "co-primary" endpoints, no multiplicity plan. More primary chances, same 5% threshold each — a design built to find something.
- A surrogate primary endpoint with no acknowledged limits. The stand-in must carry the claim; CAST is the standing reminder of what riding the marker can cost.
- Eligibility criteria that two readers apply differently. Unreproducible criteria make unreproducible populations — and undermine the comparison across sites.
- An effect size in the power calculation that nobody believes. Optimism in, underpowered trial out.
- Interim looks without pre-specified stopping rules. Planned analyses protect the trial; unpriced peeks spend its error budget invisibly.
- Analysis changes after results were visible, presented as planned. The SAP-deviation rule demands description and justification in the report; silence is the flag that outranks the rest.
- A reported primary endpoint that does not match the registered one. The registry entry predates the data — ask which was rewritten, and when.
None is automatically fatal; each is a question the protocol's authors owe an answer to. A weak trial and a bent one differ in whether the answer was written down in advance.
Module Summary
The protocol is where a trial's honesty is decided — before the first participant, in writing, in advance. You have read it as a signed three-way promise; as the container of the machinery — endpoints and estimands, randomisation and blinding, controls and populations — that keeps bias out of the answer; and as the home of the statistics that size and judge the result.
You can now:
- Read a protocol as what it is — scientific plan, ethical contract and regulatory submission — and check its content against E6(R3) Appendix B
- Judge a trial's endpoints — primary, secondary, surrogate — and recognise what an estimand pins down before the data exists
- Explain how randomisation, blinding, the choice of control and the choice of participants protect a trial's answer from bias — and decide when a placebo arm is ethically defensible
- Interpret p-values, confidence intervals, power and sample size well enough to respect them — and spot when a number is being misread
- Hold a running trial to its protocol: document deviations, route changes through amendments, and catch the red flags when you read one
Every safeguard you have studied — consent, review, monitoring, records, reporting — protects people from a trial. The protocol's machinery protects the answer, from everyone who will ever be tempted to improve it.