Surrogate outcomes · ~9 min read
Surrogate Endpoints in Surgical Trials
Radiographic fusion, implant survivorship, a biomarker change: what a surrogate outcome can and cannot show about how a patient actually did.
Published
In short
A surrogate endpoint (radiographic fusion, implant survivorship, a biomarker change) is only a valid stand-in for a clinical outcome if it meets specific statistical criteria, first formalised by Prentice in 1989, and most candidate surrogates in surgery have never been tested against them.1 In hip and knee arthroplasty RCTs published 2014-2023, trials using a surrogate as the primary outcome reported favourable results 53.1% of the time versus 34.1% for trials using a true clinical outcome.4 In rotator cuff surgery RCTs, a surrogate primary outcome was associated with roughly four times the risk of a favourable result compared with a patient-centred outcome.5 And in spinal fusion, a meta-analysis of individual patient data from four RCTs found that solid radiographic fusion was statistically associated with better disability scores but had poor predictive value at the level of the individual patient.6 None of this means surrogates are useless; it means the burden of proof for using one as a primary outcome is higher than most surgical trials currently meet.
What a surrogate endpoint actually is
A surrogate endpoint is an outcome measured as a stand-in for the outcome a patient actually cares about, chosen because it appears sooner, costs less to collect, or requires a smaller sample than the true clinical outcome would. In surgical trials the pattern is familiar: a hip resurfacing trial reports implant survivorship at two years rather than waiting for revision surgery or function at ten; a spinal fusion trial reports radiographic union rather than disability scores; a cartilage-repair trial reports a biomarker change rather than whether the patient can climb stairs without pain. The true, or "target," outcome is usually defined as something that measures how a patient feels, functions, or survives, a definition used consistently across the surrogate-endpoint literature.7 The appeal is obvious: surrogates shorten trials and shrink the sample size needed to detect an effect. The risk is just as obvious once stated plainly, that the intervention moves the surrogate without moving the thing the surrogate was supposed to predict.
The formal criteria a surrogate is supposed to meet
Surrogates are not validated by intuition or face plausibility. The foundational framework comes from Ross Prentice's 1989 paper in Statistics in Medicine, which set out operational criteria for when a statistical test on a surrogate can substitute for a test on the true endpoint.1 Stated informally, treatment has to affect the surrogate, treatment has to affect the true endpoint, the surrogate has to be associated with the true endpoint, and, the criterion that almost nothing satisfies, the full effect of treatment on the true endpoint has to be captured by its effect on the surrogate. That last condition means that once you know what the treatment did to the surrogate, knowing the treatment assignment adds no further information about the true outcome. Buyse and Molenberghs extended this in 1998, arguing that two further quantities are needed alongside Prentice's criteria: the size of the treatment's effect on the true endpoint relative to its effect on the surrogate, and the strength of association between surrogate and true outcome after adjusting for treatment, precisely because the original Prentice criteria are difficult to satisfy or even test with data from a single trial.2 In practice, validating a surrogate properly usually requires meta-analytic data across multiple trials, not the single dataset any one surgical group is likely to have. Most surgical trials that use a surrogate as a primary outcome are not running that validation at all; they are using a plausible biological or radiographic marker and assuming the relationship holds.
A well-documented case where the surrogate misled
The starkest demonstration that a plausible surrogate can point in the opposite direction from the true outcome comes from cardiology rather than surgery, but it is the canonical case the surrogate-endpoint literature returns to, and it is worth knowing regardless of specialty. Frequent premature ventricular contractions (PVCs) after a myocardial infarction were a well-established risk marker for sudden cardiac death, so it was reasonable to hypothesise that suppressing PVCs with antiarrhythmic drugs would reduce mortality. The Cardiac Arrhythmia Suppression Trial tested that hypothesis directly: encainide and flecainide did suppress PVCs, the surrogate moved exactly as predicted, but patients on the active drugs died at a substantially higher rate than those on placebo, and the trial arm was stopped early because of the excess mortality.3 Nothing about the biological plausibility of the surrogate was wrong; the drugs genuinely calmed the arrhythmia the surrogate was measuring. What failed was the assumption, never actually tested until the trial ran, that suppressing the marker would suppress the outcome it was thought to predict.
Surgery has its own version of this problem, quieter because it rarely produces a single dramatic trial but shows up reliably whenever anyone checks the correlation directly. In spinal fusion surgery, radiographic union is routinely reported as evidence that a construct "worked," yet a meta-analysis pooling individual patient data from four randomised trials in the Yale University Open Data Access (YODA) database found that although a solid fusion was statistically associated with better disability and pain scores, the predictive value of fusion status for an individual patient's clinical outcome was poor, with low specificity and a low negative predictive value.6 Put differently, plenty of patients with a "successful" radiographic fusion did not get meaningfully better, and some patients with radiographic nonunion did fine. The surrogate and the true outcome move together on average across a trial population; they diverge often enough, in either direction, that neither can be read off the other for the patient in front of you.
How often surgical trials actually run on surrogates
This is not a rare design choice confined to a handful of trials. In an analysis of 350 surgical RCTs published in 2008 and 2009, a third of trials specified a surrogate outcome as their primary endpoint, and patient-important status was not associated with a study having chosen that outcome as primary at all.8 More recent and procedure-specific numbers tell the same story with sharper contrast. Across 566 hip and knee arthroplasty RCTs published in four leading orthopaedic journals from 2014 through 2023, 42.9% used a surrogate as the primary outcome, and those trials were significantly more likely to report a favourable result for the intervention (53.1% versus 34.1% for trials using a true clinical outcome).4 Trials evaluating a new surgical technology were also more likely to lean on a surrogate primary outcome, which is exactly the situation, a novel device or technique with no long track record, where the temptation to substitute an early, favourable-looking marker for the outcome that actually matters is strongest. In rotator cuff repair RCTs, the pattern was even more pronounced: trials with a surrogate primary outcome reported a favourable result in 20 of 36 cases (55.6%), versus 10 of 71 for patient-centred outcomes (14.1%), roughly a fourfold difference in risk (relative risk 3.94, 95% CI, 2.07-7.51).5 None of these associations proves that any individual surrogate-based trial overstated its intervention's benefit. What they show is a structural bias: choosing a surrogate as the primary outcome is not a neutral methodological decision, it is a choice correlated with getting a more favourable answer.
Why this happens even without anyone intending it
Surrogates tend to be more sensitive to small, mechanistically direct effects than the downstream clinical outcome is. A new fixation technique might genuinely produce a small radiographic improvement in early stability, detectable in a modest sample, without that improvement being large enough, or occurring in the right causal position, to change whether the patient still has pain or needs a second operation two years later. The surrogate is not lying; it is measuring something real. The mistake is treating "the surrogate moved" as equivalent to "the patient benefited," when the two are connected only if the surrogate meets something close to the Prentice conditions for that specific intervention and population, not just for the biological pathway in general. A biomarker validated as a surrogate for one drug class or one disease mechanism does not automatically transfer to a different intervention acting through a different pathway, a warning that shows up repeatedly across the oncology surrogate-validation literature and applies with equal force to a novel implant or fixation method that has never had its surrogate relationship tested at all.
What to do with a surrogate endpoint in your own manuscript
Using a surrogate is not itself indefensible. Sample size, follow-up duration, and the practical reality of surgical recruitment sometimes make a true clinical endpoint infeasible for a given study. What the evidence above argues against is treating a surrogate result as interchangeable with a clinical one in how it gets discussed. If radiographic union, implant survivorship, or a biomarker change is your primary outcome, say so explicitly rather than letting the framing imply a functional or patient-reported benefit you have not measured. Where the true clinical outcome was collected as a secondary measure, even underpowered, report it alongside the surrogate rather than in a supplementary table a reader has to go looking for, and describe how the two moved relative to each other rather than only reporting each result in isolation. Where possible, cite the validation evidence, or the absence of it, for the specific surrogate you used in that specific context; "radiographic fusion is a standard outcome in spine surgery" is a statement about convention, not about validity. A surrogate that has never been checked against the outcome it claims to predict is not disqualified from use. It is, at best, an unverified assumption dressed as a result.
Frequently asked questions
What makes an outcome a valid surrogate endpoint?
It has to meet the operational criteria Prentice set out in 1989: treatment affects the surrogate, treatment affects the true clinical outcome, the surrogate is associated with the true outcome, and the full effect of treatment on the true outcome is captured by its effect on the surrogate. That last condition is the one almost nothing satisfies, and most surrogates used in surgical trials have never been formally tested against it.
Does a surrogate outcome moving mean the patient benefited?
Not necessarily. The Cardiac Arrhythmia Suppression Trial showed drugs that successfully suppressed a well-established risk marker for sudden cardiac death more than doubled patient mortality. The surrogate moved exactly as predicted; the outcome it was supposed to predict moved in the opposite direction.
Are surgical trials that use a surrogate primary outcome more likely to report a favourable result?
Yes, measurably so. Hip and knee arthroplasty trials using a surrogate as the primary outcome reported favourable results 53.1% of the time versus 34.1% for trials using a true clinical outcome, and rotator cuff trials with a surrogate primary outcome had roughly four times the risk of a favourable result compared with patient-centred outcomes. Neither finding proves any individual trial overstated its benefit, but the association is structural, not incidental.
You might also read
References
- Prentice RL. Surrogate endpoints in clinical trials: definition and operational criteria. Statistics in Medicine. 1989;8(4):431-440. https://doi.org/10.1002/sim.4780080407
- Buyse M, Molenberghs G. Criteria for the validation of surrogate endpoints in randomized experiments. Biometrics. 1998;54(3):1014-1029. https://doi.org/10.2307/2533853
- Cardiac Arrhythmia Suppression Trial (CAST) Investigators. Preliminary report: effect of encainide and flecainide on mortality in a randomized trial of arrhythmia suppression after myocardial infarction. New England Journal of Medicine. 1989;321(6):406-412. https://doi.org/10.1056/NEJM198908103210629
- Barakat N, Temple JR, Carpenter LS, Novicoff WM, Berry DJ, Browne JA. Surrogate End Points Are Associated With Favorable Results in Hip and Knee Arthroplasty Randomized Controlled Trials. Journal of Arthroplasty. 2026. https://doi.org/10.1016/j.arth.2026.04.115
- Miquel J, Salomó-Domènech M, Santana F, Torrens C. Impact of surrogate outcomes in randomized controlled trials for shoulder rotator cuff tears. Archives of Orthopaedic and Trauma Surgery. 2023;143(10):6117-6122. https://doi.org/10.1007/s00402-023-04911-0
- Noshchenko A, Lindley EM, Burger EL, Cain CM, Patel VV. What is the clinical relevance of radiographic nonunion after single-level lumbar interbody arthrodesis in degenerative disc disease?: a meta-analysis of the YODA Project database. Spine. 2016;41(1):9-17. https://doi.org/10.1097/BRS.0000000000001113
- Ciani O, et al. A framework for the definition and interpretation of the use of surrogate endpoints in interventional trials. eClinicalMedicine. 2023;65:102283. https://doi.org/10.1016/j.eclinm.2023.102283
- Adie S, Harris IA, Naylor JM, Mittal R. Are outcomes reported in surgical randomized trials patient-important? A systematic review and meta-analysis. Canadian Journal of Surgery. 2017;60(2):86-93. https://doi.org/10.1503/cjs.010616
If your study reports both a surrogate measure and the clinical outcome it's meant to stand in for, whether radiographic union alongside disability scores or implant survivorship alongside reoperation rates, StatsPlease computes the effect size for each outcome separately from your data, so your Results section shows whether they actually moved together rather than assuming it.
Try StatsPlease free