Sia is a Bayesian hierarchical model. It combines a mechanistic model of tumour dynamics with a multistate survival model, and partially pools information from historical trials. For a trial with short follow-up and most patients still censored, it forecasts the results the trial will report at maturity, with credible intervals.
Interim decisions are made a few months into follow-up, while the trial is still enrolling and most patients are censored. At that point the final survival curve is not yet known.
The usual approach is to wait for more events. Enrolment continues during the wait. In some cases, the tumour measurements already collected contained enough information to predict the final result.
Waiting is expensive. A single pivotal trial supporting a new approval costs a median of about $19m to run,1 and the median research and development cost per approved drug, estimated from public filings, is about $985m.2 Enrolment and spending continue while a decision waits for more events.
The question is how much of the final result can already be estimated from the data collected so far, and with what uncertainty.
Sia produces the range of results the trial could still report, with the probability of each given the data collected so far. A go/no-go question can be answered from that range directly, as the probability that the benefit exceeds a chosen threshold, without waiting for more events.
The same output answers related questions, such as whether to stop an arm or how to size the next study, because they depend on the same unobserved results.
Borrowing from a historical trial is an assumption. The model estimates how much to borrow from the data instead of fixing it in advance. When the trials disagree, it borrows less and the intervals are wider.
Most interim forecasting methods model the endpoint directly. They estimate a hazard from the events observed so far and extrapolate the survival curve. Sia instead models each patient's tumour burden over time and derives progression and death from it.
This lets early data contribute. A patient with 4 recorded assessments adds little to an event count, but a lot to an estimate of their tumour growth.
The model is fitted to the data collected so far in the target trial, together with the historical trial, which it partially pools with.
Each patient is simulated forward from their current state: the assessments they have not had yet, and any progression or death that follows from them.
This is repeated thousands of times, giving thousands of simulated completed trials that are consistent with the data so far. Each reported quantity, such as a KM curve, a median, a response rate or an arm difference, is calculated from these simulated trials. Formally, they are draws from the posterior predictive distribution.
A statement such as “a 70% probability the benefit exceeds 2 months” means that the benefit exceeds 2 months in 70% of the simulated trials. The result is the full set of simulated outcomes. The set is wide for two reasons: patients differ from each other, and the data so far is consistent with more than one version of the model.
The two submodels are fitted jointly. The tumour trajectory provides covariates for the event model, and the event data also informs the trajectory. Progression, death and leaving the trial are states in a single model, so progression-free and overall survival come from the same simulated trials and are consistent with each other. Dropout is modelled as part of the forecast. The diagram is redrawn from the manuscript, without the mathematical notation.
The case study3 pools a historical first-line extensive-stage small-cell lung cancer trial with a target trial, and forecasts the target trial's results at three cut-offs: trial months 4, 11 and 19, counted from when the target trial opened. The target trial is a Lilly CXCR4 trial and the historical trial is Amgen 20010145. At each cut-off the model is refitted using only the data available at that date.
Forecasts of tumour burden for 6 target-trial patients who had not progressed at their last assessment, 3 per arm. Up to a patient's last assessment, marked by the vertical line, the median and 80% interval are the model's fit to the measurements in hand. After it they are a forecast, which fades as the probability that the patient is still progression-free falls. Each panel has its own vertical scale.
The thin lines inside the band are 20 individual posterior draws for the patient. Each draw is one possible trajectory, so the draws lie inside the band without filling it. The posterior sample holds 10 such patients per arm. In each arm the three patients are the one whose tumour shrank the most relative to baseline, the one whose tumour shrank the least, and, of the remaining patients, the one with the largest forecast regrowth. These picks span the range of responses and are not a random sample.
Target-trial accrual by trial month, counted from when the trial opened. The upper panel counts patients enrolled and, separately, patients with at least one tumour assessment after baseline. A patient enters the forecast only after their first assessment, so at trial month 11, 43 patients had enrolled and 39 were included. The lower panel counts tumour assessments. Dots mark the cut-offs at trial months 4, 11 and 19, when 9, 39 and 71 patients had an assessment. The counts are pooled across arms and by month.
Each cut-off is a separate fit that uses only the data available at that date. It forecasts survival in each arm for the patients enrolled by then, because a forecast can only include patients who have been observed. The forecast is compared with the survival curve the trial reported at its final data cut.
For overall survival, the 90% interval for the median contained the reported median in both arms at trial months 4 and 11 and with the full data. At trial month 19 it contained the reported median in arm A but not in arm B. The intervals were about 18 months wide at trial month 4 and 2 to 3 months wide at trial month 19.
The page uses two time scales. Cut-offs are in trial months, counted from when the target trial opened. The charts are in follow-up months, counted from each patient's enrolment. At trial month 4, no patient has more than 4 months of follow-up.
Trial month 4, 9 patients. Each arm has 4 or 5 patients, so the intervals are wide. The 90% interval for median overall survival runs from about 6 to 23 months in arm A and about 6 to 24 months in arm B, and contains the reported medians of 9.8 months in arm A and 11.6 months in arm B. The intervals for median progression-free survival also contain the reported medians.
Trial month 11, 39 patients. The intervals for median overall survival are about 5 months wide and contain the reported medians in both arms. For progression-free survival, the forecast median in arm A is 6.4 months. This is above the reported 5.6 months, though the interval now contains it. In arm B the forecast median is 5.1 months and the interval runs from 4.4 to 5.6 months, below the reported 5.8 months.
Trial month 19, 71 patients. The intervals for median overall survival are 3.0 months wide in arm A and 2.4 months wide in arm B. The arm A interval contains the reported median. The arm B interval runs from 8.6 to 11.0 months and does not contain the reported 11.6 months. The forecast median progression-free survival in arm A is 6.1 months, still above the reported 5.6 months. In arm B it is 5.6 months, and the upper end of the interval, 5.8 months, is at the reported median.
Full data, 77 patients. With all follow-up, the model matches the reported median overall survival in both arms and the reported median progression-free survival in arm B. It estimates median progression-free survival in arm A at 5.9 months, against 5.6 months reported. This difference remains with complete data, so it is at least partly due to how the model fits arm A.
The interactive charts could not load.
Overall and progression-free survival for each arm of the target trial. The band is an 80% interval at each follow-up month, so the reported curve is expected to fall outside it at some time points.
Median overall and progression-free survival forecast at each cut-off. These are posterior summaries of the median itself, so they can differ from where the curves above cross 50%.
At trial month 11, 43 patients had enrolled, and the intervals for median overall survival contained the medians reported at the final data cut. Enrolment continued until about trial month 19, and the last tumour assessment in the data is at about trial month 30.
The intervals for median overall survival narrowed from about 18 months wide at trial month 4 to 2 to 3 months at trial month 19, and contained the reported medians in both arms except arm B at trial month 19. The model overestimates median progression-free survival in arm A from trial month 11 onward, and still does with the full data, so the difference is at least partly due to how the model fits arm A. The interval for median progression-free survival in arm B is below the reported median at trial month 11 and reaches it at trial month 19.
Whether an interval of a given width supports a decision depends on the decision. A forecast supports a go/no-go decision only against a benefit threshold and a cost of being wrong, and the trial sponsor sets both. For a given trial, the relevant question is the earliest cut-off at which the interval is narrow enough for that threshold. These results come from one disease and two trials and should be checked on a sponsor's own data before they are applied elsewhere.
A first conversation works best when it is about a specific trial: what is enrolled, what is measured, when the decision is due, and what would change if the answer came earlier. No data needs to be shared for that conversation.
Manuscript — the ES-SCLC forecasting study, with the full leave-future-out results and calibration checks. arxiv.org/abs/2607.17908.
Sia — the model core: hazards, GP knot grids, tumour dynamics, the multistate likelihood and leave-future-out scoring. The source is private, licensed under PolyForm Noncommercial 1.0.0. Access and commercial terms are available on request.
Karim Naguib runs zahrcast. At AstraZeneca he worked on patient-level models of tumour dynamics used to support Phase 3 initiation decisions. He is a coauthor of the ES-SCLC study above. Before pharma he worked in development economics, running field experiments in Kenya and South Asia.
karimn.co has the full background and publication record.