Template for Language for SAP/Protocol Incorporating Covariate Adjustment

ASA-BIOP Covariate Adjustment SWG, Standardization and Outreach Subteam

Published

July 1, 2026

1 Introduction

Covariate adjustment in randomized clinical trials is a long-established approach to improve precision and power and to provide more reliable treatment effect estimates. In practice, however, implementation and documentation have been heterogeneous across sponsors and trials, with differences in which covariates are used, how analyses are specified, and how results are interpreted. The recent FDA Guidance for Industry on “Adjusting for covariates in randomized clinical trials for drugs and biologics” (FDA 2023) and EMA Guideline on “Adjustment for baseline covariates in clinical trials” (EMA 2015) both give clear regulatory support for covariate adjustment and set expectations for pre-specification, choice of covariates, and appropriate analytic methods. They also highlight the distinction between conditional and marginal estimands and the corresponding estimators, in line with the ICH E9(R1) estimand framework (ICH E9 addendum, 2019). These concepts are central to ensuring that the estimand, analysis model, and interpretation of the treatment effect are coherent.

This document, prepared by the ASA Biopharmaceutical Section Covariate Adjustment Working Group, is intended as a cross-industry template to support consistent and transparent use of covariate adjustment in protocols and statistical analysis plans. The focus is on primary and key secondary endpoints in confirmatory individually randomized trials, with emphasis on continuous, binary, and time-to-event endpoints. For these common endpoint types, the template provides a brief contextual discussion of suitable estimands for two-arm randomized clinical trials, including both conditional estimands (for example, effects defined through regression models that condition on baseline covariates) and marginal estimands (for example, population-average effects obtained by averaging over the entire trial population). It then outlines appropriate covariate-adjusted estimation strategies for these estimands and offers concise example language that can be adapted for protocol and SAP sections.

A key aim of this template is to clarify, at a practical level, how the choice between conditional and marginal estimands relates to the planned analysis. Conventional regression models with baseline covariates naturally target conditional estimands, while approaches such as marginal standardization or weighting can be used to obtain marginal estimands when those are preferred. When estimating a conditional treatment effect through nonlinear regression, the model assumptions will generally not be exactly correct, and results can be difficult to interpret if the model is mis-specified and treatment effects substantially differ across subgroups. Interpretability increases with the quality of model specification. Sponsors are encouraged to discuss with regulators specific proposals in a protocol or statistical analysis plan containing nonlinear regression to estimate conditional treatment effects for the primary analysis. The template emphasizes that this choice should be made consciously and described explicitly, so that the estimand definition, the modeling strategy, and the reported treatment effect (for example, adjusted mean difference, odds ratio, hazard ratio, or marginal risk difference) are aligned and interpretable.

The objective is not to provide an exhaustive methodological review or to prescribe a single “correct” approach, but rather to offer concise protocol and SAP language that sponsors can reuse and adapt. By doing so, this document aims to promote principled and harmonized use of covariate adjustment, to facilitate clearer communication about estimands and analyses, and ultimately to support more efficient and transparent evaluation of treatment effects in randomized clinical trials.

In the sections that follow, we put these general principles into practical guidance. Section 2 presents the general considerations, and outlines practical approaches and example text that can be applied regardless of the endpoint type. Section 3 focuses on continuous endpoints and describes estimand definitions (conditional and marginal), statistical analysis method, sensitivity analyses and supplementary analyses, with example protocol and high level SAP language. Section 4 discusses analysis options of binary endpoints for marginal and conditional effect estimators and relevant example protocol and high level SAP language. Section 5 addresses time-to-event endpoints, discussing covariate-adjusted Cox-type analyses and the corresponding estimands and estimators.

2 General Considerations

The recently released ICH E9 (R1) by the International Council of Harmonization (ICH) (ICH E9 addendum 2019) emphasized the importance of aligning statistical analyses with the target estimand for regulatory discussion. When incorporating covariate adjustment, it is essential to distinguish between the marginal and conditional effect for appropriate estimand and estimation methodology choices.

The marginal (unconditional) treatment effect is defined by comparing the marginal outcome distributions between the treatment and control arms across all patients in the target population. On the other hand, the conditional treatment effect is defined by comparing the outcome distributions between the experimental and control arms while holding some specific baseline covariates fixed. If the conditional treatment effect is assumed to be constant across levels of these covariates, it can be represented by a single treatment effect parameter in a covariate-adjusted regression model without treatment-covariate interaction. Moreover, when the chosen measure of treatment effect is collapsible (e.g., difference in means), the conditional treatment effect, if assumed to be constant, coincides with the marginal treatment effect. When the chosen measure of treatment effect is non-collapsible (e.g., odds ratio and hazard ratio), the conditional treatment effect, even if assumed to be constant, can differ from the marginal treatment effect.

Covariate adjustment usually leads to efficiency gains when the covariates are prognostic for the outcome of interest in the trial. In some circumstances these covariates may be identified from scientific literature. In other cases, previous studies may be used to select prognostic covariates or to form prognostic indices. To ensure balance between treatment arms with respect to a selected set of prognostic covariates, randomization is often stratified by these variables. A covariate adjustment model should generally include stratification variables but can also include covariates not used for stratified randomization. Covariates used for adjustment should be prespecified, clinically justified, limited in number relative to the sample size, and restricted to baseline variables.

When covariate adjustment is implemented via a working regression model, valid inference also depends on the accuracy of the estimated standard errors. For continuous endpoints, the model-based standard error estimates can be inaccurate when the model is mis-specified or when the design departs from a simple two-arm 1:1 structure. The FDA guideline encourages sponsors to consider the use of a robust standard error method such as the Huber-White “sandwich” standard error when the model does not include treatment by covariate interactions (Rosenblum and van der Laan 2009; Lin 2013). Other robust standard error methods proposed in the literature can also cover cases with interactions (Ye et al. 2022). The asymptotically robust methods may be less applicable to very small trials conducted in rare diseases. An appropriate nonparametric bootstrap procedure can also be used (Efron and Tibshirani 1993), although its validity is only established for simple randomization rather than stratified randomization.

In addition to the primary analysis, it is important to distinguish between sensitivity analyses and supplementary analyses. Sensitivity analyses (e.g., those relying on different variance estimators or distributional assumptions) assess the robustness of conclusions for the same estimand as the primary analysis. Supplementary analyses provide complementary estimates that target different estimands (e.g., conditional vs marginal).

Sponsors that propose estimating a marginal estimand may be asked to modify the estimation strategy by, for example, omitting adjustment covariates or using an alternative pre-specified set of covariates. The estimand itself is unaffected by a change in adjustment covariates and such an alternative analysis does not interrogate deviations from the assumptions made for the primary analysis. Such an alternative analysis therefore does not satisfy the definition of sensitivity analysis. It is possible for such an alternative analysis to be non-significant and for the primary analysis to be significant (or vice versa), since adjustment covariates will affect both the estimate and its variance. However, the primary pre-specified analysis method should be given priority in assessment of the results. Requiring multiple analyses (e.g., unadjusted as well as adjusted) that target the same estimand to all be significant could substantially reduce power, thereby undermining the rationale for using covariate-adjusted analysis in the first place.

In the presence of missing values, handling of missing baseline covariate values should be prespecified and consistent with the overall missing data strategy. When multiple imputation is used, the imputation model should, where feasible, include at least the same covariates as the analysis model. Additional covariates or interaction terms could be included when needed, and inferences across imputations should be combined using Rubin’s rules, with robust variance estimators applied at the analysis stage as specified above. Single imputation may also be considered when there are missing baseline covariates, e.g., mean or median imputation for continuous covariates and mode imputation for categorical covariates (Zhao and Ding 2024; Song, Hughes and Ye 2024).

3 Continuous endpoints

For continuous endpoints, there is a general agreement among regulatory authorities about the use of analysis of covariance (ANCOVA) to adjust for baseline variables.

3.1 Estimand description

In this setting, we distinguish between conditional and marginal estimands for continuous endpoints.

A marginal estimand is the difference in population-average means between treatment groups at that timepoint. This can be viewed as the treatment effect for a “typical” patient population, rather than for a population with specific covariate values.

A conditional estimand is the mean difference in the endpoint between treatment groups at the specified timepoint, conditional on a set of prespecified baseline covariates. In practice, this usually means the expected difference in mean outcome between treatments for patients with the same value of the prespecified baseline covariates.

Protocols and SAPs should state (e.g., in the population level summary) which of these estimands is primary and if both are of interest, describe the primary and supplementary estimands. Some example protocol and SAP language is provided below.

NoteExample protocol and SAP language (marginal estimand)

The marginal (population-average) mean difference in [endpoint] at Week X between [Treatment A] and [Treatment B], defined as the difference in mean [endpoint] that would be observed if the target trial population were treated with [Treatment A] versus [Treatment B].

NoteExample protocol and SAP language (conditional estimand)

The conditional mean difference in [endpoint] at Week X between [Treatment A] and [Treatment B], at fixed values of pre-specified baseline covariates.

3.2 Statistical Analysis Methods

For continuous endpoints analyzed with a linear model (e.g., ANCOVA) including treatment and prespecified baseline covariates (with no treatment–covariate interactions), a commonly used estimator of the treatment effect is the adjusted mean difference between treatment groups at the specified timepoint. Under randomization, this adjusted mean difference targets the marginal (population-average) mean difference, a property that holds even if the linear model is misspecified. If in addition the linear model is correctly specified with no treatment–covariate interactions, it also targets the conditional mean difference (given the covariates in the model), which then coincides with the marginal effect if conditional mean difference is assumed to be constant. When there is treatment–covariate interaction, a study’s marginal treatment effect may reflect the study population’s covariate mix and may not be unbiased for the target population’s marginal effect if covariate distributions differ. This is especially important for non-continuous outcomes because interaction is scale-dependent: even with no interaction on the odds-ratio scale (constant conditional OR), there can be meaningful interaction on the absolute scale (e.g., risk differences), leading to different marginal effects across populations. As a result, marginal effects are not automatically transportable to a new population without adjustment for baseline covariate distribution (Dahabreh et al. 2019), and if interactions are expected, trials should be designed and powered to estimate treatment effects in relevant subgroups.

Under the standard linear model assumptions and 1:1 randomization, nominal model-based standard errors are generally adequate (Wang et al. 2019); however, consistent with the FDA covariate adjustment guidance, robust (e.g., Huber–White “sandwich”) standard errors are recommended when there is concern about model misspecification or when the design departs from a simple two-arm 1:1 structure. Some example protocol and SAP language is provided below.

NoteExample protocol and SAP language (either conditional or marginal estimand, under linear regression model without treatment–covariate interactions)

The primary hypothesis is that there is no difference between treatments in this mean change at Week/Month [X], more specifically: \(H_{0}:\beta_{\text{trt}} = 0\) versus \(H_{1}:\beta_{\text{trt}} \neq 0\), where \(\beta_{\text{trt}}\) is the treatment coefficient in the ANCOVA model. The primary analysis will use an analysis of covariance (ANCOVA) model including treatment group as a fixed effect and adjust for [baseline covariates]. The null hypothesis of no treatment effect will be tested using a two-sided significance level of 0.05 applied to the treatment effect in this model.

The primary estimator of the treatment effect is the adjusted mean difference in change from baseline in [Endpoint Name] at Week/Month [X] between [Treatment A] and [Treatment B], derived from the ANCOVA model. Adjusted means for each treatment group (ordinary least-squares means) and their difference will be reported, together with two-sided 95% confidence intervals and p-values. Unless otherwise specified, robust standard errors (such as robust Huber-White sandwich standard errors) will be used to construct confidence intervals and perform hypothesis tests based on normal approximation.

NoteExample protocol and SAP language (marginal estimand, under linear regression model with treatment covariate interactions)

The primary hypothesis is that there is no difference between treatments in this mean change at Week/Month [X], more specifically: \(H_{0}:\beta_{\text{trt}} = 0\) versus \(H_{1}:\beta_{\text{trt}} \neq 0\), where \(\beta_{\text{trt}}\) is the treatment coefficient in the ANCOVA model. The primary analysis will use an analysis of covariance (ANCOVA) model including treatment group as a fixed effect and adjust for [baseline covariates]. The null hypothesis of no treatment effect will be tested using a two-sided significance level of 0.05 applied to the treatment effect in this model.

Under the linear model with treatment–covariate interactions, treatment specific marginal means will be obtained by standardization: for each subject, predicted outcomes under each treatment will be generated from the fitted model and averaged over all subjects to obtain marginal mean outcomes for [Treatment A] and [Treatment B]. The difference between these marginal means will be reported, together with two-sided 95% confidence intervals based on robust standard errors or resampling methods, as well as the 95% CI for per group means.

3.3 Sensitivity analysis

For continuous endpoints with covariate adjustment, sensitivity analyses are important to assess robustness of the conclusion when varying aspects of the model or variance estimation while still targeting the same conditional or marginal mean difference specified for the primary estimand.

NoteExample protocol and SAP language (marginal estimand)

Sensitivity analyses will consist of: (i) an unadjusted analysis of change from baseline in [Endpoint Name] at Week [X] using a two sample t test and a linear model with treatment as the only predictor; (ii) alternative covariate adjusted ANCOVA models in which [specify changes, e.g., age is modeled using categories rather than as a continuous variable, or a less prognostic covariate is omitted].

NoteExample protocol and SAP language (conditional estimand)

The sensitivity analysis will re-estimate the primary model using nominal model based standard errors in addition to the primary analysis based on robust standard errors, to assess the impact of variance estimator choice.

3.4 Supplementary analysis

Supplementary analyses are additional, typically more exploratory analyses that provide complementary information beyond the primary estimand, including estimates of related but distinct estimands (e.g., conditional vs marginal) or alternative summaries of treatment effect. For continuous endpoints, supplementary analyses can help illustrate the impact of covariate adjustment (e.g., adjust for different covariates), present different effect measures, or explore treatment effects in subgroups.

4 Binary endpoints

4.1 Estimand description

Common marginal estimands for binary endpoints are the marginal risk difference or risk ratio. The conditional odds ratio is a common conditional estimand. Some example protocol and SAP language is provided below.

NoteExample protocol and SAP language (marginal estimand)

The marginal (population-average) difference in the proportion of subjects achieving [endpoint] at [Week X] between [Treatment A] and [Treatment B], defined as the difference in the proportion of subjects achieving [endpoint] if the target trial population were treated with [Treatment A] versus [Treatment B],

NoteExample protocol and SAP language (conditional estimand)

The conditional odds ratio of subjects achieving [endpoint] at [Week X] between [Treatment A] and [Treatment B], at fixed values of the stratification variables and baseline [covariates].

4.2 Statistical Analysis Methods

For binary endpoints, a common estimator of marginal treatment effects, e.g., the risk difference or risk ratio, which allows for adjustment for either continuous or categorical variables is the g-computation estimator based on a logistic regression working model. Various estimators for the standard error of the g-computation estimator have been proposed (Ye et al. 2023; Liu and Xi 2024 (for the risk difference estimator)); see Zhang et al. (2025) for an overview. The nonparametric bootstrap method for a robust standard error of the standardized estimator is an option as well (Steingrimsson et al. 2017).

Mantel-Haenszel stratum-weighted estimators are frequently used for estimating marginal treatment effects with adjustment for only stratification factors under the assumption of common risk difference or constant risk ratio across the strata defined by the covariates or stratification factors (FDA 2023; Agresti 2013). However, the assumption may not be always met in RCTs. Recent work of Qiu et al. (2025) shows that with the risk difference estimand, without this assumption, the Mantel-Haenszel risk difference estimator with the proposed modified Greenland and Robins variance estimator are able to give valid estimates and tests on the marginal risk difference.

A logistic regression model can be used to estimate an odds ratio conditional on the values of stratification variables and baseline covariates (categorical or continuous). If only categorical covariates are adjusted for, the Cochran-Mantel-Haenszel method can be used to test the existence of a treatment effect conditional on the values of the covariates.

Some example protocol and SAP language is provided below for a superiority trial.

NoteExample protocol/SAP language (marginal estimand)

The marginal risk difference/risk ratio will be estimated using the g-computation method. A logistic regression model will be fit for regressing the outcome on treatment assignments and prespecified baseline covariates, including the stratification variables. For each subject, regardless of treatment group assignment, the model-based prediction of achieving [endpoint] at [Week X] will be computed under both treatment and control, using the subject’s baseline covariates. The average response rates will then be computed under [Treatment A] and [Treatment B], by averaging these model-based predictions across all subjects in the trial. These estimates of average response will then be used to estimate the marginal risk difference/risk ratio. Standard errors and confidence intervals for the average response rates and for the marginal risk difference/risk ratio will be computed using the method of [Ye et al. 2023/Liu and Xi 2024/Zhang et al. 2025/ Steingrimsson et al. 2017].

The null hypothesis to be tested is that the marginal risk difference is 0/the marginal risk ratio is 1. This null hypothesis will be tested using a Wald test.

Adjusted response rates for each treatment group (with 95% confidence intervals) will be reported, together with the [difference/ratio] between them (with a 95% confidence interval) and the p-value from the Wald test.

NoteExample protocol/SAP language (conditional estimand)

The conditional log odds ratio comparing treatment and control and its standard error will be estimated using a logistic regression model where the outcome is whether patient achieved [endpoint] at [Week X], with the treatment indicator, the stratification variables and prespecified baseline covariates. The primary hypothesis is that there is no association between treatment and the outcome conditional on the other covariates in the model. This null hypothesis will be tested using a Wald test. The log odds ratio and its 95% confidence interval, the odds ratio and its 95% confidence interval (calculated by exponentiation), and the p-value from the Wald test will be reported.

For a non-inferiority trial, the null and alternative hypotheses must be must be specified according to the chosen effect measure and the direction of benefit. For a risk difference (RD) endpoint where larger proportions are favorable, for example, the hypotheses are \(H_0: RD\leq\Delta\) versus \(H_1: RD>\Delta\) where \(\Delta\) is the pre-specified non-inferiority margin (\(\Delta=-x\%\)). Non-inferiority will be declared if the lower bound of the two-sided 95% confidence interval for the risk difference ratio exceeds \(\Delta\).

4.3 Sensitivity Analysis

Examples of sensitivity analyses include an analysis using a different variance estimator for a marginal treatment effect estimator (Ye et al. 2023; Liu and Xi 2024; Zhang et al. 2025; Steingrimsson et al. 2017).

4.4 Supplementary Analysis

Examples of supplementary analyses include conditional treatment effect estimates (if the primary estimand is marginal, adjusted for a different set of baseline covariates) or, if the primary estimand is conditional, a marginal treatment effect estimate.

5 Time-to-event endpoints

For time-to-event endpoints, covariate adjustment is often included via stratification of the analysis for randomization stratification factors using the stratified log-rank test or the stratified Cox proportional hazards (PH) model. For some commonly used methods, e.g., the log-rank test and the Cox PH model, analytic stratification also serves a second purpose of making the analysis robust to potential differences between baseline hazard functions across the strata.

If there are additional prognostic variables beyond randomization stratification factors, covariate adjustment may be used to improve precision and power without compromising the validity of the randomized comparison, provided that covariates are prespecified and limited to baseline variables. Typical examples include a baseline laboratory measurement or a prognostic score which incorporates multiple baseline variables.

5.1 Estimand description

An example of a conditional estimand is the hazard ratio (HR) between treatments under a Cox PH model stratified by, and therefore conditional on, randomization stratification factors. This is a very common estimand in confirmatory clinical trials.

Another example of a conditional estimand is the hazard ratio (HR) between treatments under a PH model conditional on adjustment variables (and also, potentially, stratified by and therefore conditional on randomization stratification factors). This is the HR between treatments for patients with the same baseline covariate values. This estimand is not commonly used as the primary estimand in confirmatory clinical trials.

An example of a marginal (population-average) estimand is the marginal hazard ratio, defined as the treatment effect targeted by a Cox PH model that includes only the treatment indicator. Note that HR estimands may be difficult to interpret if the PH assumption is violated. Other estimands for time-to-event endpoints do not rely on a PH assumption, including the difference in restricted mean survival times between treatment groups or the difference in survival probabilities at a landmark time. This document currently focuses on HR estimands.

Some example protocol and SAP language is provided below.

NoteExample protocol and SAP language (standard estimand with stratification)

The hazard ratio for [endpoint] between [Treatment A] and [Treatment B], conditional on the values of the stratification factors.

NoteExample protocol and SAP language (conditional estimand)

The hazard ratio for [endpoint] between [Treatment A] and [Treatment B], conditional on the values of the stratification factors and the adjustment covariates.

NoteExample protocol and SAP language (marginal estimand)

The marginal (population-average) hazard ratio for [endpoint] between [Treatment A] and [Treatment B], defined as the treatment effect in a Cox PH model that includes only the treatment indicator.

5.2 Statistical Analysis Methods

If the estimand is conditional only on stratification variables, the stratified log-rank test may be used to test the null hypothesis. The conditional HR may be estimated using a stratified Cox PH model. If additional adjustment for baseline covariates beyond the stratification variables is desired, the covariate-adjusted stratified log-rank test of Ye et al. (2024) may be used to test the same null hypothesis and the corresponding covariate-adjusted estimator of the hazard ratio under the stratified Cox PH model from Ye et al. (2024) may be used for estimation.

If the estimand is conditional on adjustment variables, a Wald test of the treatment term in a stratified or unstratified Cox PH model adjusted for covariates may be used to test the null hypothesis. The conditional HR may be estimated using a Cox PH model.

If the estimand is marginal with respect to adjustment variables, the covariate-adjusted unstratified log-rank test of Ye et al. (2024) may be used to test the null hypothesis. Ye et al. (2024) also provide a covariate-adjusted estimator of the hazard ratio under the unstratified Cox PH model.

Many SAPs will contain language for collapsing sparse strata. This language is out of scope of this document.

Some example protocol and SAP language is provided below.

NoteExample protocol and SAP language (with stratification and no adjustment covariates)

The primary hypothesis is \(H_{0}:\lambda_{0}\left( t|z \right) = \lambda_{1}(t|z)\) at all times \(t\) and for all strata \(z\) where \(\lambda_{0}(t|z)\) and \(\lambda_{1}(t|z)\) are the hazards at time \(t\) in stratum \(z\) for [endpoint]. The primary analysis will test this null hypothesis with the stratified log-rank test. A two-sided significance level of 0.05 will be applied.

The conditional hazard ratio and its two-sided 95% confidence interval will be estimated using a Cox proportional hazards model with a term for treatment, stratified by the randomization stratification factors.

The hazard ratio and its 95% confidence interval and the p value from the stratified log-rank test will be reported.

NoteExample protocol and SAP language (conditional estimand with adjustment covariates)

The primary hypothesis is \(H_{0}:\beta_{trt} = 0\) where \(\beta_{trt}\) is the coefficient of the treatment term in a Cox proportional hazards model adjusted for [covariates] [and stratified by the randomization stratification factors]. A two-sided Wald test at the 0.05 significance level will be used to test \(H_{0}\).

The conditional hazard ratio and its two-sided 95% confidence interval will be estimated using a Cox proportional hazards model with a term for treatment adjusted for [covariates] [and stratified by the randomization stratification factors.]

The conditional hazard ratio and its 95% confidence interval and the p value from the Wald test will be reported.

NoteExample protocol and SAP language (marginal estimand, without stratification)

The primary hypothesis is \(H_{0}:\lambda_{0}(t) = \lambda_{1}(t)\) at all times \(t\) where \(\lambda_{0}(t)\) and \(\lambda_{1}(t)\) are the unconditional hazards at time \(t\) for [endpoint] in the control and treatment group, respectively. The primary analysis will test this null hypothesis with the covariate-adjusted log-rank test of Ye et al. (2024), adjusted for the values of [covariates]. A two-sided significance level of 0.05 will be applied.

The marginal hazard ratio and its two-sided 95% confidence interval will be estimated using the covariate-adjusted estimator of Ye et al. (2024), which estimates the hazard ratio and its standard error under the Cox proportional hazards model. This marginal hazard ratio has the same interpretation as the hazard ratio from a Cox proportional hazards model with only a term for treatment.

The marginal hazard ratio and its 95% confidence interval and the p value from the covariate-adjusted log-rank test will be reported.

NoteExample protocol and SAP language (stratified estimand, with additional covariate adjustment)

The primary hypothesis is \(H_{0}:\lambda_{0}\left( t|z \right) = \lambda_{1}(t|z)\) at all times \(t\) and for all strata \(z\) where \(\lambda_{0}(t|z)\) and \(\lambda_{1}(t|z)\) are the hazards at time \(t\) in stratum \(z\) for [endpoint] in the control and treatment group, respectively. The primary analysis will test this null hypothesis with the covariate-adjusted stratified log-rank test of Ye et al. (2024), adjusted for the values of [covariates]. A two-sided significance level of 0.05 will be applied.

The hazard ratio and its two-sided 95% confidence interval will be estimated using the covariate-adjusted estimator of Ye et al. (2024), which estimates the hazard ratio and its standard error under the stratified Cox proportional hazards model. This estimated hazard ratio is conditional on the strata, and has the same interpretation as the hazard ratio from a stratified Cox proportional hazards model with only a term for treatment.

The hazard ratio and its 95% confidence interval and the p value from the covariate-adjusted stratified log-rank test will be reported.

5.3 Sensitivity analysis

Sensitivity analyses may assess sensitivity to assumptions about censoring, including the assumption of no informative censoring.

5.4 Supplementary analysis

Potential supplementary analyses include:

  • An unadjusted analysis, with or without stratification, to illustrate the impact of covariate adjustment on the point estimate and precision.

  • If the primary analysis is stratified, an unstratified analysis. Note that standard errors should account for stratified randomization even if the analysis is not stratified (Ye et al. 2024).

  • An analysis with a Cox PH model adjusted for covariates and stratified by randomization stratification factors, if the primary analysis is covariate-adjusted but marginal with respect to the covariates.

  • Estimates of different estimands, e.g., differences of the restricted mean survival times or differences of survival probabilities at landmark time points. Such analyses are particularly useful if the PH assumption is violated.

6 References

  1. Agresti A. Categorical Data Analysis. 3rd ed. Hoboken, NJ: John Wiley & Sons. 2013.

  2. Dahabreh IJ, Robertson SE, Steingrimsson JA, Stuart EA and Hernán MA. Extending inferences from a randomized trial to a target population. European Journal of Epidemiology. 2019;34:719–722.

  3. EMA. “Guideline on adjustment for baseline covariates in clinical trials.” Committee for Medicinal Products for Human Use (CHMP), EMA/CHMP/295050/2013. European Medicines Agency. 2015.

  4. Efron B and Tibshirani RJ. An Introduction to the Bootstrap. New York: Chapman and Hall/CRC. 1993. doi:10.1201/9780429246593.

  5. FDA. “Adjusting for Covariates in Randomized Clinical Trials for Drugs and Biological Products Guidance for Industry”. 2023.

  6. ICH E9 (R1). “Addendum on Estimands and Sensitivity Analysis in Clinical Trials to the Guideline on Statistical Principles for Clinical Trials.” International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. 2019.

  7. Lin W. Agnostic notes on regression adjustments to experimental data: Reexamining Freedman’s critique. Annals of Applied Statistics. 2013;7(1):295–318. doi:10.1214/12-AOAS583.

  8. Liu J and Xi D. Covariate adjustment and estimation of difference in proportions in randomized clinical trials. Pharmaceutical Statistics. 2024;23(6):884–905. doi:10.1002/pst.2397.

  9. Qiu X, Qian Y, Yi J, Wang J, Du Y, Yi Y and Ye T. Clarifying the role of the Mantel–Haenszel risk difference estimator in randomized clinical trials. Biometrics. 2025;81(4).

  10. Rosenblum M and van der Laan MJ. Using regression models to analyze randomized trials: asymptotically valid hypothesis tests despite incorrectly specified models. Biometrics. 2009;65(3):937–945.

  11. Song Y, Hughes JP and Ye T. Adjusting for incomplete baseline covariates in randomized controlled trials: a cross-world imputation framework. Biometrics. 2024;80:ujae094.

  12. Steingrimsson JA, Hanley DF and Rosenblum M. Improving precision by adjusting for prognostic baseline variables in randomized trials with binary outcomes, without regression model assumptions. Contemporary Clinical Trials. 2017;54:18–24. doi:10.1016/j.cct.2016.12.026.

  13. Wang B, Ogburn EL and Rosenblum M. Analysis of covariance in randomized trials: more precision and valid confidence intervals, without model assumptions. Biometrics. 2019;75:1391–1400.

  14. Ye T, Yi Y and Shao J. Inference on the average treatment effect under minimization and other covariate-adaptive randomization methods. Biometrika. 2022;109(1):33–47. doi:10.1093/biomet/asab015.

  15. Ye T, Bannick M, Yi Y and Shao J. Robust variance estimation for covariate-adjusted unconditional treatment effect in randomized clinical trials with binary outcomes. Statistical Theory and Related Fields. 2023;7(2):159–163.

  16. Ye T, Shao J and Yi Y. Covariate-adjusted log-rank test: guaranteed efficiency gain and universal applicability. Biometrika. 2024;111:691–705.

  17. Zhang X, Chu H, Liu L and Roychoudhury S. A robust score test in g-computation for covariate adjustment in randomized clinical trials leveraging different variance estimators via influence functions. Statistics in Medicine. 2025;44:e70080.

  18. Zhao A and Ding P. To adjust or not to adjust? Estimating the average treatment effect in randomized experiments with missing covariates. Journal of the American Statistical Association. 2024;119:450–460.