Abstract
Advances in oncology drug development are driving the emergence of novel therapies, challenging traditional dose-efficacy assumptions in dose-finding oncology trials. Traditional trial designs aim to identify a maximum tolerated dose (MTD) by assessing patients’ dose-limiting toxicities (DLTs) – adopting traditional dose-efficacy paradigms that efficacy increases with treatment dose. With these new therapies in mind, emphasis should shift toward methodological advancements in trial designs aimed at identifying optimal doses, rather than solely determining MTDs. Incorporating patient-reported outcomes (PROs) within dose-finding oncology trials is increasingly recommended to better understand treatments’ tolerability profiles, especially given the extended tolerability assessment windows for novel immunotherapies and targeted therapies. This article introduces PRO-ADD (
Introduction
Phase I dose-finding oncology trials typically look to assess the safety of new clinical treatments emerging from drug development across a range of doses. Traditional cytotoxic treatments often demonstrate a positive correlation between dosage, toxicity, and efficacy – with a higher dose exhibiting greater activity and an increased probability of toxicity. Consequently, traditional dose-finding oncology trial designs look to identify a maximum tolerated dose (MTD) by estimating the probability a patient experiences a dose-limiting toxicity (DLT) during the trial. Trial designs often look to recommend the largest dose with a DLT probability closest to the target level – recommending the (assumed) most efficacious dose whilst safeguarding patient’s safety.
Advances in oncology drug development are driving the emergence of novel therapies that challenge these traditional assumptions. 1 Research suggests the implicit monotonicity assumption embedded in MTD determination may be violated for these new treatments, particularly in the case of immunotherapies and molecularly targeted agents (MTAs). Unlike cytotoxic treatments, these novel agents may exhibit a plateau in efficacy beyond a certain dose, so increasing the dose does not necessarily increase efficacy. Recent recommendations therefore suggest that dose-finding oncology trials should focus on determining minimum effective doses or optimal biologically active dose rather than MTDs for these therapies. 2
Whilst MTDs are almost always identified for cytotoxic treatments in phase I trials, more than one third of phase I trials investigating MTAs observe few DLTs and fail to establish the MTD. 3 The rapid development of immune-oncology agents emerging from drug development has motivated the rise of seamless Phase I/II trial designs to assess both preliminary efficacy and toxicity endpoints within early phase trials. 4 Whilst such designs look to satisfy this specific dose-efficacy relationship for novel therapies, the majority of these trial designs continue to assess toxicity solely using clinician-assessed DLTs.5,6
For cytotoxic agents, treatment is often administered over a relatively short period of time – traditionally guided by DLTs that occur in the first cycle of treatment (usually 28 days). Such assessment windows are deemed appropriate for the investigation of cytotoxic chemotherapies as DLTs are often observable soon after treatment commencement.7,8 Conversely, MTAs and immunotherapies are administered for a longer period of time, often until disease progression or resistance is observed. 9 These treatments may induce low-grade toxicities 1 which, though not definable as a DLT, become intolerable to patients over prolonged dose administration periods. Immune-related adverse events associated with immunotherapy regimens are often observed beyond the traditional DLT assessment window, 2 and research has suggested that for MTAs, approximately 57% of grade 3 or 4 toxicities associated with treatment administration occur beyond the first cycle of treatment. 9 For example, Durvalumab, an approved immunotherapy treatment, and targeted cancer drug Ibrutinib have been associated with early discontinuation of treatment due to treatment adverse events.10,11 Trial designs with a time-to-event component have been recommended to account for such late-onset toxicities. 2
The utilisation of patient-reported outcomes (PROs) is being increasingly endorsed for use within dose-finding trial designs to refine our understanding of a treatment’s tolerability profile.12–15 As defined by the FDA, a PRO is a report of the status of a patient’s health condition coming directly from the patient, without interpretation of the patient’s response by a clinician or anyone else. 16 Standardised measures (PROMs), such as the PRO-CTCAE, allow patients to evaluate the tolerability of up to 78 symptomatic adverse events to assess the safety and tolerability of a novel therapy from a patient’s perspective. 17 The inclusion of PROs within dose-finding oncology trials has been advocated to support the assessment of subjective toxicities such as fatigue, pain and anxiety,13,18,19 which may lead to a conflicting assessment of treatment tolerability between patient and clinician.20,21
The use of PROs in dose-finding oncology trials has significantly increased over recent years, but is still emerging, with only 5.3% of dose-finding oncology trials including PROs as an endpoint. 22 A review of published dose-finding trials incorporating PROs found that PROs informed dose recommendations in just 11.4% of eligible trials (4/35). 23 Whilst it is recognised that PRO data in early phase trials may influence subsequent clinical development, 24 existing trial designs focus solely on binary patient-informed DLT endpoints observed within a DLT assessment window.25–27
The FDA are encouraging trialists to consider the inclusion of PROs for the assessment of tolerability in dose-finding and subsequent trials. 15 The integration of electronic systems for ePRO 28 collection facilitates tolerability assessment beyond the DLT assessment window, reducing the patient and site burden associated with paper-based data collection and handling. 29 Such advancements, alongside encouragement from regulators to assess PROs regularly within cancer clinical trials, 30 set the stage for longitudinal analysis of PRO data across multiple time points. Such analysis is increasingly motivated for immunotherapies and MTAs administered over prolonged periods, 24 where adverse events may be less clinically severe but increasingly bothersome to the patient over treatment duration. 31
Existing dose-finding trial designs do not incorporate longitudinal patient-reported outcome (PRO) assessments into dose-decision making. This represents a critical limitation of current approaches, with the longitudinal modelling of PROs more sensitively capturing patient-experienced toxicities over time compared to existing approaches. In this article, we introduce the first dose-finding trial design to integrate longitudinal PRO endpoints alongside preliminary efficacy and clinician-assessed toxicity. In doing so, we strengthen PRO-integrated trial designs’ ability to robustly evaluate tolerability and further align PRO-integrated trial designs with patient priorities for tolerability and therapeutic benefit. 32
Specifically, we introduce PRO-ADD (PRO-
In Section 2 we introduce the PRO-ADD modular dose-optimisation trial design. Section 3 details the simulation study set-up. Results from the simulation studies are detailed in Section 4 and discussed in Section 5.
Methods
Modular trial design framework
PRO-ADD is a modular trial design framework that separates the trial into three key modules: initial dose-escalation, dose-optimisation, and final dose selection. The components of each module can be customised, allowing flexibility in selecting the dose-escalation method, endpoints, and interim analyses to tailor the design to the goals of a given trial. As illustrated in Figure 1, each component functions as an individual jigsaw piece, which can be assembled to build a tailored trial design.

A figurative representation of the PRO-ADD modular trial design.
The PRO-ADD framework is intentionally conceptually simple, with modularity allowing additional elements where appropriate and feasible, supporting implementation across dose-finding trials of varying size and complexity whilst maintaining robustness and interpretability. This modular trial design firstly focuses upon identifying the MTD, before optimising dose with respect to additional key endpoints, and finally recommending a dose(s) for further evaluation. In the subsequent sections, we describe the endpoints and estimation approaches used in our illustrative implementation of PRO-ADD. Section 2.6 then details the trial design, randomisation procedure and final dose-decision criterion utilised for this example. Rather than prescribing a specific design, this example illustrates one possible implementation of PRO-ADD. Simpler alternatives may also be used if deemed appropriate.
Several binary, ordinal and continuous endpoints have been proposed to capture and summarise PRO data.25,33
Patient dose-limiting toxicity
The patient dose-limiting toxicity (P-DLT), first proposed by Lee et al. 25 is a binary endpoint similar to the clinician assessed DLT typically used to assess toxicities within dose-finding oncology trials. Presently, PROs have only been incorporated within dose-finding trial designs as a binary endpoint, with current work focussing solely upon the determination of a MTD.25–2734
The limitations of condensing the multidimensional nature of tolerability assessment into a binary DLT have motivated the development of novel endpoints aimed at more comprehensively accounting for the overall severity and clinical relevance of multiple toxicities within tolerability assessments. 35 Limitations, such as overlooking moderate toxicities which fall short of the DLT definition, and the number of DLTs which occur, 35 impact both clinician-assessed DLT and P-DLT endpoints.
Normalised PRO-adverse event (PRO-nAE) burden score
A number of continuous burden scores have been proposed to summarise adverse events utilising CTCAE33,35–39 and many can be readily adapted for PRO measures including the PRO-CTCAE. A more in-depth discussion of an exemplar summary measure, the toxicity index score, along with its applicability to dose-finding trials, is provided in Section 1.1 of the Supplementary Materials.
In this article, we present and utilise the normalised PRO-Adverse event (PRO-nAE) burden score, adapted from the Adverse Event (AE) burden score of Le-Rademacher et al.
38
to provide a single quantitative summary measure that captures both the frequency and severity of toxicities experienced by a patient. Although originally developed for clinician-assessed adverse events, we have adapted this approach for patient-assessed toxicities using instruments such as the PRO-CTCAE. We refer to this endpoint as the PRO-nAE burden score from henceforth.
For severity grades
A discussion linking the AE burden score and Total Toxicity Profile score can be found in Section 1.2 of the Supplementary Materials.
In this article, we illustrate the use of PRO-nAE burden score with the PRO-CTCAE severity attribute, which mirrors the CTCAE, with values 0–4 corresponding to an observed toxicity with grade
In Figure 2 we present an exemplar PRO questionnaire completed by a patient. Supposing a patient is asked about three adverse events and reports two moderate toxicities and one severe toxicity, their PRO-nAE burden score is equal to

Exemplar short PRO questionnaire filled in by a patient.
One may view the PRO-nAE burden score as the normalised average severity of side effects experienced by a patient during the trial. We can translate the normalised PRO-nAE burden score to the average mean severity grade experience by a patient as per Table 1. For example, a PRO-nAE score of 0.25 indicates that the patient is experiencing a level of severity equivalent to mild across all assessed side effects. If a patient has a PRO-nAE burden score of 0.60, this indicates the patient is experiencing a level of severity equivalent to moderate to severe side effects across all assessed side effects. For trialists interested in the exact mean distribution of side-effect severity associated with a given PRO-nAE burden score, please refer to Section 2 of the supplementary materials which provides examples for PRO-nAE burden scores of 0.06, 0.18, 0.24, 0.39, 0.52, and 0.59. Other PRO-nAE scores can be translated in the same manner using the code provided in this article.
Relationship between average (mean) side effect severity experienced by a patient and PRO-nAE burden score.
Relationship between average (mean) side effect severity experienced by a patient and PRO-nAE burden score.
We model PRO-nAE burden score longitudinally across multiple treatment cycles or timepoints. A generalised linear mixed effect model is employed, defined by a set of covariates
To predict the PRO-nAE burden score, we use a Bayesian mixed-effect beta regression model
42
with a logit link function and beta likelihood. This approach ensures that model predictions remain bounded in [0,1], aligning with the properties of the PRO-nAE burden score.
42
In this framework, we apply uninformative priors to the model parameters to allow the data to predominantly inform the posterior distribution. The model assumes,
Bounded PRO-nAE responses
The uninformative priors placed on each covariate are,
Trialists may also wish to consider utilising weakly informative priors to incorporate modest prior knowledge.
Aligning with the OPTIMISE-ROR guidance,
43
which identifies key PRO research objectives for early phase dose-finding oncology trials, PROs are analysed within PRO-ADD to inform final dose-selection decision-making. We utilise this model to estimate
Prior research has demonstrated that PRO score can vary non-linearly over the course of treatment. 44 To account for this, the inclusion of the quadratic time covariate allows the model to capture plateauing of PRO-nAE burden score. Methodological standards for PRO-CTCAE usage in oncology trials indicates that most adverse events occur during the initial cycles of treatment, with symptom severity typically plateauing during longer term treatment administration as symptomatology stabilises. 45 Prospective longitudinal studies of patient-reported symptoms within oncology trials similarly identified plateauing symptom severity toward the later stages of they study period. 46 As with any model, it is important to evaluate the trade-off between a model’s predictive performance and its parsimony to ensure computational stability. We consider this model particularly suitable for larger, dose-optimisation trials where reliable estimation of longitudinal symptom trajectories can support informed dose selection.
Many dose-finding trial designs which integrate preliminary efficacy into dose recommendations estimate the probability of a binary response outcome, often defined by conventional criteria such as Response Evaluation Criteria in Solid Tumors version 1.1 (RECIST v1.1).
47
A common approach employs a simple beta-binomial conjugate model to estimate the probability of response
Research indicates that more dosage is not necessarily correlated with greater treatment efficacy.
48
Brock et al’s review of dose-finding oncology trials found that whilst some cancer treatments (particularly cytotoxic treatments) do exhibit monotonic dose-response relationships, plateauing efficacy beyond critical doses or consistent efficacy response across all doses are possible dose-response relationships within cancer trials.
48
The FDA’s Project Optimus initiative has recognised that for targeted drugs, higher doses beyond a critical dose may not increase efficacy response.
49
Whilst immunotherapies such as Ipilimumab have shown increasing dose-response relationships,
50
other immunotherepies such as anti-PD-1/PD-L1 therapies have alternatively exhibited flat dose-response relationships.
51
Similar paradigms have been observed for CAR-T cell therapies.
52
Under these schemas, we utilise iPIPE
53
(inverse product-of-independent-probability-escalation) to estimate the probability of response for each dose
PIPE was first presented to identify MTDs within a dose-escalation trial design.
54
Cheung et al.
53
extends this methodology such that, rather than classifying amongst doses to identify the MTD, the inverse PIPE-classifier (iPIPE) estimates response rates
For data
To estimate the probability of binary response, we utilise the beta-binomial conjugate model with
iPIPE gives us access to the inverse cumulative distribution function for the posterior distribution of
As we can evaluate the inverse cumulative distribution function for the posterior distribution of
Dose-optimisation using a modular trial design
Illustrative example of a PRO-ADD design
Figure 3 illustrates an example of a PRO-ADD modular trial design. The design initially identifies the MTD based on clinician- reported DLTs encompassing unacceptable toxicities and treatment related death, followed by randomisation among admissible doses. Early stopping rules for futility or excessive toxicity are incorporated. Final dose selection is guided by evaluating the trade-off between patient-reported tolerability and preliminary efficacy outcomes. Our proposed trial utilises a Bayesian Optimal Interval (BOIN)
55
backbone to identify the dose with a probability of DLT closest to an elicited target DLT rate

An illustration of a PRO-ADD trial design.
Within PRO-ADD, repeated measure modelling is utilised to inform estimated PRO-nAE burden score as per Section 2.3 and iPIPE is employed to estimate the probability of a binary response endpoint as per Section 2.4.
Stage 1 follows exactly the dose-escalation routine of BOIN,
55
and is set forth as follows. Dose is escalated by evaluating the empirical DLT rate (
At each dose allocation, the safety stopping rule is applied to remove doses with an unacceptably high probability of DLT.
58
For
Patients are escalated toward the MTD as per Stage 1 of the trial design until the same dose is recommended consecutively for a pre-specified number of cohorts or the first futility interim analysis occurs at the enrolment of a specific cohort.
At Stage 2, we introduce an additional futility stopping rule to remove dose(s) which have shown inadequate activity based on the complete response data collected up to that point. The admissible dose set
For a null response rate of 10% (deemed futile) and an alternative response rate of 25% (deemed effective), the futility stopping rule protects Type I error rate at 20% and power at 74%. 59 For trialists wishing to consider other criterion, PRO-ADD provides trialists flexibility to consider other priors aligning with their own criteria for futility.
Note that a patient with partial response data are excluded from futility stopping rule decision making. Thus a dose stopped for futility may be reinstated later if the partial data, once fully collected, suggests that the dose may in fact be efficacious. To ensure sufficient sample size for the assessment of futility, the stopping rule is evaluated only twice within the trial – once we have at least six cohorts of complete response data and only assessing doses with at least six patients on treatment. This approach allows for a simpler futility rule which avoids the need for methods that incorporate partial data into the futility decision making. Futility stopping rules incorporating partial data are particularly valuable when the futility rule is evaluated at each patient enrolment. For trialists who wish to utilise partial response data, a review of methods is provided by Zhou et al. 60
Within Stage 2, patients are allocated to doses using the following criteria: For Evaluate Repeat accordingly until the pre-specified sample size has been reached.
The trial is terminated once a maximum trial size is reached or if no doses are deemed admissible. The final admissible set contains all admissible doses at most as large as the MTD as identified by the BOIN design using isotonic regression. 55
Final dose recommendation
At this point, the longitudinal PRO data are utilised within this PRO-ADD design. The final recommended dose is assessed by quantifying the trade-off between dose response and PRO tolerability data using a loss function amongst admissible doses
In PRO-ADD we investigate
Due to the small sample sizes of early phase dose-finding trials, it is to be expected that estimates of response and PRO-nAE burden scores are subject to significant uncertainty. Bayesian estimation valuably provides posterior distributions for parameters, capturing the uncertainty of response and PRO-nAE burden score posterior estimates. We recommend that we take advantage of this property by evaluating the expected (empirical) loss for each dose
We estimate PRO-nAE burden score at the final assessment timepoint. Given the posterior distribution for the PRO-nAE burden score
Figure 4 shows potential samples

1000 sampled points and densities of the posterior distribution of iPIPE efficacy estimates and estimated PRO-nAE burden score for 5 doses under Scenario 5 of the simulation study, with true preliminary efficacy rates and PRO-nAE scores marked with a star. The unacceptable loss region (with a loss greater than 0.9) is coloured in dark blue.
Using these samples, the empirical loss for each dose
The optimal dose recommendation
We recommend dose
This decision rule prioritises the three endpoints by first identifying admissible doses using DLT data only, before choosing the optimal dose using preliminary efficacy and patient-reported tolerability endpoints. This ensures any recommended dose is suitable from a safety perspective first, before it is then evaluated in terms of its preliminary efficacy and tolerability.
Clinically,
Once a loss function has been chosen, the choice of
Acceptable dose(s)
Trialists may also wish to identify a set of acceptable doses. The set of acceptable doses
In the subsequent simulation study, we set an acceptable dose as one with a loss of no more than 0.90 (equivalent to a 10% response rate and 0 PRO-nAE burden score). This loss region is highlighted in Figure 4.
Simulation study
Case study
Our simulation study is motivated by KEYNOTE-001, a first-in-human Phase I dose-escalation study of Pembrolizumab in patients with advanced solid tumours investigating toxicity and activity across three treatment doses (ClinicalTrials.gov identifier: NCT01295827). 62
Whilst KEYNOTE-001 was initially designed as a dose-finding trial using the
Subsequently published case reports on immune-related adverse events associated with Pembrolizumab indicated that adverse events including severe colitis, pneumonitis, and organ damage appeared on average of 5 to 15 weeks after commencement of treatment. 64 This reiterates the motivation for the assessment of adverse events beyond clinician DLT assessment windows. Follow up findings published three years after the KEYNOTE-001 study indicated that there was no association between the dose of pembrolizumab (2 mg/kg or 10 mg/kg every 3 weeks or 10 mg/kg every 2 weeks) and the activity nor toxicity of the immunotherapy. 51
Data generation
Binary endpoints
Binary clinician-assessed DLT observation and response outcomes are sampled from a Bernoulli distribution.
PRO-nAE burden score
The synthesis of PRO-CTCAE data has thus far has been limited to the simulation of a binary Patient-DLT. As such, within this article we take particular care with the simulation of the continuous PRO-nAE burden score.
The normalised adverse event score is bounded on [0,1], naturally lending itself to simulation via Beta sampling. However, to assess whether the distributional properties of the AE burden score abides by such a data-generating scheme, AE scores were first generated for each patient toxicity-by-toxicity using a non-parametric data generating schema. The synthesis of PRO-nAE data is motivated by a dataset of 219 patient’s PRO-CTCAE scores with advanced solid tumours who were enrolled on Phase I clinical trials at Princess Margaret Cancer Centre in Canada from 1 May 2017 to 1 January 2019. 65
The synthesised PRO-CTCAE data was generated under the following assumptions: Dose-dependent toxicity: For a toxicity type Time-dependent toxicity: For a toxicity type Correlation between toxicities: For each patient
A patient’s PRO-nAE burden score was synthesised non-parametrically using latent Beta random variables. Data generation therefore relied on a
Following this investigation, we conclude that the simulated PRO-nAE burden scores can be appropriately sampled using a Beta distribution. Thus, for ease of computation, PRO-nAE scores are sampled from a Beta distribution, with shape and rate parameters determined via the maximum likelihood estimation of the non-parametric data generation using package ‘
Correlation between binary clinician-DLT and the repeated continuous PRO-nAE scores for each patient is induced using a Gaussian copula. Figure 5 demonstrates how DLT is correlated with PRO-nAE burden score within this simulation study. 67 Whilst evidence indicates that DLT occurrence is more likely as dose increases, research suggests the probability of response may not increase similarly. 48 As such, in the main manuscript we suppose DLT and efficacy response are independent. We subsequently investigate PRO-ADD performance when DLT and efficacy responses are mildly correlated in sensitivity analyses.

Simulated PRO-nAE burden scores for Scenario 1 with observed DLT and no DLT across 5 doses.
In the following simulation study, sample size is fixed at 60 patients, enrolled in cohorts of 3. We evaluate the first two tumour responses at week 8 and week 16. Patients are sequentially enrolled every 4 weeks. Whilst PROs were not collected in the original KEYNOTE-001 study, we schedule PRO collection every two weeks to align with current literature which recommends regular (weekly or fortnightly) PRO data collection over initial cycles of treatment to accurately capture the high number of adverse events likely to occur at the start of a trial. 45 Figure 6 highlights this schedule.

Scheduling timetable for the simulation study, inspired by the KEYNOTE-001 dose-finding trial. 62
Patients are escalated toward the MTD as per Stage 1 of the trial design until the same dose is recommended consecutively for two cohorts or the first futility interim analysis takes place at the enrolment of the eleventh cohort. A futility stopping rule is used to remove doses with insufficient response rate at the end of the tenth cohort and sixteenth cohort (using the complete response data for cohorts 1–6 and cohorts 1–12, respectively).
The target DLT rate deemed admissible is 0.25. Dose (de-)escalation boundaries
Eight relationships between PRO-nAE burden score and preliminary efficacy are investigated in this simulation study. These scenarios are labelled in Table 2. Of particular note, Scenario 7 explores a unimodal dose-efficacy curve and Scenario 8 explores a setting where two doses have an equivalent loss (both with a loss of 0.75). iPIPE is only an appropriate analysis approach to analyse preliminary efficacy if the dose-efficacy relationship is monotonic or plateauing, however we present Scenario 7 as a sensitivity analysis to illustrate PRO-ADD’s performance in the unlikely event that a treatment presumed to have a monotonic or plateauing dose-efficacy relationship instead exhibits a unimodal relationship. Should trialists suspect a treatment has a unimodal dose-efficacy relationship, an alternate analysis approach (such as a beta-binomial model) should be alternatively used.
PRO-nAE burden score and efficacy simulation scenarios investigated within this simulation study.
PRO-nAE burden score and efficacy simulation scenarios investigated within this simulation study.
Two potential tolerability scenarios for PRO-nAE burden score are investigated. We consider the monotonic increasing of PRO-nAE burden score over treatment administration, which may reflect the growing symptom burden and accumulating moderate toxicities occurring over extended treatment windows (Scenarios 1–4 and 7–8). Secondly, we consider a scenario where a patient’s symptomatology stabilises over treatment administration. This is characterised by a plateau in observed PRO-nAE burden score over treatment cycles (Scenarios 5–6). 45
An example of the time trends to be investigated in this simulation study are presented in Figure 7. Section 4 of the Supplementary materials graphically displays simulation scenarios 1–6.

The PRO-nAE burden score time trends investigated in the simulation study. (a) Monotonically increasing time trend and (b) Plateauing time trend.
These simulation scenarios were investigated under two MTD scenarios, the first where dose 5 is the MTD and another where dose 3 is the MTD. The DLT rates for both scenarios are shown below, with MTD indicated in bold: DLT Scenario M5: 0.01, 0.05, 0.10, 0.15, DLT Scenario M3: 0.06, 0.13,
In the subsequent simulation study, we compare PRO-ADD to U-BOIN. 60 Similarly to PRO-ADD, U-BOIN is a two stage design – firstly identifying the MTD using DLT data alone before incorporating an efficacy endpoint into decision-making to identify the optimal biological dose. What’s more, like PRO-ADD, U-BOIN utilises a trade-off framework. Importantly, in U-BOIN’s case, decision making is guided by evaluating the utility of each dose based on its DLT and response rate, without incorporating patient-reported tolerability measures. Whilst U-BOIN can take ordinal toxicity and response endpoints, we compare PRO-ADD to the design with binary DLT and efficacy responses. U-BOIN simulation performance is assessed using the online web app www.trialdesign.org. The exact inputs used to run the U-BOIN simulation study are provided in Section 5 of the Supplementary materials. In line with current practice, to ensure fair comparison between U-BOIN and PRO-ADD, in the main manuscript we evaluate PRO-ADD performance supposing complete data collection. As a sensitivity analysis, we evaluate the performance of PRO-ADD in light of intercurrent events.
Investigated sensitivity analyses
We explore many sensitivity analyses for PRO-ADD in Section 6 and 7 of the Supplementary Materials. Investigations include, Performance of PRO-ADD when dose 1 is the MTD or no doses are safe. Performance of PRO-ADD when DLT and efficacy response are mildly correlated on a patient level. In this instance, patient’s responses are correlated using a Gaussian copula with covariance 0.15. Performance of PRO-ADD with a different acceptable loss to results presented in the main manuscript, with Performance of PRO-ADD in light of intercurrent events including dose discontinuation due to DLT and death unrelated to treatment by utilising a hypothetical handling strategy.
69
Performance is evaluated both when patient DLT and activity responses are independent of one another or mildly correlated.
Results
For PRO-ADD, we present the probability of selecting the single optimal dose (i.e. the dose with the smallest loss amongst admissible doses) and the probability of identifying an acceptable dose based on
Table 3 shows PRO-ADD and U-BOIN performance across eight scenarios when the MTD is dose 5 and all doses are admissible for safety. The probability of correct selection of the optimal dose for PRO-ADD ranges from 43% to 77%. The optimal dose is also recognised as an acceptable dose (with a loss of at most 0.9) 71%–96% of the time. PRO-ADD performs better or approximately as well as U-BOIN in all scenarios. PRO-ADD most significantly improves upon U-BOIN performance (by at least 18%) in Scenarios 2, 4, 6, 7, and 8 where efficacy plateaus. In these scenarios, the inclusion of PRO-nAE burden score successfully constrains the decision criterion to choose the most effective dose with smallest tolerability burden. In these scenarios, U-BOIN often identifies higher, more intolerable doses with no improved efficacy. In Scenario 3 where doses 3–5 have equal PRO-nAE burden score, utilisation of iPIPE estimation of efficacy ensures PRO-ADD correctly identifies dose 5 as the best dose 13% more often than U-BOIN, which relies on a beta-binomial model to estimate response rate.
Proportion of optimal dose recommendations (bold) for each dose level under 8 scenarios for 5,000 simulated trials dose 5 is the MTD (Scenario M5), with probability no dose selected also indicated for PRO-ADD and U-BOIN.
Proportion of optimal dose recommendations (bold) for each dose level under 8 scenarios for 5,000 simulated trials dose 5 is the MTD (Scenario M5), with probability no dose selected also indicated for PRO-ADD and U-BOIN.
Note: Acceptable doses are highlighted in green and have a maximum loss of 0.90. Individual patient DLT and activity responses are simulated to be independent and full patient data are collected.
Scenario 1 highlights the strength of PRO-ADD. Whilst all doses are considered safe, patient tolerability to treatment decreases as dosage increases. However, the treatment’s efficacy begins to level off at dose 4, with a marginal improvement in the probability of response between doses 4 and 5 of 2%. Traditional trial designs that target the MTD would typically recommend dose 5 as the recommended Phase II dose, with U-BOIN identifying dose 5 as best 36% of the time. However, with the addition of the PRO-nAE endpoint, the PRO-ADD design effectively determines that dose 4 is optimal using our pre-specified
What’s more, in Scenario 2 PRO-ADD reliably identifies that dose 3 is the optimal dose approximately 74% of the time. As such, PRO-ADD successfully concludes that doses 4 and 5 have an increased tolerability burden but do not offer patient’s any additional efficacy benefit. PRO-ADD does well estimating the PRO-nAE burden score when burden score monotonically increases across cycles (Scenarios 1–4) and when PRO-nAE burden score plateaus beyond a certain cycle of treatment (Scenarios 5–8).
Whilst iPIPE is not a suitable analysis method for unimodal dose-efficacy relationships, PRO-ADD can still perform well in this instance as per Scenario 7. In this case, though iPIPE would estimate the efficacy rate for dose 5 to be at least that of dose 4, the increased PRO-nAE burden score associated with dose 5 ensures that this dose is not recognised as optimal. Whilst U-BOIN utilises a beta-binomial model to estimate preliminary efficacy rate, PRO-ADD still performs 25% better than U-BOIN.
In Scenario 8, we regard doses 1 and 2 as equally optimal. PRO-ADD selects doses 1 and 2 84% of the time, an improvement of 25% over U-BOIN.
Table 4 demonstrates PRO-ADD and U-BOIN performance when dose 3 is the MTD across scenarios 1–8. Dose 3 is recognised as the optimal dose between 44%–69% of the time. Across all scenarios, at most 9% of trials recommend a dose higher than the true MTD and optimal dose – showcasing the designs effective safety overdosing control. PRO-ADD once again performs equally well or better than U-BOIN for each scenario.
Proportion of optimal dose recommendations (bold) for each dose level under 8 scenarios for 5,000 simulated trials dose 3 is the MTD (Scenario M3), with probability no dose selected also indicated for PRO-ADD and U-BOIN.
Note: Acceptable doses are highlighted in green and have a maximum loss of 0.90. Inadmissible doses with a DLT rate above the target are highlighted in red. Individual patient DLT and activity responses are simulated to be independent and full patient data are collected.
A trial can fail to identify an optimal dose for two reasons. Firstly, the trial may end prematurely as futility and safety stopping rules are activated and no doses are deemed safe or efficacious. Alternatively, the trial may reach completion without identifying any acceptable doses (i.e., no dose has an estimated loss below 0.9). In Tables S3 and S4 of the Supplementary materials we breakdown the reason for failure to identify an optimal dose into these two factors for PRO-ADD. A larger proportion of trials are halted when dose 3 is the MTD, driven by an increased likelihood of stopping for safety. What’s more, under the M3 scenario, it is more difficult to identify that the loss of dose 3 is below 0.9. The optimal dose’s loss under M3 ranges from 0.70–0.86 rather than 0.68–0.83 under scenario M5.
The safety and futility rules remove inadmissible doses throughout the trial. For PRO-ADD, the mean number of patients allocated to each dose for MTD scenario M5 is (11.16, 11.47, 13.92, 13.26, 12.32). The mean number of patients allocated to each dose for MTD scenario M3 is (15.40, 16.11, 18.64, 11.70, 8.63). Table S5 in the Supplementary materials details PRO-ADD patient allocation for each dose and each simulation scenario. For both MTD scenarios, the safety and futility stopping rules ensure fewer patients are allocated to futile doses (doses 1 and 2 in Scenario M5) and to unsafe doses (doses 4 and 5 in scenario M3). For Scenario 1 where dose 5 is the MTD, the first futility stopping rule identifies at least one dose as futile 35.6% of the time, increasing to 81.5% and 87.1% at the second and final futility assessment at the end of the trial respectively. Specific detail on the proportion of trials which recognise each dose as futile for Scenario 1 is presented in Table S7 in the Supplementary Materials. Further discussion on sensitivity to futility stopping rules is presented in Section 6 of the Supplementary Materials. Section 7 of the Supplementary Materials presents results for other sensitivity analyses referred to in Section 3.6.
We present the novel PRO-ADD dose-finding trial design – a modular framework for dose-optimisation trials. This innovative approach integrates three key outcomes: clinician-assessed DLTs, patient-assessed tolerability, and efficacy to determine the optimal dose. PRO-ADD dynamically adjusts dosing based on clinician-assessed DLTs and efficacy, and subsequently incorporates PRO-nAE burden score at the final analysis to recommend the most appropriate dose. The inclusion of patient-assessed tolerability provides PRO-ADD with superior probability of correct selection compared to the U-BOIN design across the majority of simulation scenarios. To our knowledge, PRO-ADD is the first trial design to introduce these three critical endpoints, including the longitudinal assessment of PROs, within a dose-optimisation design, offering a comprehensive approach to determining the optimal dose.
Practical design considerations
Selection of final PRO assessment timepoint
When modelling longitudinal PRO data with PRO-ADD, trialists should carefully consider the timing of PRO assessments, including the selection of the final analysis timepoint. In line with OPTIMISE-ROR guidance, 43 timing should be defined at the commencement of the trial in collaboration with stakeholders including clinicians, statisticians, PRO methodologists, and patient partners. Selection of the final assessment time point should be informed by a range of considerations including the treatment’s tolerability profile, mechanism of action, and investigated administration schedules. 43 For treatments such as immunotherapies and MTAs1,9 which can produce low-grade toxicities that accumulate over time, extended tolerability assessments beyond the initial treatment cycles may be warranted. PRO-ADD provides trialists with the flexibility to select a final assessment timepoint that aligns with the requirements of the investigational treatment.
Selecting loss functions and acceptable losses
Selection of an appropriate loss function is essential to ensure numerical loss translates to a clinical equivalence. Practical considerations for elicitation of utility functions have been provided for dose-finding trials and can be similarly utilised for trials incorporating loss functions. 70 In particular, the elicitation of a loss function relies on identifying three points that are judged to be equally desirable, with this approach effectively utilised in practice. 71 In our PRO-ADD example using the PRO-nAE burden score, we may elicit two of these points simply by asking “Supposing a patient experiences no side effects, what is the minimal probability of preliminary efficacy which would be acceptable for a dose?” and “Supposing efficacy is guaranteed, what is the maximal PRO-nAE burden score which you would deem acceptable for a given dose?”. The final point to be elicited must be a compromise of both PRO-nAE burden score and probability of preliminary efficacy, with equal attractiveness to the other elicited points. To aid such decision making, the mean distribution of side effect severity for varying PRO-nAE burden scores may be presented to stakeholders, translating PRO-nAE burden score to a dose’s average tolerability profile (see Figure S2 of the Supplementary Materials). Loss functions should be elicited in collaboration with clinical teams and patients to ensure that decision making reflects both clinical and patient priorities.
PRO-ADD in the presence of intercurrent events
Within the supplementary materials of this manuscript, we investigate PRO-ADD design performance considering two intercurrent events though other intercurrent events may occur in practice. To maintain design performance in light of trial-specific intercurrent events which may occur, trialists are encouraged to identify and select handling strategies for relevant intercurrent events upfront and evaluate design performance in light of such approaches. General guidance for the handling of intercurrent events within dose-finding and dose-optimisation trials should also be followed.69,72,73
Additional design considerations
The modular framework of this proposed design allows trialists to tailor design characteristics to their individual needs.
Stages 1 and 2
To complete the dose-escalation and optimisation routines in Stages 1 and 2 of PRO-ADD, we may consider any dose-escalation/optimisation trial design including extensions to the original BOIN design. Since its original publication, BOIN has been extended to introduce efficacy endpoints within interim and final decision making, including BOIN12, 61 BOIN-ET 74 and U-BOIN. 60 BOIN12 and BOIN-ET consider binary DLT and binary efficacy responses within a one and two stage design respectively. U-BOIN extends toxicity and efficacy to categorical responses with a utility function introduced to inform interim and final dose decision making.
The PRO-ADD trial design was developed in line with recommendations for the incorporation of PROs within early phase trials13,43 which recommends that PROs be integrated at final analyses to inform dose-selection. Whilst ePROs are an effective way to capture PROs, 28 paper-based PROs are still common. The extended administrative process and data collection period for paper-based PROs may reduce the current potential for PROs to be effectively used in adaptive decision-making. However, as ePROs collection becomes increasingly common, future work may wish to introduce PROs within adaptive, interim decision making. 13 For example, a trialist may also wish to introduce the continuous PRO-nAE burden score within the defining of admissible doses in Stage 2 of PRO-ADD. By utilising conjugacy models for continuous data, a tolerability stopping rule can be introduced similarly to the futility and safety stopping rules presented in this paper.
Final dose recommendation
The decision criterion defined in this article utilises a loss to make a final dose recommendation. Loss functions are implicitly embedded within the majority of early phase trials. For example, the
By utilising a loss for the determination of the recommended dose, PRO-ADD performance depends on good predictive accuracy for PRO-nAE burden score and response rate for each dose. To ensure good convergence of the linear mixed effect model for PRO-nAE burden score, the presented PRO-ADD formulation is recommended for larger dose-optimisation trials. Trialists may wish to draw on existing knowledge of dose–response and toxicity relationships for the novel therapy to inform model building. Without such prior knowledge, the PRO-ADD model presented here provides a generic framework for estimation.
Trialists who wish to utilise PRO-ADD with a smaller sample size should consider a more parsimonious model or simpler conjugacy models. In such instances, trialists should take particular care evaluating model performance in light of missing data as a vital sensitivity analysis.
The modular framework enables trialists to select the most appropriate analysis methods for their specific needs. This could include the analysis of efficacy endpoints within interim and final analyses. Whilst use of binary efficacy endpoints is currently common practice, there is flexibility to adapt and modify these methods as practices evolve – such as incorporating alternate conjugate models for the analysis of continuous response data rather than discrete data. Whilst the majority of response endpoints remain binary, there is increasing interest and guidance supporting the use of continuous biomarkers in early phase drug development. 4 Draft guidance has been developed by the FDA to support the inclusion of the Circulating Tumor DNA (CtDNA) biomarker within cancer clinical trials, 81 and this biomarker has previously been utilised within novel dose-finding trial designs. 82 Whilst iPIPE has been implemented in this paper to estimate the probability of binary response, it can be used similarly with a normal conjugate model for continuous response data. Further details of this approach are provided in Section 8 of the Supplementary materials.
Point estimation for endpoints
Although the estimate of expected loss is the primary focus for final dose recommendation in PRO-ADD, trialists may still find value in the point estimates for efficacy and PRO-nAE burden score estimated in the trial. Estimates of bias and MSE for iPIPE and beta regression PRO-nAE burden score estimates are presented in Section 9 of the Supplementary Materials.
PRO summary scores
Whilst we consider the PRO-nAE burden score within this manuscript, PRO-ADD can be alternatively utilised when PROs are summarised using the Total Toxicity Profile score. 35 In such a setting, collaboration with patients and key stakeholders can help customise a weight matrix to emphasise toxicities of greater concern. However, in cases where eliciting a weight matrix may be challenging, the proposed PRO-nAE burden score provides a straightforward alternative, with equal weighting of all toxicity types being sufficient.
Trialists may also consider using alternative PRO measures to define the PRO-nAE burden score. Whilst PRO-CTCAE has been employed as a case study in this article to summarise patients’ symptomatic adverse events, other PRO measures that assess additional tolerability concepts can also be explored. Examples could include EORTC QLQ-C30 83 which is used to assess quality of life and identified as one of the most common PROMs within early phase dose-finding oncology trials. 23
The longitudinal assessment of PROs will become increasingly important as we look to assess patients’ tolerability to treatment beyond the traditionally short DLT assessment periods.
Further considerations
Having demonstrated the core operating characteristics of the proposed PRO-ADD design, we are now building on this work to explore the design’s performance under additional practical considerations. This includes the evaluation of analytical strategies for handling intercurrent events (e.g. treatment discontinuation) and missing data – both of which can impact outcome interpretation and design performance. These ongoing efforts aim to support the framework’s robust application and facilitate its adoption in real-world early phase trial settings.
Conclusion
This article presents PRO-ADD, a new modular trial design framework for dose-optimisation integrating clinician-assessed DLTs, PROs, and preliminary efficacy. By incorporating PROs, PRO-ADD ensures the selected dose is not only active but also considered safe and tolerable from both clinician and patient perspectives, setting a new standard for patient-centred dose-optimisation. We illustrate an example application of PRO-ADD, demonstrating the design’s strong performance identifying the most active and tolerable dose whilst avoiding unnecessary escalation to higher doses offering no additional benefit. As clinical development advances, incorporating patient-centric outcomes such as PROs will be valuable for refining dose-finding strategies – ensuring that dose decisions balance clinical benefit, patient-experienced tolerability, and quality of life.
Supplemental Material
sj-pdf-1-smm-10.1177_09622802261435969 - Supplemental material for PRO-ADD: Patient-empowered dose-finding trials integrating safety, preliminary efficacy and patient-reported outcomes for optimal dose selection
Supplemental material, sj-pdf-1-smm-10.1177_09622802261435969 for PRO-ADD: Patient-empowered dose-finding trials integrating safety, preliminary efficacy and patient-reported outcomes for optimal dose selection by Emily Alger, Sumithra J Mandrekar, Jun Yin and Christina Yap in Statistical Methods in Medical Research
Footnotes
Acknowledgements
Ethical approval and informed consent
Not applicable.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: EA has been supported to undertake this work as part of a PhD studentship from the Institute of Cancer Research within the MRC/NIHR Trials Methodology Research Partnership. CY receives programmatic infrastructure funding from Cancer Research UK (CTUQQR-Dec22/100004), which supported this work.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
