Beyond Biomarkers: What Truly Defines Therapeutic Benefit in Clinical Trials?

In drug development, regulatory evaluation, and health economics and outcomes research (HEOR), the definition of clinical evidence is undergoing a profound shift. Historically, demonstrating a statistical improvement in a laboratory biomarker or imaging metric—such as tumor shrinkage, reduced serum protein levels, or lower blood pressure—was often deemed sufficient to claim therapeutic success. Because these surrogate metrics were objective and easy to quantify, clinical trials naturally organized around them.

However, in oncology, chronic diseases, and rare disorders, clinicians and researchers frequently encounter a difficult reality: a tumor may shrink on a scan or a biomarker may normalize, while the patient’s physical vitality, functional capacity, and daily well-being sharply decline. Today, regulators, payers, and patients ask a simpler, more essential question: How does an intervention actually affect how a patient feels, functions, or survives in daily life?

To answer this question, researchers rely on a structured methodological framework: Clinical Outcome Assessments (COAs).

This article marks the first of a four-part series exploring the science and strategy behind COAs—moving from conceptual foundations to qualitative development, psychometric validation, and clinical trial implementation.

Redefining “Treatment Benefit”

Every clinical trial aims to demonstrate a treatment benefit. Yet in regulatory science, this phrase carries a strict and specific definition.

According to consensus frameworks established by the International Society for Pharmacoeconomics and Outcomes Research (ISPOR) and regulatory agencies, a treatment benefit is:

A favorable effect on a meaningful aspect of how a patient feels or functions in their typical life, or on their survival.

Two elements of this definition deserve particular attention:

  • A Meaningful Aspect of Health: The clinical effect must target something the patient cares about and actively wants to improve, stabilize, or avoid. This definition matters because it prevents sponsors from claiming clinical benefit based solely on biological activity that patients never actually experience in daily life.
  • In Their Typical Life: An artificial task completed only within a clinic has little regulatory value unless researchers can prove it directly corresponds to how the patient functions in everyday life outside the medical center.

Biomarkers vs. Clinical Outcome Assessments (COAs)

To design sound endpoints, investigators must distinguish between biomarkers and COAs:

  • Biomarkers: Objective physical, chemical, or imaging measurements—such as blood pressure, HbA1c, or tumor dimensions on an MRI. They operate independently of patient volition, effort, or rater judgment. While biomarkers indicate biological activity or serve as surrogate endpoints, they do not directly evaluate how a patient feels or functions. A heart failure patient’s BNP levels may improve, for example, even while they remain too short of breath to climb a flight of stairs.
  • Clinical Outcome Assessments (COAs): Measurement tools that generate a rating or score representing a patient’s health status. Unlike biomarkers, all COAs rely on human judgment, motivation, or active effort.

The Four Quadrants of COAs

COAs fall into four distinct categories based on who provides the rating and how judgment is applied:

1. Patient-Reported Outcome (PRO)

  • Rater: The patient.
  • Core Focus: Direct reporting of health status without third-party interpretation.
  • Application: The gold standard for subjective experiences such as pain, nausea, fatigue, and psychological distress.

2. Clinician-Reported Outcome (ClinRO)

  • Rater: A healthcare professional with specialized training.
  • Core Focus: Uses professional clinical judgment to observe, interview, and score a patient’s condition (e.g., psychiatric rating scales, physical examination rubrics).
  • Considerations: Prone to inter- and intra-rater variability; requires rigorous rater training and standardized scoring definitions.

3. Observer-Reported Outcome (ObsRO)

  • Rater: A non-clinician observer, such as a parent, spouse, or caregiver.
  • Core Focus: Reports observable actions and events in daily life without requiring medical expertise (e.g., infant feeding behavior, nighttime wandering in dementia).
  • The Rule Against “Proxy” Reporting: Observers must strictly record only what they can directly observe (e.g., “The patient stayed in bed for 10 hours”), rather than attempting to infer internal states (e.g., “The patient felt depressed”). Regulators reject proxy inference because caregivers cannot reliably quantify another person’s subjective feelings.

4. Performance Outcome (PerfO)

  • Rater: None (standardized task execution).
  • Core Focus: Quantifies a standardized physical or cognitive task (e.g., the 6-Minute Walk Test, timed pegboard tasks).
  • Considerations: While observer judgment is eliminated, the score remains highly dependent on the patient’s effort, volition, and motivation on that day.

Direct vs. Indirect Measurement and the Concept of Interest (COI)

Investigators can structure an endpoint to measure a health benefit either directly or indirectly:

  • Direct Measurement: The COA assesses the meaningful health aspect itself. A PRO measuring pain intensity directly evaluates the patient’s lived reality.
  • Indirect Measurement: When a meaningful daily activity is too complex or environmentally variable to standardize directly—such as a dementia patient’s capacity for independent community living or a COPD patient’s daily physical activity—researchers isolate a measurable, standardized functional ability: the Concept of Interest (COI).

The appropriate COI differs across diseases: gait speed for neuromuscular disorders, delayed memory recall for early Alzheimer’s disease, or inspiratory muscle strength for COPD. When using indirect tools, sponsors must empirically show that changes in the COI translate to meaningful improvements in daily life—demonstrating, for instance, that a 30-meter gain on a 6-Minute Walk Test genuinely enables a patient to shop independently or navigate stairs at home.

Context of Use (COU): Defining the Boundaries

No outcome measure is universally valid. A COA is reliable and fit-for-purpose only within its specified Context of Use (COU), which defines:

  • Target Population: Disease severity, subtype, age, culture, and language.
  • Study Design & Role: Primary vs. secondary endpoint, timing, and assessment frequency.
  • Setting & Mode: In-clinic vs. decentralized/home-based, paper vs. electronic devices (eCOA).

In real-world trials, altering any key parameter of the COU introduces practical risks. Migrating a validated paper questionnaire to an electronic device (eCOA) can alter how patients interact with response scales, requiring formal measurement equivalence testing. Similarly, translating an instrument for global trials without rigorous cultural adaptation can compromise content validity.

In Part 2 of this series, we will examine the qualitative research methods and conceptual modeling techniques required to build robust, patient-centered measures from the ground up.


Note: This series draws on and adapts core concepts from Genentech’s Coursera course “Data Sciences in Pharma: Patient Centered Outcomes Research,” together with the FDA’s Patient-Focused Drug Development (PFDD) guidance.