A published responder rate is not a finished patient-benefit claim
When a cosmetic clinic brochure, manufacturer press release, or aesthetic conference presentation declares that an injectable filler, energy device, or neurotoxin achieved an “88% responder rate” or a “93% patient response,” prospective aesthetic patients naturally assume they are reading an unambiguous statement of clinical efficacy. In the consumer imagination, an 88% responder rate sounds like an 88% probability of achieving an obvious, flattering, and durable aesthetic transformation. In reality, a raw responder percentage is an incomplete statistical calculation. Without its surrounding methodological architecture, a responder rate does not communicate whether the change was a clinician’s scale step, a patient’s questionnaire score, or a percentage computed after missing follow-up was dropped from the denominator.
To evaluate any published cosmetic responder percentage, a patient or clinician must reconstruct six foundational variables that protocol authors and biostatisticians specify before a clinical trial begins:
Responder definition: The exact grading scale used and the precise score-change threshold required for a participant or anatomical site to be categorized as a “responder.”
Rater and assessment type: Who performed the evaluation—a trained clinician after observation (ClinRO), the patient without clinician interpretation (PRO), a non-clinician observer (ObsRO), a scored performance task (PerfO)—or whether the rater is unknown. An unblinded treating injector is not interchangeable with a protocol-defined ClinRO.
Baseline severity and eligibility criteria: The entry score required to enroll in the study, which dictates whether a numerical change was mathematically and clinically achievable.
Denominator and unit of analysis: The exact cohort evaluated (n/N), distinguishing individual human participants from treated anatomical subunits such as nasolabial folds, periocular areas, or hands, and separating full intent-to-treat groups from per-protocol completers.
Assessment timepoint: The precise post-treatment elapsed duration when the evaluation occurred, distinguishing immediate post-procedural edema from true mid-term or long-term structural tissue response.
Missing follow-up and retreatment censoring rules: How the statistical analysis accounted for missed visits, dropouts, asymmetry corrections, and touch-up retreatments.
Without these six operational variables, two studies quoting identical “85% responder rates” may describe fundamentally incompatible clinical realities. TrialEndpoints is a publication, not a contract research organization (CRO), statistician, regulator, or source of trial results. A related sponsor-facing page, fit-for-purpose COA selection under FDA PFDD Guidance 3, is written for protocol authors; that instrument-selection worksheet is not this patient reconstruction job. The worksheet below is for a prospective aesthetic patient who needs to name missing fields before comparing any published cosmetic responder percentage.
This decoder is distinct from broad trial landscape surveys, such as our analysis of the aesthetic clinical trial pipeline, or our regulatory review of regenerative aesthetics trial evidence. Rather than cataloging active investigational pipelines, this guide equips patients to dissect any published percentage they encounter on a med spa menu, brand website, or trial registry.
What FDA means by an endpoint, a COA, and a responder
To understand why published responder percentages cannot be compared at face value, one must examine how regulatory agencies define clinical evidence. In its Patient-Focused Drug Development (PFDD) Glossary (content current as of 5 March 2026), the U.S. Food and Drug Administration (FDA) defines an endpoint as a precisely defined variable intended to reflect an outcome of interest that is statistically analyzed to address a particular research question. The agency explicitly notes that a precise definition typically specifies the type of assessments made, the timing of those assessments, the assessment tools used, and detailed instructions for how multiple assessments within an individual are to be combined. A percentage that lacks any of these components is not a complete regulatory endpoint; it is merely an isolated mathematical numerator divided by an arbitrary denominator.
The PFDD framework categorizes all Clinical Outcome Assessments (COAs) into four distinct types based on who provides the information:
Patient-Reported Outcome (PRO): A measurement based on a report that comes directly from the patient about the status of a patient’s health condition without interpretation of the patient’s response by a clinician or anyone else.
Clinician-Reported Outcome (ClinRO): A measurement based on a report that comes from a trained health-care professional after observation of a patient’s health condition. Crucially, FDA notes that a ClinRO cannot directly assess symptoms or feelings known only to the patient.
Observer-Reported Outcome (ObsRO): A measurement based on a report of observable signs, events, or behaviors by someone other than the patient or a healthcare professional, such as a parent or caregiver.
Performance Outcome (PerfO): A measurement based on a standardized task performed by a patient that is administered and evaluated by an appropriately trained individual or completed independently. PerfOs require patient cooperation and motivation.
Aesthetic filler and toxin labels commonly report ClinRO wrinkle-scale ratings, PRO questionnaires, or both. Because a ClinRO cannot directly assess symptoms known only to the patient, and a PRO does not apply a clinician’s photonumeric grading rule, the two assessment classes measure different phenomena. They are not interchangeable sources of the same percentage.
The concept of a responder definition is formalized in FDA’s December 2009 guidance, Patient-Reported Outcome Measures: Use in Medical Product Development to Support Labeling Claims. FDA defines a responder definition as a score change in a measure, experienced by an individual patient over a predetermined time period, that has been demonstrated in the target population to have a significant treatment benefit. The guidance instructs sponsors to formulate an a priori responder definition—established before examining study data—rather than engaging in post hoc selective picking (“cherry-picking”) of favorable score thresholds to craft advertising claims. Furthermore, FDA emphasizes that an empirical responder definition is derived using anchor-based clinical methods and is evaluated strictly within the context of each specific trial. A responder definition is not a universal constant that transfers freely across different medical devices, drug formulations, or anatomical indications.
In October 2025, FDA finalized PFDD Guidance 3, Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments. Guidance 3 reinforces the vital distinction between a COA (the complete measurement system, including user instructions, items, and scoring algorithms), a COA score (the numerical rating generated), and a COA-based endpoint (the variable analyzed in a clinical trial protocol). Guidance 3 states that using COAs to construct trial endpoints, including how scores become responder definitions, is discussed in PFDD Guidance 4 (April 2023). FDA’s Guidance 4 landing page, content current as of 5 April 2023, still labels the document not for implementation and states that it contains nonbinding recommendations. This article mentions that draft status only to keep endpoint-construction advice out of the patient worksheet; it does not implement Guidance 4. There is no uniform responder formula that transfers across aesthetic products, scales, or follow-up times.
Finally, the PFDD glossary defines fit-for-purpose as a conclusion that the level of validation associated with a tool is sufficient to support its specific context of use (COU). Fit-for-purpose is therefore not a finding that a threshold, scale, rater, population, or follow-up time transfers to another study.
The same glossary defines clinical relevance as the extent to which an endpoint can capture and measure an aspect of a potential clinical benefit—improvement in how the patient feels, functions, and/or survives—that is important from a clinical perspective and from the patient’s perspective. Statistical significance of a responder percentage is not that determination.
Who rated it is a different analysis, not a second copy of the same percentage
One of the most frequent sources of confusion in aesthetic marketing is the conflation of clinician ratings with patient self-assessments. When an advertisement states that “95% of subjects responded at Month 6,” the reader cannot know whether that percentage is a clinician’s one-grade change on a wrinkle scale or a patient’s report of improvement on a questionnaire.
The divide between clinician and patient ratings was heavily scrutinized during the FDA’s March 23, 2021 General and Plastic Surgery Devices Panel meeting on dermal fillers. In its executive summary for the 23 March 2021 meeting, FDA CDRH stated that the primary effectiveness evaluation is traditionally a ClinRO that does not incorporate input from the patient. Dermal-filler studies typically include patient input as secondary or ancillary endpoints. The panel cites FACE-Q as a used PRO and states that GAIS is not a validated PRO.
A clear real-world demonstration of how rater perspective alters reported percentages appears in the FDA Summary of Safety and Effectiveness Data (SSED) for Juvéderm Vollure XC (PMA P110033/S020). In the pivotal clinical trial supporting approval for moderate-to-severe nasolabial folds (NLFs), the study tracked both blinded investigator ratings and direct patient appraisals at the same visits:
Investigator-rated response (ClinRO): At Month 6, the live Evaluating Investigator responder rate—defined as at least a 1-point improvement on the 5-point Nasolabial Fold Severity Scale (NLFSS)—was 93.2% (109/117 folds).
Patient-reported improvement (PRO): At the same Month 6 visit, a separate FACE-Q Appraisal of Nasolabial Folds analysis reported that 95.7% (112/117 subjects) reported improvement, with mean score increasing from 32.1 at baseline to 72.8. The SSED reads those scores as subjects being less bothered by nasolabial-fold depth; they are not the satisfaction rate below.
Patient satisfaction rate: When asked about overall treatment satisfaction on the same survey, 82.1% (96/117 subjects) reported being “very satisfied.”
Notice that within this single public SSED, three distinct percentages describe the Month 6 outcome: 93.2%, 95.7%, and 82.1%. If a clinic quotes 95.7%, it is highlighting patient-perceived improvement; if it quotes 93.2%, it is reporting fold-level investigator scoring; if it quotes 82.1%, it is reporting categorical high satisfaction. None of these numbers is false, but neither are they interchangeable. A ClinRO measures anatomical change against a standardized visual scale; a PRO measures subjective experiential impact. Neither can substitute for the other.
Baseline, n/N, and unit of analysis
Even when the rater and the scale are known, a published percentage remains meaningless without its denominator (n/N), its unit of analysis, and its baseline inclusion threshold. The National Library of Medicine’s ClinicalTrials.gov Results Quality Control Review Criteria explicitly mandate that every submitted outcome measure must report a comprehensive title, description, time frame, analysis population, and exact counts analyzed. The QC criteria require that whenever the unit of analysis is not participants—for example eyes, lesions, or implants—the type of units and the numbers of units analyzed must be provided. A fold-level cosmetic percentage is the same class of incomplete result if those unit counts are omitted.
In facial aesthetics, this unit-of-analysis distinction is frequently obscured in consumer summaries. Returning to the Vollure XC pivotal study, the primary trial design was split-face: subjects received the investigational filler in one nasolabial fold and an active control gel in the contralateral fold. Consequently, the primary effectiveness endpoint evaluated 117 nasolabial folds, drawn from a modified intent-to-treat (mITT) cohort within 123 randomized human subjects. Consumer summaries sometimes restate that figure as “93% of patients responded.” The SSED defines the co-primary as a percentage of nasolabial folds, not a percentage of people. Even when the fold count matches the number of analyzed subjects—one VOLLURE XC–treated fold per person in this split-face design—the unit of analysis is still the fold. Restating a fold-level rate as a people-level rate is a unit error, not a harmless shorthand.
Baseline severity requirements represent another critical filter. In the Vollure XC trial, protocol eligibility required that a participant exhibit a live Evaluating Investigator score of either 2 (moderate) or 3 (severe) on both nasolabial folds on the 0-to-4 NLFSS, yielding a study population baseline mean of 2.6. Patients with mild folds (grade 1) or extreme folds (grade 4) were excluded from enrollment.
This baseline threshold dictates the mathematical feasibility of achieving a “response.” A participant starting at grade 3 has ample room to improve by 1 full point (shifting to grade 2 or grade 1). If a clinic promotes this same 93.2% figure to someone who currently presents with mild grade 1 folds, the trial population does not match. Enrollment required a live Evaluating Investigator score of 2 on both folds or 3 on both. A one-point change from grade 1 to grade 0 was not the studied context of use, and the SSED is not a universal visible-benefit rule for milder lines.
Finally, the construction of the denominator (N) can dramatically distort published success rates. Consider three common statistical denominators:
Intent-to-Treat (ITT): Includes all participants who were randomized or received treatment, regardless of whether they completed the study, attended follow-up visits, or experienced protocol violations.
Modified Intent-to-Treat (mITT): Typically restricts the denominator to participants who received treatment and completed at least one post-baseline evaluation.
Per-Protocol Completers (PP): Restricts the denominator exclusively to subjects who completed all scheduled visits without protocol deviations, discarding missed visits and early dropouts.
In a hypothetical enrollment of 100 people, if 20 miss follow-up and 72 of the 80 completers meet the responder rule, a per-protocol caption can say “90%” (72/80). If missed visits are counted as non-response in the enrolled cohort, the same data are 72% (72/100). Whenever a published claim omits n/N and the analysis population, the reader cannot tell which denominator was used. That arithmetic is a reconstruction check, not a typical cosmetic result.
Timepoint and missing follow-up change the same study's later percentages
A responder percentage is attached to a visit. ClinicalTrials.gov results quality-control review criteria require each outcome measure to include a time frame a general reader can understand, including the time point(s) at which the measurement was assessed. Quoting a cosmetic responder rate without that visit is an incomplete public result, not a duration guarantee.
The evolution of responder percentages over time is vividly illustrated by tracking the long-term Evaluating Investigator NLFSS responder rates across the entire 18-month follow-up window in the Vollure XC SSED:
Month 6: 93.2% (109/117 folds assessed)
Month 9: 84.6% (99/117 folds assessed)
Month 12: 57.5% (65/113 folds assessed)
Month 15: 61.7% (50/81 folds assessed)
Month 18: 59.4% (57/96 folds assessed)
Examining this trajectory reveals two vital methodological lessons. First, the same SSED’s later Evaluating Investigator NLF percentages are not the Month 6 figure: they fall from 93.2% at six months to 59.4% at 18 months, with changing denominators. Product-and-area duration ceilings are a separate reader job in how long Juvéderm lasts on labeled duration ceilings; this article does not recast that duration table as a general percentage decoder.
Second, the denominator is not fixed: 117 folds at Month 6, 113 at Month 12, 81 at Month 15, and 96 at Month 18. The SSED states that N is the number of NLFs assessed within the analysis window. A later percentage that omits n/N is not the same public result as Month 6.
Crucially, the Vollure XC protocol enforced a rigorous censoring rule: any nasolabial fold that received touch-up correction for asymmetry or repeat treatment was classified as a non-responder at all subsequent time points. Under that SSED rule, later NLF percentages are not the Month 6 result restated, and they are not interchangeable with a caption that is silent on retreatment. If a published late-stage percentage omits the missing-follow-up or repeat-treatment rule, that field stays unknown and blocks comparison.
Across the wider aesthetic device and injectable market, timepoint discrepancies make cross-study comparisons impossible. FDA’s 2021 dermal-filler panel executive summary noted that across 13 PMA approvals for nasolabial-fold indications, 7 different effectiveness scales were used, with primary evaluations ranging from 12 weeks to 13 months after treatment, using a mix of live and photographic evaluation. Comparing an unnamed 85% at 12 weeks with an unnamed 85% at 12 months is not a product ranking; the panel’s point is that same-indication responder rates often cannot be compared.
A p-value and a one-grade rule do not travel
A common advertising tactic is to pair a high responder percentage with a statistically significant p-value (e.g., “93% response, p < 0.001”), creating the impression that the product was proven overwhelmingly superior to competing options. However, in biostatistics, statistical significance simply measures the probability that an observed result occurred by chance under a specific null hypothesis. It does not determine clinical magnitude, visible superiority, or patient satisfaction.
The Vollure XC SSED provides an instructive lesson in statistical interpretation. The protocol’s co-primary effectiveness hypothesis was tested against a pre-specified performance standard: specifically, whether the Month 6 responder rate was statistically greater than 50%. The observed rate of 93.2% easily surpassed 50%, yielding a dramatic p < 0.001. That p-value did not compare VOLLURE XC with the split-face control on mean change; it tested the fold-level responder rate against a pre-specified 50% standard.
When the trial analyzed the co-primary endpoint comparing the mean NLFSS score change between Vollure XC and its active split-face control gel, the results were far closer:
Vollure XC mean NLFSS improvement at Month 6: 1.4 points
Active control mean NLFSS improvement at Month 6: 1.3 points
Between-group comparison: p = 0.097 (not statistically significant)
While the responder rate against the 50% benchmark was highly significant (p < 0.001), the difference in mean NLFSS change versus the split-face control was 0.1 points (1.4 versus 1.3) and was not statistically significant (p = 0.097). Highlighting only p < 0.001 does not establish visible superiority over the control, and it does not make a p-value equal clinical relevance as defined in the PFDD glossary.
A related fallacy is the belief that a “1-grade change” is a universal standard of meaningful clinical improvement. The 2021 panel states that on the WSRS, an improvement greater than or equal to one grade has been used as a responder rule. That is a description of a rule that has been used, not a transferable meaningful-change threshold. Panel Table 9 also notes that ClinRO scales are typically 4- to 6-grade photonumeric instruments; a one-grade step is therefore not the same interval from scale to scale. Some labeled neurotoxin programs use a different rule entirely, as unpacked in Xeomin versus Botox labeled endpoints: Xeomin’s glabellar studies counted a responder only when both the investigator and the subject independently recorded at least a 2-grade improvement—not a one-grade wrinkle-scale change, and not the investigator-only “none or mild” rule used in Botox Cosmetic’s older glabellar trials. Those labeled endpoints are incompatible; they are not a ranking and not a universal cosmetic threshold.
Reconstruction worksheet and labeled hypothetical studies
To protect themselves against misleading commercial claims, prospective aesthetic patients and conscientious clinicians can utilize a structured evidence worksheet. Whenever a clinic, device brochure, or media report cites a responder rate, the reader should populate each column of the matrix below. If any critical cell must be marked “Unknown,” the published percentage cannot be compared to any other treatment.
The same TrialEndpoints page on fit-for-purpose outcome measures and context of use is adjacent sponsor reading on why COA type and context of use are named fields; it is still not a source of trial results and is not this decoder. The table below reconstructs named fields from the JUVÉDERM VOLLURE XC SSED as a public-document illustration, then three labeled hypothetical records. The SSED excerpt is not a treatment ranking, not a typical result, and not a recommendation. Hypothetical rows are internally consistent teaching examples, not comparative clinical evidence.
| Study Record / Scenario | Responder Definition & Scale | Rater & Assessment Type | Baseline Severity | Denominator & Unit (n/N) | Timepoint & Censoring Rule |
|---|---|---|---|---|---|
| SSED excerpt (not a ranking): JUVÉDERM VOLLURE XC P110033/S020, Month 6 | ≥1-point improvement on 5-point NLFSS (0–4 scale) | Live Evaluating Investigator (ClinRO); secondary FACE-Q PRO reported separately | Bilateral grade 2 (moderate) or grade 3 (severe); baseline mean 2.6 | 93.2% (109/117 folds); mITT analysis of 123 randomized subjects (unit = fold) | Month 6; tested vs 50% performance standard (p<0.001); mean change vs control was 1.4 vs 1.3 (p=0.097) |
| SSED excerpt (not a ranking): JUVÉDERM VOLLURE XC P110033/S020, Month 18 | ≥1-point improvement on 5-point NLFSS from original baseline | Live Evaluating Investigator (ClinRO) | Same baseline cohort (initial baseline mean 2.6) | 59.4% (57/96 folds); denominator reflects folds assessed at visit (unit = fold) | Month 18; folds receiving touch-up or repeat injection strictly categorized as non-responders |
| Hypothetical Case A: Med Spa Social Media Claim | Unknown (“Overall aesthetic success” on informal clinic survey) | Treating injector post-procedure review (unblinded clinician, not ClinRO) | Unrestricted / unrecorded; patients seeking voluntary cheek enhancement | Reported as “92%”; raw participant count and unit of analysis unstated | Immediate post-procedure (Day 0–14); dropouts and dissatisfied clients uncounted |
| Hypothetical Case B: Submental Tightening Device | ≥1-grade improvement on 4-point Submental Laxity Scale | Blinded independent expert photographic review (photographic ClinRO) | Enrolled subjects with baseline grade 2 (moderate) or grade 3 (severe) | 87.5% (70/80 completers); per-protocol cohort (ITT denominator was 100 enrolled) | Week 12 post-final session; subjects lost to follow-up excluded from denominator |
| Hypothetical Case C: Topical Skin-Booster Study | ≥10-point improvement on patient satisfaction questionnaire | Patient-reported outcome (PRO) via electronic patient diary | Self-reported skin dullness; no objective barrier or severity threshold | 85.0% (85/100 participants); unit = individual human subject | Month 3; missing survey responses imputed using last-observation-carried-forward |
Analyzing the five rows of this matrix illustrates why surface-level percentages are entirely deceptive:
Deconstructing Hypothetical Case A: In this labeled hypothetical, the scale is unnamed, the rater is the unblinded treating injector, the visit is immediate post-procedure, and dropouts are uncounted. With those fields unknown or incompatible, the “92%” cannot be compared with a protocol-defined ClinRO percentage.
Deconstructing Hypothetical Case B: This labeled hypothetical uses blinded photographic ClinRO review and a named ≥1-grade rule, but the 87.5% (70/80) caption is a completer denominator. The table also records an enrolled ITT denominator of 100. If the 20 missing participants are not in the numerator, 70/100 is 70%, not 87.5%. The point is the missing n/N and analysis-population fields, not a ranking of tightening devices.
Deconstructing Hypothetical Case C: This study reports an “85% response,” but the rater is the patient responding to a subjective satisfaction questionnaire (PRO) rather than an objective clinical rater. A high score indicates patient enjoyment of the regimen, but it cannot be compared against a clinician-graded wrinkle score.
SSED excerpt (rows 1 and 2): The SSED excerpt shows named fields in a public document: the NLFSS ≥1-point rule, the live Evaluating Investigator ClinRO, a separate FACE-Q PRO, the baseline enrollment rule, fold-level n/N, the Month 6 versus-50% test beside the 1.4 versus 1.3 mean-change comparison, later denominators, and the rule that folds receiving asymmetry correction or repeat treatment were non-responders afterward. Those fields are label-specific. They are not a typical result and not a product ranking.
Sources
U.S. Food and Drug Administration: Patient-Focused Drug Development Glossary (Current as of March 5, 2026; defines clinical outcome assessments, clinical relevance, endpoints, and fit-for-purpose context of use).
U.S. Food and Drug Administration CDER/CBER/CDRH: Patient-Reported Outcome Measures: Use in Medical Product Development to Support Labeling Claims (December 2009) (Nonbinding 2009 guidance: a responder definition is an individual score change over a predetermined time period demonstrated in the target population; use an a priori definition rather than cherry-picking post hoc PRO labeling claims).
U.S. Food and Drug Administration CDER/CBER/CDRH: Patient-Focused Drug Development: Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments (October 2025) (Final guidance distinguishing COAs, COA scores, and trial endpoints; references draft Guidance 4 for endpoint construction).
U.S. Food and Drug Administration: Patient-Focused Drug Development: Incorporating Clinical Outcome Assessments Into Endpoints for Regulatory Decision-Making (Draft Guidance, April 2023) (Draft regulatory guidance designated not for implementation).
ClinicalTrials.gov: Results Quality Control Review Criteria (U.S. National Library of Medicine criteria requiring explicit specification of outcome measure titles, time frames, analysis populations, and non-participant units of analysis).
U.S. Food and Drug Administration Center for Devices and Radiological Health: Summary of Safety and Effectiveness Data for JUVÉDERM VOLLURE XC (PMA P110033/S020) (Pivotal clinical study data reporting NLFSS investigator-rated responder rates, split-face control comparisons, FACE-Q PRO scores, and long-term censoring rules).
U.S. Food and Drug Administration Center for Devices and Radiological Health: Executive Summary for the General Issues Panel Meeting on Dermal Fillers (March 23, 2021) (Examines variability across 13 PMA approvals, heterogeneity of wrinkle rating scales, and the fundamental separation of clinician-reported and patient-reported outcomes).




