Skip to main navigation Skip to search Skip to main content

From Clinical Trials to Real-World Impact: Introducing a Computational Framework to Detect Endpoint Bias in Opioid Use Disorder Research

  • The ENDPOINT Consortium
  • Florida International University
  • City University of New York
  • Stony Brook University
  • University of Wisconsin-Madison

Research output: Contribution to journalArticlepeer-review

Abstract

Introduction: Clinical trial endpoints are a ‘finite sequence of instructions to perform a task’ (measure treatment effectiveness), making them algorithms. Consequently, they may exhibit algorithmic bias: internal and external performance can vary across demographic groups, impacting fairness, validity and clinical decision-making. Methods: We developed the open-source Detecting Algorithmic Bias (DAB) Pipeline in Python to identify endpoint ‘performance variance’—a specific algorithmic bias—as the proportion of minority participants changes. This pipeline assesses internal performance (on demographically matched test data) and external performance (on demographically diverse validation data) using metrics including F1 scores and area under the receiver operating characteristic curve (AUROC). We applied it to representative opioid use disorder (OUD) trial endpoints. Results: F1 scores remained stable across minority representation levels, suggesting consistency in precision-recall balance (F1) despite demographic shifts. Conversely, AUROC measures were more sensitive, revealing significant performance variance. Training on demographically homogeneous populations boosted internal performance (accuracy within similar cohorts) but critically compromised external generalisability (accuracy within diverse cohorts). This pattern reveals an ‘endpoint bias trade-off’: optimising performance for homogeneous populations vs. having generalisable performance for the real world. Discussion and Conclusions: Consistently performing endpoints for one demographic profile may lose generalisability during population shifts, potentially introducing endpoint bias. Increasing minority representation in the training data consistently improved generalisability. The endpoint bias trade-off reinforces the importance of diverse recruitment in OUD trials. The DAB Pipeline helps researchers systematically pinpoint when an endpoint may suffer ‘performance variance’ (i.e., bias). As an open-source tool, it promotes transparent endpoint evaluation and supports selecting demographically invariant OUD endpoints.

Original languageEnglish
Article numbere70085
JournalDrug and Alcohol Review
Volume45
Issue number1
DOIs
StatePublished - Jan 2026

Keywords

  • algorithmic bias
  • demographic parity
  • open-source software
  • opioid use disorder
  • performance variance

Fingerprint

Dive into the research topics of 'From Clinical Trials to Real-World Impact: Introducing a Computational Framework to Detect Endpoint Bias in Opioid Use Disorder Research'. Together they form a unique fingerprint.

Cite this