Creation of a database of case-identifying algorithms developed using Canadian administrative health data

At Broadstreet we do a lot of work with real world evidence (RWE), frequently using administrative health data sources. Being able to standardize how cohorts for these types of studies are identified is important as it ensures that research is comparable across different databases and jurisdictions. These types of studies rely on case-identifying algorithms, to define study population, exposure, outcome, and covariate definitions. Validated case-identifying algorithms are useful and necessary because they help researchers reliably define study definitions and increase reproducibility and transparency.1 These algorithms are programmed to look for specific combinations of administrative data fields, including diagnostic codes (International Classification of Diseases [ICD] 9 or 10 codes) or prescription dispensing records. These algorithms are generally validated through a process of chart review and/or clinical validation to then generate performance metrics including sensitivity, specificity, positive predictive value, negative predictive value, to inform decisions on the most suitable algorithm for each study need.

Life-cycle of the development, validation, and applications of case-identifying algorithms in RWE studies

However, despite the importance for researchers to standardize how cohorts are identified, no centralized resource currently exists to summarize existing algorithms and their performance in the Canadian context. Because of this, a Broadstreet team lead by Christina Qian and Alexa Bowie conducted a targeted literature review with the aim of identifying validated algorithms to include in a living database of algorithms developed using Canadian administrative health data for future research use. The review identified studies that developed and/or validated algorithm(s) to identify cases of diseases, conditions, tests, procedures, prescriptions, and/or healthcare resource use using Canadian administrative health data.2,3 In the process, 270 studies were identified, and 82 eligible studies reporting 1453 algorithms were summarized. Of the 82 studies, 48 (58.5%) used data from Ontario which is Canada’s most populous province. Chronic kidney disease, diabetes, hypertension, and rheumatoid arthritis were the most frequent areas of focus (each representing ~4% of included studies).

Flowchart of living database lifecycle and application

The team is now pleased to make the findings of the review available as a public resource on the Broadstreet website. The data have been collected in a downloadable spreadsheet. For each publication, data are provided on the ICD-10 or other code used in the algorithm, the disease/condition, the geographical source of the data, population age, the URL of the publication. Our intention is for this to be a living database, which will be regularly updated in order to support researchers using Canadian administrative health data in real-world data studies.

If you have a study you feel should be added to the database or would like to discuss how Broadstreet could aid your company in the development of a real-world data study, get in touch.

References

  1. Wang SV, Pottegård A. Building transparency and reproducibility into the practice of pharmacoepidemiology and outcomes research. Am J Epidemiol. 2024;193(11):1625-1631. doi:10.1093/aje/kwae087
  2. Bowie, A. C., Tinajero, M. G., & Qian, C. (2024). EPH193 Validated Case-Identifying Algorithms Using Canadian Administrative Health Data: A Targeted Literature Review. Value in Health, 27(6), S187.
  3. Tinajero, M. G., Bowie, A. C., & Qian, C. (2024). [1369] Validated Case-Identifying Algorithms Using Canadian Administrative Health Data: A Targeted Literature Review. Pharmacoepidemiology and Drug Safety, 33(S2), 619.