TriNetX and Real-World Evidence: A Critical Review of Its Strengths, Limitations, and Bias Considerations in Clinical Research
Mahmoud NassarⒾ1*, Hazem AbosheaishaaⒾ2, Khaled ElfertⒾ3, Azizullah BeranⒾ4, Abdellatif IsmailⒾ5, Mouhand MohamedⒾ6, Anoop MisraⒾ7, Muhammed Amir EssibayiⒾ8, David J. AltschulⒾ8, Ahmed Y. AzzamⒾ9
- 1Department of Medicine, Jacobs School of Medicine and Biomedical Sciences, University at Buffalo, Buffalo, NY, USA
- 2Internal Medicine Department, Icahn School of Medicine at Mount Sinai, NYC H+H Queens, New York, NY, USA
- 3Division of Gastroenterology, West Virginia University School of Medicine, Morgantown, WV, USA
- 4Division of Gastroenterology and Hepatology, Indiana University, Indianapolis, IN, USA
- 5Department of Internal Medicine, University of Maryland Medical Center Midtown, Baltimore, MD, USA
- 6Division of Gastroenterology and Hepatology, Mayo Clinic, Rochester, MN, USA
- 7National Diabetes, Obesity and Cholesterol Foundation (N-DOC), New Delhi, Delhi, India
- 8Department of Neurological Surgery, Montefiore Medical Center, Albert Einstein College of Medicine, Bronx, NY, USA
- 9Montefiore-Einstein Cerebrovascular Research Lab, Montefiore Medical Center, Albert Einstein College of Medicine, Bronx, NY, USA
Abstract
Introduction: The increasing utilization of real-world data platforms in medical research necessitates a comprehensive understanding of their methodological strengths and limitations. TriNetX has emerged as a significant platform for exploring large healthcare datasets. This review aims to critically evaluate TriNetX's methodological framework and limitations, assess the impact of electronic health record coding accuracy on data reliability, and analyze the platform's capacity for generating generalizable real-world evidence in clinical research.
Methods: We conducted a comprehensive review examining TriNetX's data architecture, quality metrics, and research applications, focusing on data integrity, platform architecture, and the external validity of research findings.
Results: The analysis reveals significant methodological considerations. TriNetX's reliance on retrospective data introduces biases such as selection bias and confounding variables. The coding accuracy of electronic health records, which have not been independently validated, is a critical determinant of data reliability. The demographic representation is limited, affecting the generalizability of results.
Discussion: Despite its extensive use, TriNetX's effective utilization requires careful consideration of its inherent limitations. The platform's data, predominantly from insured populations in academic and acute care settings, may not fully represent broader demographic groups. Addressing these methodological constraints is crucial for enhancing the reliability and applicability of research findings derived from TriNetX.
Conclusions: TriNetX is a valuable resource for healthcare research. However, its limitations must be acknowledged, and future research should focus on standardizing data collection and enhancing data validation processes to mitigate platform-specific biases and improve the quality and applicability of the findings.
Keywords: TriNetX, Clinical Research, Data Quality, Real-World Evidence, Research Methodology
Article information
Introduction
The increasing reliance on real-world data platforms in medical research has transformed the landscape of clinical evidence generation. Among these platforms, TriNetX has emerged as a significant player, providing real-time access to anonymized electronic medical records (EMR) and claims data from millions of patients globally [1]. While several studies have utilized TriNetX for large-scale observational research, a comprehensive systematic evaluation of its methodological framework, limitations, and research applications remains notably absent from the literature [2]. Current knowledge gaps in real-world data platforms center around three critical areas that demand systematic investigation. These encompass the impact of data quality variations on research outcomes, the methodological considerations in managing platform-specific biases, and the external validity of findings derived from such platforms [3]. Despite the growing adoption of TriNetX in medical research, existing literature has primarily focused on individual study outcomes rather than systematically evaluating the platform’s capabilities and limitations. This gap is particularly significant given the platform’s increasing influence on clinical decision-making and research methodology [4].
The novelty of our systematic review lies in its comprehensive review of TriNetX’s architecture, data quality metrics, and research applications through the methodological lens. Unlike previous platform-specific analyses, which typically focus on individual aspects such as data completeness or specific disease outcomes, our review provides an integrated assessment of the entire research ecosystem. This approach enables a deeper understanding of how various components, from data capture to analysis, interact and influence research outcomes [5,6,7,8].
TriNetX has revolutionized evidence-based medical research through its extensive network of healthcare systems. The platform facilitates rapid cohort identification and supports comprehensive research capabilities, including instant dataset queries for patient demographics, comorbidities, and treatment histories. This technological framework enhances study design efficiency and execution while implementing advanced analytical tools, such as propensity score matching, to minimize bias in observational studies and improve finding reliability [9,10,11].
However, the platform’s dependence on aggregated, real-world data presents significant methodological challenges that warrant careful consideration. The integrity of research outcomes may be compromised by potential issues in data completeness and accuracy, as information is sourced from EMR and claims data that may contain inherent errors or inconsistencies [12,13]. These challenges are further complicated by variations in coding practices, documentation standards, and data collection methods across healthcare systems [14,1,15]. While TriNetX offers sophisticated analytical tools, observational studies remain vulnerable to unmeasured confounding, limiting causal inference compared to randomized controlled trials. Furthermore, the representativeness of the data warrants careful consideration, as it reflects specific patient populations that may not adequately represent broader demographic groups [16,17].
This systematic review addresses three fundamental objectives in understanding and optimizing TriNetX’s research applications. First, it comprehensively evaluates the platform’s methodological framework and its implications for research quality. Second, it assesses the impact of data quality variations, coding practices, and technological disparities on research outcomes. Third, it establishes evidence-based recommendations for optimizing the platform’s utilization in medical research.
Our review fills critical gaps in understanding real-world data platforms through several innovative approaches. We systematically analyze the methodological considerations unique to TriNetX-based research while evaluating the platform’s capacity for generating reliable and generalizable evidence. Additionally, we provide a framework for assessing and mitigating platform-specific limitations and developing standardized approaches for managing data quality variations and potential biases. Through this comprehensive analysis, we aim to enhance researchers’ ability to effectively utilize TriNetX while maintaining methodological rigor and ensuring the validity of their clinical research findings.
Methods:
Review Approach and Scope
We aimed to review the TriNetX platform’s capabilities, limitations, and implications for medical research. Our study focused on understanding the platform’s architecture, data quality considerations, and application in generating real-world evidence. We reviewed the platform’s features, technical specifications, and documented research applications to provide a thorough understanding of its utility and limitations in healthcare research.
Thematic Framework
The organization of this review follows a thematic structure derived from critical aspects of research platform evaluation. Our analysis framework emerged from examining key operational and methodological aspects of TriNetX, including data collection processes, platform architecture, and research applications. The primary themes were developed to address fundamental aspects of the platform: inherent limitations of retrospective studies, EMR coding accuracy, diagnostic coding precision, platform-specific data limitations, network representativeness, risk of bias, technological disparities, practice variations, and external validity considerations.
Analysis Structure
Our review begins with a review of the inherent limitations of retrospective observational studies, providing context for understanding TriNetX’s operational framework. We then progress through interconnected themes, analyzing EMR coding accuracy and its implications for data reliability. The review extends to examine specific coding impacts on TriNetX, platform data limitations, and network representativeness. We further explore the risk of bias considerations, technological disparities, and practice variations that affect research outcomes. The analysis concludes with an assessment of external validity and recommendations for optimal platform utilization.
Synthesis Approach
The review synthesizes and highlights the current understanding of TriNetX’s capabilities and limitations, drawing from published literature and platform documentation. We examine how various methodological aspects interact and influence research outcomes, providing a view of the platform’s utility in medical research. This synthesis aims to guide researchers in effectively utilizing TriNetX while maintaining awareness of its methodological constraints and opportunities for optimization.
Discussion
Inherent limitations of retrospective observational studies
Susceptibility to bias in retrospective studies is considerable due to the fact that exposure and outcome data have already occurred, and researchers cannot control participant selection, leading to potential selection bias [18,19]. Additionally, the inability to randomize subjects leaves these studies vulnerable to confounding factors that may obscure the true relationship between variables, complicating causal inferences [20,21]. Data quality is another critical issue, as the information is often collected for purposes other than research, like medical billing, resulting in data that may be incomplete, inconsistent, or inaccurate, thus affecting the reliability of the study [22,23]. Recalling bias further complicates matters, as patient-reported outcomes or historical records may be skewed by imperfect recollection or subjective interpretation [24,25]. Temporal ambiguity in these studies often makes it challenging to establish a clear sequence between exposure and outcome, thereby hindering causality determination [26,27]. Since retrospective studies depend on existing data, researchers have limited control over variables and cannot adjust for all relevant factors or unforeseen confounders [28,19]. The representativeness of the study population may also not mirror the general population, limiting the generalizability of the findings [29,30]. Despite the minimal harm to participants from using existing data, ethical considerations such as data privacy, consent, and the use of sensitive information still pose significant concerns [31,32].
Accuracy of EMR coding
Issues surrounding the accuracy of EMR coding are multifaceted, reflecting significant challenges across various aspects of the coding process. Studies like that by Horsky et al. (2017) have demonstrated considerable variability in coding, often resulting in omissions or incorrect entries, with only half of the entered codes for specific diagnoses being appropriate and a high omission rate for secondary diagnoses such as nicotine dependence or dialysis dependence [33]. Automated coding tools, such as the TOLBIAC control for ICD-10 codes, show moderate accuracy (micro-average F-measure of 0.76 for drug prescriptions) but require enhancements in text analysis capabilities [34]. Reliance on billing codes can lead to substantial false positives, as evidenced by a study identifying lung cancer patients with an ICD-based method achieving only 65% precision [35]. Comparatively, automated systems have proven completer and more accurate than manual coding, highlighted by their superior performance in coding PQRI quality measures [36]. Accuracy also varies between specialists and generalists, with problem lists in EHRs being more accurate when managed by primary care providers than specialists [37]. Phenotyping challenges are significant, such as chronic rhinosinusitis, where reliance solely on billing codes results in low precision and necessitates additional clinical validation [38]. Administrative coding errors are prevalent, too; an audit of emergency medical admissions revealed changes in the primary diagnosis in over 16% of cases, impacting outcome metrics like morbidity indices [39]. The quality of ICD-10 coding heavily depends on the experience of the coders, as shown in initiatives to improve morbidity and mortality data accuracy in Lagos hospitals [40]. AI-based methods like adversarial autoencoders have been proposed to address missing or incomplete clinical codes, demonstrating improvements in predictive performance [41]. Furthermore, big data concerns, such as inconsistencies in administrative coding for conditions like hospital-acquired venous thromboembolism, underscore the need for more robust diagnostic validation methods [42].
Errors in diagnostic coding
The impact of diagnostic coding errors on research outcomes is profound and multifaceted. Schrodi (2017) highlighted that errors in diagnostic coding could lead to misclassification of patient conditions, resulting in biased results in epidemiological and clinical studies, particularly diminishing the statistical power in genetic association studies, which underscores the critical need for accurate diagnostic code utilization [43]. Farzandipour and Sheikhtaheri (2009), through a systematic review, revealed that inaccuracies in ICD-10 codes significantly affect the validity of diagnoses in health databases, impacting both research findings and administrative decisions [44]. Usher et al. (2018) found that diagnostic coding errors during inter-hospital transfers are associated with increased inpatient mortality, emphasizing the necessity for robust health information exchange systems to enhance diagnostic accuracy [45]. Zafirah et al. (2018) noted that coding inaccuracies could lead to significant financial consequences by misassigning Diagnosis-Related Groups (DRGs), resulting in substantial revenue losses for hospitals [46]. Lorence and Ibrahim (2003) observed that studies using unvalidated diagnostic codes often lack reliability, with up to 87% of medical records in some datasets containing errors significant enough to alter research findings [47]. Diagnostic errors also complicate epidemiological research, as they can misrepresent disease prevalence and associated outcomes, necessitating validation studies to ascertain true associations [48]. Furthermore, Nouraei et al. (2015) discussed how variability in coding accuracy across healthcare facilities complicates the benchmarking of patient outcomes, affecting the comparability and utility of large datasets for outcomes management [39]. Table 1 reviewed the classification of diagnostic coding errors commonly encountered in EMR systems, their implications for research outcomes, and possible mitigation strategies.
| Error Category | Description | Research Implications | Mitigation Strategies |
|---|---|---|---|
| Diagnostic Omission | Missing secondary diagnoses or comorbidities documentation | Underestimation of disease prevalence and comorbidity burden | Implementation of comprehensive validation protocols for secondary diagnoses |
| Code Misclassification | Incorrect assignment of diagnostic codes or procedure codes | Biased estimation of treatment effects and outcome measures | Development of standardized coding protocols and automated validation checks |
| Temporal Inconsistency | Incorrect chronological ordering of diagnoses and treatments | Compromised analysis of treatment patterns and outcomes | Implementation of temporal validation algorithms |
| Documentation Variability | Inconsistent coding practices across healthcare providers | Reduced reliability in multi-center comparisons | Standardization of documentation practices across network institutions |
| Billing-Driven Coding | Coding optimized for reimbursement rather than clinical accuracy | Potential overestimation of certain conditions | Integration of clinical validation alongside billing codes |
| System Migration Errors | Data inconsistencies from ICD-9 to ICD-10 transition | Disrupted longitudinal analysis and trend assessment | Development of robust code mapping protocols |
| Incomplete Documentation | Missing or partial clinical information | Limited ability to adjust for confounders | Implementation of mandatory documentation fields |
| Granularity Loss | Use of non-specific codes when specific ones exist | Reduced precision in phenotype definitions | Training programs for accurate code selection |
Coding Impact Analysis on TriNetX
Analyzing EMR and ICD-10 coding issues affecting TriNetX reveals several key challenges. Palestine et al. (2018) noted imprecision in how ICD-10 codes are applied across different EMR systems, leading to inconsistencies in large-scale data analysis, particularly highlighting disparities in coding diseases like uveitis, which potentially introduces bias into research outcomes derived from pooled EHR data [49]. Kortüm et al. (2016) discussed the increased complexity post-EHR implementation, where the introduction of EMR systems with ICD-10 coding led to increased diagnosis diversity and exposed suboptimal coding precision, affecting the accuracy and utility of clinical research databases like TriNetX. Stewart et al. (2019) found that transitioning from ICD-9 to ICD-10 disrupted the recording of specific mental health disorders due to mismatches in diagnostic coding frameworks, which could impact longitudinal studies utilizing databases such as TriNetX where such transitions lead to data inconsistencies [50]. Quan et al. (2005) highlighted challenges in comorbidity analysis, as coding algorithms for defining comorbidities in ICD-10 differ in their accuracy compared to previous ICD-9 models, influencing the quality of comorbidity data within databases like TriNetX and impacting their predictive utility [51]. Caskey et al. (2013) pointed out that misclassification and information loss are common when mapping ICD-9 codes to ICD-10, with up to 26% of pediatric diagnosis codes categorized as convoluted, posing challenges to the reliability of TriNetX data in pediatric research [52]. Additionally, Horsky et al. (2017) emphasized that high variation in coding precision across institutions due to human and system factors leads to inconsistencies in databases, complicating the interpretation of multicenter data like those aggregated in TriNetX [33].
Data Limitation in the Platform
One major limitation is the restricted data scope, as critical clinical and demographic details such as treatment adherence, lifestyle factors, and social determinants of health are often not captured. These omissions can lead to incomplete analyses and potential biases in understanding patient outcomes. Additionally, the relatively short observation periods inherent in the database further constrain its utility for studying long-term outcomes. Patients who leave the institution, relocate, or otherwise fall out of the network are no longer tracked, leading to gaps in longitudinal data and potential underestimation of adverse events or delayed outcomes. These limitations necessitate cautious interpretation of findings and, when possible, supplementation with external data sources to ensure more robust and generalizable conclusions.
Network Representativeness
Considerations of the representativeness of the TriNetX network reveal several limitations affecting the breadth and validity of its research outcomes. Topaloglu and Palchuk (2018) noted that although TriNetX connects healthcare organizations worldwide, its representativeness is primarily influenced by the geographic and demographic composition of its network, initially dominated by patients from the United States with scant representation from low- and middle-income countries, restricting the generalizability of research findings to global populations [4]. Despite significant growth in international representation, as TriNetX expanded from 55 healthcare organizations in seven countries in 2017 to over 220 organizations in 30 countries by 2022, disparities in data representation persist, impacting the validity of cross-regional comparisons [1]. Furthermore, including healthcare organizations with specific specialties or academic affiliations can skew the dataset towards populations frequently treated in those settings, such as cancer centers, limiting generalizability for broader disease populations [4]. González et al. (2020) highlighted that variability in data harmonization across institutions might lead to inconsistent quality, impacting the representativeness and reliability of aggregated data for multicenter research [53]. Moreover, Haudenschild et al. (2021) pointed out that the self-selection of participating healthcare organizations and reliance on EMR for data may introduce biases that disproportionately represent urban and technologically advanced healthcare systems, raising concerns about selection bias [54].
Risk of Bias
An in-depth examination of selection bias and unaccounted confounding factors presents significant challenges across various research designs. Cohort studies may show selection bias when inclusion criteria are linked to both exposure and outcomes, skewing results as higher hazard rates during early cohort inclusion periods suggest significant bias [55]. In Mendelian randomization, bias can arise from colliders—effects influenced by both genetic variants and confounders—with large selection effects amplifying this bias in simulations [56]. Unaccounted confounders in observational research can obscure causal relationships when variables influence both exposure and outcome [57]. Notably, selection bias impacts a study’s external validity, while confounding affects internal validity, which is crucial for comparative effectiveness research [58]. Missing data related to exposure or outcomes can also introduce bias, mitigated by techniques such as inverse probability weighting [59]. In environmental studies, selection bias in case-crossover designs occurs when reference periods don’t accurately represent exposure times, exacerbating bias [60]. Conditioning on colliders can create spurious associations in observational research [61], and in longitudinal studies, loss to follow-up can lead to bias if dropouts differ systematically from those who remain, with causal diagrams aiding in bias mitigation [62].
Technological Disparities
A discussion of technological disparities among hospitals reveals significant variations in access and utilization based on racial, ethnic, and regional differences, which in turn impact health outcomes. Kim et al. (2010) found that racial and ethnic minorities, such as Hispanic patients, are significantly less likely to utilize hospitals with advanced technological services like MRI and trauma units compared to white patients, underscoring systemic inequities [63]. Groeneveld et al. (2005) reported that hospitals serving higher proportions of minority populations often lag in adopting new medical technologies, contributing to racial gaps in access to advanced procedures such as coronary artery bypass grafting [64]. Sequist (2011) noted that while Health Information Technology (HIT) has the potential to reduce disparities, it is less accessible to safety-net hospitals, perpetuating inequities in the quality of care among underserved populations [65]. Economic constraints also play a critical role; according to David & Jahnke (2018), smaller or rural hospitals often face financial barriers to investing in new technologies, leading to disparities in equipment availability and maintenance compared to urban institutions [66]. Newton et al. (2010) highlighted that hospitals in low-resource settings struggle with challenges such as poorly trained staff and inadequate infrastructure, further widening the technological divide and compromising patient safety [67]. Lastly, Walker et al. (2020) observed disparities in adopting digital health tools like patient portals, with older and African American patients utilizing these technologies less frequently, indicating a digital divide even within technologically equipped hospitals [68].
Practice Disparities
Healthcare disparities and practices significantly limit the generalizability and applicability of research outcomes. Chin et al. (2012) identified that disparities in healthcare access and quality create heterogeneity in patient populations, leading to research outcomes that may not be generalizable across diverse groups, particularly when studies fail to account for racial and ethnic disparities [69]. Kilbourne et al. (2006) noted that research often focuses on settings with adequate resources, excluding under-resourced healthcare facilities and their patient populations, thus introducing selection bias and restricting the relevance of findings to broader, diverse settings [70]. Rust & Cooper (2007) highlighted that studies neglecting socioeconomic determinants of health, such as poverty or education, fail to provide actionable insights for reducing health disparities, further limiting the practical applicability of research[71]. Chinman et al. (2017) pointed out that research focusing narrowly on single interventions without considering multilevel determinants often falls short of addressing complex healthcare disparities, limiting the applicability of findings in real-world scenarios [72]. Koh et al. (2010) argued that research not tailored to cultural or linguistic contexts may result in ineffective interventions when applied to populations with specific needs, such as non-English speakers [73]. Lane et al. (2016) discussed how disparities in healthcare practices, such as differences in provider training and resource availability, complicate the translation of research into practice, reducing the impact of evidence-based interventions [74]. Dankwa-Mullan et al. (2010) observed that studies conducted in resource-rich environments often overlook the challenges faced by underfunded facilities, making their recommendations impractical for these settings [75]. Finally, Baker et al. (2001) emphasized that research lacking community engagement fails to address local needs and contexts, reducing the effectiveness of interventions when applied in diverse settings [76].
External Validity Critique
Critiquing the external validity and applicability of TriNetX study findings highlights several limitations in generalizing results to the general population. Stuart et al. (2015) noted that studies utilizing datasets might face challenges in population representativeness due to variability in patient demographics and institutional contributions, often not reflecting real-world diversity, which can limit generalizability[77]. Murad et al. (2018) found that the generalizability of findings is further complicated by differences in coding practices, healthcare delivery models, and patient characteristics among contributing healthcare organizations, leading to inconsistent applicability of results [78]. Kennedy-Martin et al. (2015) pointed out that many datasets exclude elderly patients or those with multiple comorbidities, which restricts the ability to generalize results to these high-need groups, affecting external validity [78]. Pearl & Bareinboim (2014) discussed how the inherent structure of datasets, relying on real-world data from diverse settings, can introduce selection bias and complicate causal interpretations, further impacting generalizability [79]. Leviton (2017) highlighted that interventions tested in highly specialized or resource-intensive institutions contributing to large databases may not be practical for low-resource or rural healthcare settings, limiting the utility of such findings [80]. Ling et al. (2023) mentioned that applying study results to populations not represented in the original datasets requires advanced statistical adjustments, with the feasibility and accuracy of such transportability efforts being an area of ongoing research [81].
TriNetX Strengths
Despite its limitations, TriNetX offers significant strengths that enhance research capabilities across various medical and healthcare fields. Topaloglu & Palchuk (2018) highlighted that TriNetX aggregates extensive real-world patient data from diverse healthcare institutions, enabling large-scale observational studies that provide insights beyond what is possible with randomized controlled trials[4]. The TriNetx platform facilitates rapid cohort identification, allowing researchers to define specific criteria and quickly identify patient cohorts, thereby streamlining the design and initiation of studies and accelerating evidence generation. Murad et al. (2018) emphasized that by connecting academic institutions, healthcare providers, and pharmaceutical companies, TriNetX fosters collaboration, advancing research across multidisciplinary teams [78]. Kennedy-Martin et al. (2015) pointed out that big database plays a critical role in generating real-world evidence for comparative effectiveness studies, health economics research, and post-marketing surveillance, complementing traditional trial data [82]. Leviton (2017) observed that it provides tools for propensity score matching and other statistical methods to minimize bias in observational research, enhancing the reliability of findings [80]. Ling et al. (2023) added that the platform’s data is regularly updated, offering near-real-time insights, which is particularly valuable for monitoring emerging health trends or disease outbreaks [81]Lastly, with a growing network of international contributors, TriNetX supports multinational studies, which are crucial for addressing global health challenges and understanding regional variations in care. This underscores its scalability for multinational research.
Balanced Insights
Recognizing both the strengths and limitations of platforms like TriNetX is crucial for guiding future research effectively. Topaloglu & Palchuk (2018) emphasize that acknowledging the real-world data scope of TriNetX can guide the development of more representative and practical study designs, leveraging its capabilities to include diverse populations and real-world conditions [4]. Murad et al. (2018) point out that recognizing issues such as selection bias and confounding factors is essential for implementing robust methodologies to mitigate these issues, thus enhancing the reliability of findings [78]. Understanding both strengths and limitations helps policymakers evaluate the applicability of findings for healthcare interventions, ensuring decisions are grounded in balanced evidence [83,10]. Leviton (2017) highlights that the platform’s ability to produce real-world evidence is crucial for informing guidelines and improving clinical practice, supporting its integration into translational research [80]. Ling et al. (2023) state that recognizing advanced tools for data analysis encourages the continuous improvement and development of methodologies, such as AI-driven analytics, to maximize the platform’s utility [81]. Pearl & Bareinboim (2014) argue that awareness of both strengths and ethical challenges, such as data privacy, ensures responsible use and compliance with regulations, protecting patient rights while advancing research [79]. Finally, Chin et al. (2012) discuss how balancing the platform’s capabilities and limitations helps prioritize research areas where its strengths are most applicable, such as in rare disease studies or post-market drug surveillance [69].
Database Comparisons
Comparing TriNetX with other databases and research methods reveals both unique strengths and shared challenges. Topaloglu & Palchuk (2018) explain that, unlike centralized databases like SEER (Surveillance, Epidemiology, and End Results), TriNetX operates through a federated network model, facilitating multicenter data analysis while preserving data ownership at contributing sites, which reduces privacy risks but can introduce heterogeneity in data quality [4]. González et al. (2020) contrast TriNetX with the I2B2 platform, which requires more intensive efforts for semantic normalization and data integration, whereas TriNetX’s daily data refresh supports timelier research opportunities [53]. Singh et al. (2021) highlight that TriNetX supports enhanced clinical trial design by linking EMRs with pharmaceutical data, unlike systems like NSQIP (National Surgical Quality Improvement Program), which focus narrowly on surgical outcomes [3]. Hernandez et al. (2022) point out that TriNetX integrates genomic data using FHIR standards, enabling advanced pharmacogenomics and personalized medicine studies—a strength not commonly found in competing platforms [84]. Finally, Evans et al. (2021) observe that while TriNetX facilitates efficient multicenter collaborations, other federated systems like PCORnet face interoperability challenges due to variations in EMR platforms, highlighting different operational dynamics [22].
Limitation Mitigation
Minimizing the impact of limitations in future studies using TriNetX involves several strategic approaches. Topaloglu & Palchuk (2018) suggest standardizing data collection practices to ensure consistency in data entry and harmonization across participating healthcare organizations, which can reduce variability and improve dataset reliability [4]. Evans et al. (2021) recommend employing validation tools and monitoring data completeness, such as analyzing diagnosis-medication couplets to ensure robust medication data to enhance data quality [22]. Evans et al. (2023) emphasize developing methodologies to detect and correct date-shifting practices in real-world datasets to ensure temporal accuracy and validity of clinical event timelines [85]. Antoine et al. (2023) advocates using advanced statistical methods, such as stabilized inverse-probability weighting and G-computation, to mitigate biases from confounding variables, thus improving causal inference [86]. Palchuk et al. (2023) highlights the importance of incorporating multilevel and longitudinal data analysis to better account for regional and temporal differences in healthcare practices, improving the generalizability of study findings [10]. Furler et al. (2012) note that expanding the TriNetX network to include more healthcare organizations from underserved and rural areas can improve the representativeness of its datasets[87]. Monjas et al. (2023) propose implementing automated validation and quality control mechanisms at the point of data entry to minimize errors and inconsistencies, enhancing the accuracy of datasets [88]. Lastly, Antoine et al. (2023) discussed the advantages of leveraging synthetic control arms in real-world evidence studies to enhance comparative effectiveness research by reducing reliance on incomplete or biased datasets [86].
Databases Comparison
Unlike the Surveillance, Epidemiology, and End Results (SEER) database, which maintains a centralized repository focused primarily on cancer-specific outcomes, TriNetX employs a federated network model that preserves data ownership at contributing sites while facilitating multicenter analysis. This architectural difference significantly impacts data access patterns and privacy management, with TriNetX offering more dynamic data updates compared to SEER’s annual reporting cycle [89]. Compared to PCORnet (Patient-Centered Clinical Research Network), TriNetX demonstrates different data standardization and interoperability approaches. While PCORnet employs a common data model requiring extensive local data transformation, TriNetX’s approach to data harmonization allows for more rapid implementation across new sites. However, this efficiency comes with potential trade-offs in data granularity. PCORnet’s more rigorous standardization process may provide better consistency in certain research applications, particularly in longitudinal studies requiring detailed clinical parameters [90,91]. The i2b2 (Informatics for Integrating Biology and the Bedside) platform presents another interesting comparison. While i2b2 excels in providing detailed phenotypic data and supports sophisticated query building, it requires more intensive efforts for semantic normalization and data integration compared to TriNetX’s streamlined approach. TriNetX’s daily data refresh capability offers advantages for time-sensitive research, though i2b2’s more granular data model may be preferable for certain types of clinical research [92]. TriNetX has demonstrated superior scalability capabilities in managing large-scale, multi-institutional studies compared to traditional research databases. Its integration of genomic data using FHIR standards enables advanced pharmacogenomics and personalized medicine studies, a functionality not commonly found in competing platforms [93,22,94,84,95,10,4]. However, this advantage must be weighed against the potential for data quality variations across participating institutions. When reviewing data quality metrics, each platform presents distinct strengths and limitations. SEER’s rigorous data collection protocols ensure high-quality cancer-specific data but limit its utility for broader healthcare research. PCORnet’s emphasis on patient-centered outcomes provides rich patient-reported data but may face challenges in standardizing information across diverse healthcare settings. TriNetX’s approach to data quality focuses on rapid accessibility and broad coverage, though this may sometimes come at the expense of granular clinical detail [22,84,10,4]. Usability comparisons reveal that TriNetX generally offers more intuitive interfaces for cohort identification and study design compared to traditional research databases. Its real-time query capabilities and integrated analytical tools provide rapid hypothesis testing and study feasibility assessment advantages. However, platforms like i2b2 may offer more flexibility for complex phenotype definitions, albeit with a steeper learning curve [92].
Conclusions
Evaluating TriNetX as a research platform reveals significant strengths and notable limitations that researchers must carefully consider. While TriNetX offers unprecedented access to large-scale, real-world patient data and facilitates rapid cohort identification across multiple healthcare organizations, several critical limitations warrant attention. The platform’s reliance on EHR coding introduces potential inaccuracies, with studies indicating considerable variability in coding precision and completeness across institutions. Selection bias remains a significant concern, as the network predominantly represents insured patients from academic and acute care settings, potentially limiting generalizability to broader populations. These limitations are compounded by technological and practice disparities among participating healthcare organizations, which can affect data quality and representation. The transition between coding systems (ICD-9 to ICD-10) has introduced additional complexity, potentially impacting longitudinal studies and data consistency. However, TriNetX’s strengths in facilitating large-scale observational studies and supporting advanced statistical methods like propensity score matching partially mitigate these challenges. Future developments should focus on standardizing data collection practices, expanding network representation to include more diverse healthcare settings, and implementing robust validation tools to enhance the platform’s utility. Researchers should employ advanced statistical methods to address confounding variables and selection bias while clearly acknowledging study limitations in their findings. Integrating automated validation mechanisms and quality control at the point of data entry could significantly improve data accuracy. Despite its limitations, TriNetX remains a valuable tool for medical research, particularly in areas such as rare disease studies, post-market surveillance, and comparative effectiveness research. Success in utilizing the platform requires a balanced approach that leverages its strengths while actively addressing its limitations through appropriate methodological choices and careful interpretation of results. Future enhancements should prioritize expanding network diversity, improve data validation mechanisms, and developing more sophisticated tools for bias mitigation.
Conflicts of Interest
The authors declare no competing interests that could have influenced the objectivity or outcome of this research.
Funding Source
The National Center for Advancing Translational Sciences (NCATS), National Institutes of Health, supported the project described through CTSA award number UM1TR004400. The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIH.
Acknowledgments
None
Institutional Review Board (IRB)
This review analyzed only previously published literature and aggregate, de-identified data from a commercial federated research network. No human-subjects research was conducted and no institutional review board approval was required.
Large Language Model
The manuscript was language-edited using an LLM strictly to refine clarity, grammar, and readability. No new content was created or collected during this process, ensuring the original scientific content remained unchanged.
Authors Contribution
MN and AA conceptualized the idea; MN, HA, KE, AB, AI, MM, AM, ME, DA, and AA equally contributed by reviewing, editing, performing data analysis, and refining the manuscript.
Data Availability
This review article contains no new primary data. All information discussed is derived from previously published sources and publicly available databases, as cited in the manuscript.
References
- Palchuk M. B., London J. W., Perez-Rey D., Drebert Z. J., Winer-Jones J. P., Thompson C. N., Esposito J., Claerhout B.. A global federated real-world data and analytics platform for research. JAMIA Open. 2023;6(2):ooad035. doi:10.1093/jamiaopen/ooad035 PMID: 37193038 PMCID: PMC10182857
- Sáez Marín Adolfo Jesús, Hernández-Ibarburu Gema, Alonso Fernández Rafael, Sanchez Pina Jose Maria, Lázaro Del Campo Paula, Jimenez Ubieto Ana, Cuellar Clara, Tamayo Soto Andrea, Mas Babio Reyes, Medina Lucia, Calbacho Maria, Ayala Rosa, Martinez Lopez Joaquin. Large Scale Analysis of Autologous Stem Cell Transplantation for Multiple Myeloma Patients Older Than 65 Years. Blood. 2023;142(Supplement 1):7064-7064. doi:10.1182/blood-2023-180923
- Singh D., Slavin B. R., Holton T.. Comparing Surgical Site Occurrences in 1 versus 2-stage Breast Reconstruction via Federated EMR Network. Plast Reconstr Surg Glob Open. 2021;9(1):e3385. doi:10.1097/GOX.0000000000003385 PMID: 33564597 PMCID: PMC7862101
- Topaloglu U., Palchuk M. B.. Using a Federated Network of Real-World Data to Optimize Clinical Trials Operations. JCO Clin Cancer Inform. 2018;2:1-10. doi:10.1200/CCI.17.00067 PMID: 30652541 PMCID: PMC6816049
- Badhiwala J. H., Karmur B. S., Wilson J. R.. Propensity Score Matching: A Powerful Tool for Analyzing Observational Nonrandomized Data. Clin Spine Surg. 2021;34(1):22-24. doi:10.1097/BSD.0000000000001055 PMID: 32804684
- Kane L. T., Fang T., Galetta M. S., Goyal D. K. C., Nicholson K. J., Kepler C. K., Vaccaro A. R., Schroeder G. D.. Propensity Score Matching: A Statistical Method. Clin Spine Surg. 2020;33(3):120-122. doi:10.1097/BSD.0000000000000932 PMID: 31913173
- Prasad A., Shin M., Carey R. M., Chorath K., Parhar H., Appel S., Moreira A., Rajasekaran K.. Propensity score matching in otolaryngologic literature: A systematic review and critical appraisal. PLoS One. 2020;15(12):e0244423. doi:10.1371/journal.pone.0244423 PMID: 33382777 PMCID: PMC7774981
- Reiffel J. A.. Propensity Score Matching: The 'Devil is in the Details' Where More May Be Hidden than You Know. Am J Med. 2020;133(2):178-181. doi:10.1016/j.amjmed.2019.08.055 PMID: 31618617
- London Jack W., Orourke Julia, Warnick Jeff, Doole John, De Keyser Luc, Drebert Zuzanna, Wan Olivia, Thompson Courtney, Palchuk Matvey. Deriving breast cancer chemotherapy patterns from real-world data. Journal of Clinical Oncology. 2023;41(16_suppl):e13586-e13586. doi:10.1200/JCO.2023.41.16_suppl.e13586
- Mohamed Y., Song X., McMahon T. M., Sahil S., Zozus M., Wang Z., Greater Plains Collaborative, Waitman L. R.. Electronic health record data quality variability across a multistate clinical research network. J Clin Transl Sci. 2023;7(1):e130. doi:10.1017/cts.2023.548 PMID: 37396818 PMCID: PMC10308424
- Sandhu Namneet, Whittle Sarah, Varela Lucia Otero, Eastwood Cathy A., Southern Danielle A., Quan Hude. Assessing Hospital Data Quality: Application of a Data Quality Tool in 15 Countries. International Journal of Population Data Science. 2024;9(5). doi:10.23889/ijpds.v9i5.2889
- Langworthy B., Wu Y., Wang M.. An overview of propensity score matching methods for clustered data. Stat Methods Med Res. 2023;32(4):641-655. doi:10.1177/09622802221133556 PMID: 36426585 PMCID: PMC10119899
- Segalas C., Leyrat C., Carpenter J. R., Williamson E.. Propensity score matching after multiple imputation when a confounder has missing data. Stat Med. 2023;42(7):1082-1095. doi:10.1002/sim.9658 PMID: 36695043 PMCID: PMC10946973
- Anand P., Zhang Y., Merola D., Jin Y., Wang S. V., Lii J., Liu J., Lin K. J.. Comparison of EHR Data-Completeness in Patients with Different Types of Medical Insurance Coverage in the United States. Clin Pharmacol Ther. 2023;114(5):1116-1125. doi:10.1002/cpt.3027 PMID: 37597260 PMCID: PMC10919452
- Wilson J. L., Betensky M., Udassi S., Ellison P. R., Lilienthal R., Stahl L. R., Palchuk M. B., Zia A., Town D. A., Kimble W., Goldenberg N. A., Morizono H.. Leveraging a global, federated, real-world data network to optimize investigator-initiated pediatric clinical trials: the TriNetX Pediatric Collaboratory Network. JAMIA Open. 2024;7(3):ooae077. doi:10.1093/jamiaopen/ooae077 PMID: 39224867 PMCID: PMC11368118
- Onesimu J. Andrew, J Karthikeyan, Eunice Jennifer, Pomplun Marc, Dang Hien. Privacy Preserving Attribute-Focused Anonymization Scheme for Healthcare Data Publishing. IEEE Access. 2022;10:86979-86997. doi:10.1109/access.2022.3199433
- Zuo Z., Watson M., Budgen D., Hall R., Kennelly C., Al Moubayed N.. Data Anonymization for Pervasive Health Care: Systematic Literature Mapping Study. JMIR Med Inform. 2021;9(10):e29871. doi:10.2196/29871 PMID: 34652278 PMCID: PMC8556642
- Goldstein N. D., Kahal D., Testa K., Burstyn I.. Inverse probability weighting for selection bias in a Delaware community health center electronic medical record study of community deprivation and hepatitis C prevalence. Ann Epidemiol. 2021;60:1-7. doi:10.1016/j.annepidem.2021.04.011 PMID: 33933628 PMCID: PMC8355055
- Talari K., Goyal M.. Retrospective studies - utility and caveats. J R Coll Physicians Edinb. 2020;50(4):398-402. doi:10.4997/JRCPE.2020.409 PMID: 33469615
- D'Onofrio B. M., Sjolander A., Lahey B. B., Lichtenstein P., Oberg A. S.. Accounting for Confounding in Observational Studies. Annu Rev Clin Psychol. 2020;16:25-48. doi:10.1146/annurev-clinpsy-032816-045030 PMID: 32384000
- Verbeek J. H., Whaley P., Morgan R. L., Taylor K. W., Rooney A. A., Schwingshackl L., Hoving J. L., Vittal Katikireddi S., Shea B., Mustafa R. A., Murad M. H., Schunemann H. J., GRADE Working Group. An approach to quantifying the potential importance of residual confounding in systematic reviews of observational studies: A GRADE concept paper. Environ Int. 2021;157:106868. doi:10.1016/j.envint.2021.106868 PMID: 34530289
- Evans L., London J. W., Palchuk M. B.. Assessing real-world medication data completeness. J Biomed Inform. 2021;119:103847. doi:10.1016/j.jbi.2021.103847 PMID: 34161824
- Mashoufi M., Ayatollahi H., Khorasani-Zavareh D., Talebi Azad Boni T.. Data Quality in Health Care: Main Concepts and Assessment Methodologies. Methods Inf Med. 2023;62(1-02):5-18. doi:10.1055/s-0043-1761500 PMID: 36716776
- Crutchfield C. R., Givens R. R., O'Connor M., deMeireles A. J., Lynch T. S.. Recall Bias in the Retrospective Collection of Common Patient-Reported Outcome Scores in Hip Arthroscopy. Am J Sports Med. 2022;50(12):3190-3197. doi:10.1177/03635465221118375 PMID: 35993555
- Zini M. L. L., Banfi G.. A Narrative Literature Review of Bias in Collecting Patient Reported Outcomes Measures (PROMs). Int J Environ Res Public Health. 2021;18(23). doi:10.3390/ijerph182312445 PMID: 34886170 PMCID: PMC8656687
- Valentin Simon, Bramley Neil, Lucas Christopher. Learning Hidden Causal Structure from Temporal Data. Cognitive Science. 2020:1906-1912.
- Zhao Helen Hailin, Shipp Abbie J., Carter Kameron, Gonzalez‐Mulé Erik, Xu Erica. Time and change: A meta‐analysis of temporal decisions in longitudinal studies. Journal of Organizational Behavior. 2024;45(4):620-640. doi:10.1002/job.2771
- Kinter S., Delaney J. A., Susarla S., McKinney C.. Retrospective Cohort Studies in Craniofacial Outcomes Research: An Epidemiologist's Approach to Mitigating Bias. Cleft Palate Craniofac J. 2025;62(6):1061-1067. doi:10.1177/10556656241233234 PMID: 38389276
- Degtiar Irina, Rose Sherri. A Review of Generalizability and Transportability. Annual Review of Statistics and Its Application. 2023;10(1):501-524. doi:10.1146/annurev-statistics-042522-103837
- Thomas Darren S., Collin Simon, Berrocal-Almanza Luis C., Stirnadel-Farrant Heide, Zhang Yiduo, Sun Ping. Extending Inferences From Sample To Target Populations: On The Generalizability Of A Real-World Clinico-Genomic Database Non-Small Cell Lung Cancer Cohort. medRxiv. 2025. doi:10.1101/2023.06.15.23291372
- Latha Narayanan Valli, Sujatha N., Mukul Mech, Lokesh V. S.. Ethical considerations in data science: Balancing privacy and utility. International Journal of Science and Research Archive. 2024;11(1):011-022. doi:10.30574/ijsra.2024.11.1.1098
- Pina Eduardo, Ramos José, Jorge Henrique, Váz Paulo, Silva José, Wanzeller Cristina, Abbasi Maryam, Martins Pedro. Data Privacy and Ethical Considerations in Database Management. Journal of Cybersecurity and Privacy. 2024;4(3):494-517. doi:10.3390/jcp4030024
- Horsky J., Drucker E. A., Ramelson H. Z.. Accuracy and Completeness of Clinical Coding Using ICD-10 for Ambulatory Visits. AMIA Annu Symp Proc. 2017;2017:912-920. PMID: 29854158 PMCID: PMC5977598
- Chaux R., Treussier I., Audeh B., Pereira S., Hengoat T., Paviot B. T., Bousquet C.. Automated Control of Codes Accuracy in Case-Mix Databases by Evaluating Coherence with Available Information in the Electronic Health Record. Stud Health Technol Inform. 2019;264:551-555. doi:10.3233/SHTI190283 PMID: 31437984
- Yu Yue, Ruddy Kathryn Jean, Leventakos Konstantinos, Liu Bolun, Huo Nan, Pachman Deirdre R., Zong Nansu, Xiao Guohui, Chute Christopher, Pfaff Emily, Cheville Andrea L., Jiang Guoqian. Using EHR data and machine learning approach to facilitate the identification of patients with lung cancer from a pan-cancer cohort. Journal of Clinical Oncology. 2023;41(16_suppl):e13552-e13552. doi:10.1200/JCO.2023.41.16_suppl.e13552
- McColm D., Karcz A.. Comparing manual and automated coding of physicians quality reporting initiative measures in an ambulatory EHR. J Med Pract Manage. 2010;26(1):6-12. PMID: 20839502
- Luna D., Franco M., Plaza C., Otero C., Wassermann S., Gambarte M. L., Giunta D., Gonzalez Bernaldo de Quiros F.. Accuracy of an electronic problem list from primary care providers and specialists. Stud Health Technol Inform. 2013;192:417-21. doi:10.3233/978-1-61499-289-9-417 PMID: 23920588
- Hsu J., Pacheco J. A., Stevens W. W., Smith M. E., Avila P. C.. Accuracy of phenotyping chronic rhinosinusitis in the electronic health record. Am J Rhinol Allergy. 2014;28(2):140-4. doi:10.2500/ajra.2014.28.4012 PMID: 24717952 PMCID: PMC5517777
- Nouraei S. A., Hudovsky A., Frampton A. E., Mufti U., White N. B., Wathen C. G., Sandhu G. S., Darzi A.. A Study of Clinical Coding Accuracy in Surgery: Implications for the Use of Administrative Big Data for Outcomes Management. Ann Surg. 2015;261(6):1096-107. doi:10.1097/SLA.0000000000000851 PMID: 25470740
- Olagundoye O., van Boven K., Daramola O., Njoku K., Omosun A.. Improving the accuracy of ICD-10 coding of morbidity/mortality data through the introduction of an electronic diagnostic terminology tool at the general hospitals in Lagos, Nigeria. BMJ Open Qual. 2021;10(1). doi:10.1136/bmjoq-2020-000938 PMID: 33674344 PMCID: PMC7939013
- Yordanov T. R., Abu-Hanna A., Ravelli A. C. J., Vagliano I.. Autoencoder-Based Prediction of ICU Clinical Codes. Artificial Intelligence in Medicine, Aime 2023. 2023;13897:57-62. doi:10.1007/978-3-031-34344-5_8
- Pellathy T., Saul M., Clermont G., Dubrawski A. W., Pinsky M. R., Hravnak M.. Accuracy of identifying hospital acquired venous thromboembolism by administrative coding: implications for big data and machine learning research. J Clin Monit Comput. 2022;36(2):397-405. doi:10.1007/s10877-021-00664-6 PMID: 33558981 PMCID: PMC8349368
- Schrodi S. J.. The Impact of Diagnostic Code Misclassification on Optimizing the Experimental Design of Genetic Association Studies. J Healthc Eng. 2017;2017:7653071. doi:10.1155/2017/7653071 PMID: 29181145 PMCID: PMC5664372
- Farzandipour M., Sheikhtaheri A.. Accuracy of diagnostic coding based on ICD-10. KAUMS Journal. 2009;12:67-76.
- Usher M., Sahni N., Herrigel D., Simon G., Melton G. B., Joseph A., Olson A.. Diagnostic Discordance, Health Information Exchange, and Inter-Hospital Transfer Outcomes: a Population Study. J Gen Intern Med. 2018;33(9):1447-1453. doi:10.1007/s11606-018-4491-x PMID: 29845466 PMCID: PMC6109004
- Zafirah S. A., Nur A. M., Puteh S. E. W., Aljunid S. M.. Potential loss of revenue due to errors in clinical coding during the implementation of the Malaysia diagnosis related group (MY-DRG((R))) Casemix system in a teaching hospital in Malaysia. BMC Health Serv Res. 2018;18(1):38. doi:10.1186/s12913-018-2843-1 PMID: 29370785 PMCID: PMC5784726
- Lorence D. P., Ibrahim I. A.. Benchmarking variation in coding accuracy across the United States. J Health Care Finance. 2003;29(4):29-42. PMID: 12908652
- Re V.. Validation of health outcomes of interest in healthcare databases. Pragmatic Randomized Clinical Trials: Using Primary Data Collection and Electronic Health Records. 2021:207-218. doi:10.1016/B978-0-12-817663-4.00022-2
- Palestine A. G., Merrill P. T., Saleem S. M., Jabs D. A., Thorne J. E.. Assessing the Precision of ICD-10 Codes for Uveitis in 2 Electronic Health Record Systems. JAMA Ophthalmol. 2018;136(10):1186-1190. doi:10.1001/jamaophthalmol.2018.3001 PMID: 30054618 PMCID: PMC6583860
- Stewart C. C., Lu C. Y., Yoon T. K., Coleman K. J., Crawford P. M., Lakoma M. D., Simon G. E.. Impact of ICD-10-CM Transition on Mental Health Diagnoses Recording. EGEMS (Wash DC). 2019;7(1):14. doi:10.5334/egems.281 PMID: 31065557 PMCID: PMC6484373
- Quan H., Sundararajan V., Halfon P., Fong A., Burnand B., Luthi J. C., Saunders L. D., Beck C. A., Feasby T. E., Ghali W. A.. Coding algorithms for defining comorbidities in ICD-9-CM and ICD-10 administrative data. Med Care. 2005;43(11):1130-9. doi:10.1097/01.mlr.0000182534.19832.83 PMID: 16224307
- Caskey R., Zaman J., Nam H., Chae S. R., Williams L., Mathew G., Burton M., Li J. J., Lussier Y. A., Boyd A. D.. The transition to ICD-10-CM: challenges for pediatric practice. Pediatrics. 2014;134(1):31-6. doi:10.1542/peds.2013-4147 PMID: 24918217 PMCID: PMC4531279
- Gonzalez L., Perez-Rey D., Alonso E., Hernandez G., Serrano P., Pedrera M., Gomez A., De Schepper K., Crepain T., Claerhout B.. Building an I2B2-Based Population Repository for Clinical Research. Stud Health Technol Inform. 2020;270:78-82. doi:10.3233/SHTI200126 PMID: 32570350
- Haudenschild C., Vaickus L., Levy Joshua. Configuring a federated network of real-world patient health data for multimodal deep learning prediction of health outcomes. bioRxiv. 2021. doi:10.1145/3477314.3507007
- Torner A., Dickman P., Duberg A. S., Kristinsson S., Landgren O., Bjorkholm M., Svensson A.. A method to visualize and adjust for selection bias in prevalent cohort studies. Am J Epidemiol. 2011;174(8):969-76. doi:10.1093/aje/kwr211 PMID: 21920949
- Gkatzionis A., Burgess S.. Contextualizing selection bias in Mendelian randomization: how bad is it likely to be?. Int J Epidemiol. 2019;48(3):691-701. doi:10.1093/ije/dyy202 PMID: 30325422 PMCID: PMC6659463
- Grimes D. A., Schulz K. F.. Bias and causal associations in observational research. Lancet. 2002;359(9302):248-52. doi:10.1016/S0140-6736(02)07451-2 PMID: 11812579
- Haneuse S.. Distinguishing Selection Bias and Confounding Bias in Comparative Effectiveness Research. Med Care. 2016;54(4):e23-9. doi:10.1097/MLR.0000000000000011 PMID: 24309675 PMCID: PMC4043938
- Valeri L., Coull B. A.. Estimating causal contrasts involving intermediate variables in the presence of selection bias. Stat Med. 2016;35(26):4779-4793. doi:10.1002/sim.7025 PMID: 27411847
- Bateson T. F., Schwartz J.. Selection bias and confounding in case-crossover analyses of environmental time-series data. Epidemiology. 2001;12(6):654-61. doi:10.1097/00001648-200111000-00013 PMID: 11679793
- Infante-Rivard C., Cusson A.. Reflection on modern methods: selection bias-a review of recent developments. Int J Epidemiol. 2018;47(5):1714-1722. doi:10.1093/ije/dyy138 PMID: 29982600
- Howe C. J., Cole S. R., Lau B., Napravnik S., Eron J. J., Jr.. Selection Bias Due to Loss to Follow Up in Cohort Studies. Epidemiology. 2016;27(1):91-7. doi:10.1097/EDE.0000000000000409 PMID: 26484424 PMCID: PMC5008911
- Kim T. H., Samson L. F., Lu N.. Racial/ethnic disparities in the utilization of high-technology hospitals. J Natl Med Assoc. 2010;102(9):803-10. doi:10.1016/s0027-9684(15)30677-5 PMID: 20922924
- Groeneveld P. W., Laufer S. B., Garber A. M.. Technology diffusion, hospital variation, and racial disparities among elderly Medicare beneficiaries: 1989-2000. Med Care. 2005;43(4):320-9. doi:10.1097/01.mlr.0000156849.15166.ec PMID: 15778635
- Sequist T. D.. Health information technology and disparities in quality of care. J Gen Intern Med. 2011;26(10):1084-5. doi:10.1007/s11606-011-1812-8 PMID: 21809173 PMCID: PMC3181307
- David Yadin, Jahnke Ernest Gus. Planning Medical Technology Management in a Hospital. Global Clinical Engineering Journal. 2018:23-32. doi:10.31354/globalce.v0i1.23
- Newton R. C., Mytton O. T., Aggarwal R., Runciman W. B., Free M., Fahlgren B., Akiyama M., Farlow B., Yaron S., Locke G., Whittaker S.. Making existing technology safer in healthcare. Qual Saf Health Care. 2010;19 Suppl 2:i15-24. doi:10.1136/qshc.2009.038539 PMID: 20693212
- Walker D. M., Hefner J. L., Fareed N., Huerta T. R., McAlearney A. S.. Exploring the Digital Divide: Age and Race Disparities in Use of an Inpatient Portal. Telemed J E Health. 2020;26(5):603-613. doi:10.1089/tmj.2019.0065 PMID: 31313977 PMCID: PMC7476395
- Chin M. H., Clarke A. R., Nocon R. S., Casey A. A., Goddu A. P., Keesecker N. M., Cook S. C.. A roadmap and best practices for organizations to reduce racial and ethnic disparities in health care. J Gen Intern Med. 2012;27(8):992-1000. doi:10.1007/s11606-012-2082-9 PMID: 22798211 PMCID: PMC3403142
- Kilbourne A. M., Switzer G., Hyman K., Crowley-Matoka M., Fine M. J.. Advancing health disparities research within the health care system: a conceptual framework. Am J Public Health. 2006;96(12):2113-21. doi:10.2105/AJPH.2005.077628 PMID: 17077411 PMCID: PMC1698151
- Rust G., Cooper L. A.. How can practice-based research contribute to the elimination of health disparities?. J Am Board Fam Med. 2007;20(2):105-14. doi:10.3122/jabfm.2007.02.060131 PMID: 17341746
- Chinman M., Woodward E. N., Curran G. M., Hausmann L. R. M.. Harnessing Implementation Science to Increase the Impact of Health Equity Research. Med Care. 2017;55 Suppl 9 Suppl 2(Suppl 9 2):S16-S23. doi:10.1097/MLR.0000000000000769 PMID: 28806362 PMCID: PMC5639697
- Koh H. K., Oppenheimer S. C., Massin-Short S. B., Emmons K. M., Geller A. C., Viswanath K.. Translating research evidence into practice to reduce health disparities: a social determinants approach. Am J Public Health. 2010;100 Suppl 1(Suppl 1):S72-80. doi:10.2105/AJPH.2009.167353 PMID: 20147686 PMCID: PMC2837437
- Lane M., Bell R., Latham-Sadler B., Bradley C., Foxworth J., Smith N., Millar A. L., Hairston K. G., Roper B., Howlett A.. Translational Research Training at Various Levels of Professional Experience to Address Health Disparities. J Best Pract Health Prof Divers. 2016;9(1):1178-1187. PMID: 28804792 PMCID: PMC5552239
- Dankwa-Mullan I., Rhee K. B., Stoff D. M., Pohlhaus J. R., Sy F. S., Stinson N., Jr., Ruffin J.. Moving toward paradigm-shifting research in health disparities through translational, transformational, and transdisciplinary approaches. Am J Public Health. 2010;100 Suppl 1(Suppl 1):S19-24. doi:10.2105/AJPH.2009.189167 PMID: 20147662 PMCID: PMC2837422
- Baker E. L., White L. E., Lichtveld M. Y.. Reducing health disparities through community-based research. Public Health Rep. 2001;116(6):517-9. doi:10.1093/phr/116.6.517 PMID: 12196610 PMCID: PMC1497391
- Stuart E. A., Bradshaw C. P., Leaf P. J.. Assessing the generalizability of randomized trial results to target populations. Prev Sci. 2015;16(3):475-85. doi:10.1007/s11121-014-0513-z PMID: 25307417 PMCID: PMC4359056
- Murad M. H., Katabi A., Benkhadra R., Montori V. M.. External validity, generalisability, applicability and directness: a brief primer. BMJ Evid Based Med. 2018;23(1):17-19. doi:10.1136/ebmed-2017-110800 PMID: 29367319
- Pearl Judea, Bareinboim Elias. External Validity: From Do-Calculus to Transportability Across Populations. Statistical Science. 2014;29(4):579-595. doi:10.1214/14-sts486
- Leviton L. C.. Generalizing about Public Health Interventions: A Mixed-Methods Approach to External Validity. Annu Rev Public Health. 2017;38:371-391. doi:10.1146/annurev-publhealth-031816-044509 PMID: 28125391
- Ling A. Y., Montez-Rath M. E., Carita P., Chandross K. J., Lucats L., Meng Z., Sebastien B., Kapphahn K., Desai M.. An Overview of Current Methods for Real-world Applications to Generalize or Transport Clinical Trial Findings to Target Populations of Interest. Epidemiology. 2023;34(5):627-636. doi:10.1097/EDE.0000000000001633 PMID: 37255252 PMCID: PMC10392887
- Kennedy-Martin T., Curtis S., Faries D., Robinson S., Johnston J.. A literature review on the representativeness of randomized controlled trial samples and implications for the external validity of trial results. Trials. 2015;16:495. doi:10.1186/s13063-015-1023-4 PMID: 26530985 PMCID: PMC4632358
- Evans D.. Hierarchy of evidence: a framework for ranking evidence evaluating healthcare interventions. J Clin Nurs. 2003;12(1):77-84. doi:10.1046/j.1365-2702.2003.00662.x PMID: 12519253
- Hernandez S., Fairchild K., Pemberton M., Dahmer J., Zhang W., Palchuk M. B., Topaloglu U.. Applying FHIR Genomics for Research - From Sequencing to Database. AMIA Jt Summits Transl Sci Proc. 2022;2022:236-243. PMID: 35854733 PMCID: PMC9285172
- Evans L., London J. W., Palchuk M. B.. The Detection of Date Shifting in Real-World Data. Appl Clin Inform. 2023;14(4):763-771. doi:10.1055/a-2130-2197 PMID: 37459888 PMCID: PMC10533217
- Antoine A., Perol D., Robain M., Delaloge S., Lasset C., Drouet Y.. Target trial emulation to assess real-world efficacy in the Epidemiological Strategy and Medical Economics metastatic breast cancer cohort. J Natl Cancer Inst. 2023;115(8):971-980. doi:10.1093/jnci/djad092 PMID: 37220893 PMCID: PMC10407701
- Furler J., Magin P., Pirotta M., van Driel M.. Participant demographics reported in "Table 1" of randomised controlled trials: a case of "inverse evidence"?. Int J Equity Health. 2012;11:14. doi:10.1186/1475-9276-11-14 PMID: 22429574 PMCID: PMC3379950
- Munoz Monjas A., Rubio Ruiz D., Perez-Rey D., Palchuk M.. Automatic Outlier Detection in Laboratory Result Distributions Within a Real World Data Network. Stud Health Technol Inform. 2023;302:88-92. doi:10.3233/SHTI230070 PMID: 37203615
- Enewold L., Parsons H., Zhao L., Bott D., Rivera D. R., Barrett M. J., Virnig B. A., Warren J. L.. Updated Overview of the SEER-Medicare Data: Enhanced Content and Applications. J Natl Cancer Inst Monogr. 2020;2020(55):3-13. doi:10.1093/jncimonographs/lgz029 PMID: 32412076 PMCID: PMC7225666
- Forrest C. B., McTigue K. M., Hernandez A. F., Cohen L. W., Cruz H., Haynes K., Kaushal R., Kho A. N., Marsolo K. A., Nair V. P., Platt R., Puro J. E., Rothman R. L., Shenkman E. A., Waitman L. R., Williams N. A., Carton T. W.. PCORnet(R) 2020: current state, accomplishments, and future directions. J Clin Epidemiol. 2021;129:60-67. doi:10.1016/j.jclinepi.2020.09.036 PMID: 33002635 PMCID: PMC7521354
- Qualls L. G., Phillips T. A., Hammill B. G., Topping J., Louzao D. M., Brown J. S., Curtis L. H., Marsolo K.. Evaluating Foundational Data Quality in the National Patient-Centered Clinical Research Network (PCORnet(R)). EGEMS (Wash DC). 2018;6(1):3. doi:10.5334/egems.199 PMID: 29881761 PMCID: PMC5983028
- Ganslandt T., Mate S., Helbing K., Sax U., Prokosch H. U.. Unlocking Data for Clinical Research - The German i2b2 Experience. Appl Clin Inform. 2011;2(1):116-27. doi:10.4338/ACI-2010-09-CR-0051 PMID: 23616864 PMCID: PMC3631913
- Adnyana I. Made Dwi Mertha, Utomo Budi, Eljatin Dwinka S., Setyawan Muhamad F.. Developing and Establishing Attribute-based Surveillance System: A Review. Preventive Medicine: Research & Reviews. 2024;1(2):76-83. doi:10.4103/pmrr.Pmrr_54_23
- Feroz Anam Shahil, Hussaini Anum Shiraz, Seto Emily. Feasibility and Ethical Considerations for Conducting Online versus In-person Interviews for a Qualitative Study. Preventive Medicine: Research & Reviews. 2024;1(6):321-323. doi:10.4103/pmrr.Pmrr_197_24
- Karol Sunidhi, Thakare Meenal M.. Strengthening Immunisation Services in India through Digital Transformation from Co-WIN to U-WIN: A Review. Preventive Medicine: Research & Reviews. 2024;1(1):25-28. doi:10.4103/pmrr.Pmrr_18_23