2026 Global Symposium Tutorials

Morning Sessions (8 am - 12 pm ET)

An Introduction to the Journey from Data to Evidence Using OHDSI

Faculty: Erica Voss (Johnson & Johnson)

This tutorial describes the journey from raw health data to reliable evidence, covering core concepts like the OMOP Common Data Model, Standardized Vocabularies, and open-source tools such as ATLAS and HADES. This tutorial is for newcomers and aim to explain how data standards, tools, and community practices enable large-scale, open-science research, from data transformation (ETL) to study execution and interpretation. 

An Introduction to ATLAS

Faculty: Chris Knoll (Johnson & Johnson), Sajjan Madappady (EPAM Systems), Peter Hoffmann (Data4Life), Jack Murphy (EPAM)

Each year many OHDSI attendees are new to the community and need a practical, hands-on introduction to the OMOP Common Data Model and the primary tools for reproducible observational research. This tutorial will teach attendees how to use Atlas to define vocabularies and cohorts, explore data, and assemble study-ready outputs. The session will include hands-on walkthroughs, a brief high-level introduction to Strategus concepts (how Atlas outputs can feed distributed execution), and a preview of an updated Atlas user interface that streamlines study creation and review. Note: we will not perform full Strategus study design in the tutorial, and Strategus developer availability for detailed questions will be limited.

Topics:

  • Datasource characterization in Atlas to identify which concepts appear in a dataset over time
  • Searching and browsing the OMOP vocabulary and saving reusable concept sets
  • Building and validating cohort definitions using Atlas’ cohort editor
  • Estimating descriptive measures such as incidence rates and cohort prevalence to understand outcome/exposure distributions
  • Discovering event patterns and treatment pathways with Atlas visualizations
  • Brief, high-level overview of Strategus concepts: how Atlas outputs map to Strategus inputs, the typical execution flow, and considerations when preparing a study for distributed execution (no full Strategus study design)

 

Preliminary Agenda

8:00am – 8:15am

Introduction

  • Who should be using Atlas

Chris

8:15am – 9:00am

The CDM & Data Sources

  • The CDM Data model
  • Database Characterization in Atlas

Chris

9:00am – 9:30am

Vocabulary, Concepts and Concept Sets

  • Vocabulary Overview (Domains, Codes)
  • Concepts: Standard vs. Non-Standard
  • Searching and Filtering
  • Record Counts (Record and Descendants)
  • Building Concept Sets

Jack

9:30am – 11am

Cohort Definitions

  • What is a Phenotype?
  • Cohort Definitions: Logic to build a Phenotype
  • Generation and Reports

Peter

11:00am – 11:45am

Analytics

  • Characterization
  • Incidence Rates
  • Pathways

Jack

11:45am – 12:00pm

Question & Answer

All

Bringing FAIR to Imaging Research with the Medical Imaging OMOP Extension

Faculty: Jen Park (Johns Hopkins University), Kyulee Jeon (Yonsei University), Teri Sippel (Johns Hopkins University), Blake Dewey (Johns Hopkins University), Seng Chan You (Yonsei University), Paul Nagy (Johns Hopkins University)

Medical imaging is an essential source of clinical information, yet imaging data often remain siloed from EHR data and difficult to make Findable, Accessible, Interoperable, and Reusable (FAIR). This hands-on tutorial introduces the Medical Imaging OMOP extension (MI-CDM) and demonstrates an end-to-end workflow for integrating DICOM metadata and imaging-derived features into the OMOP Common Data Model for multimodal observational research.

Participants will learn how to:

  • Understand the design principles of the Medical Imaging OMOP extension (MI-CDM)
  • Index and characterize DICOM archives from clinical or research repositories
  • Standardize DICOM metadata using OMOP-compatible terminology
  • Load imaging metadata into searchable OMOP tables
  • Represent imaging-derived features while preserving provenance
  • Perform multimodal analyses by combining imaging, clinical, and demographic data using standard OHDSI tools such as ATLAS
  • Apply FAIR principles to imaging data within the OMOP ecosystem

Who should attend?

  • OHDSI community members interested in medical imaging
  • Researchers conducting observational or multimodal studies
  • Data engineers and informaticians responsible for imaging ETL pipelines
  • Institutions planning to integrate imaging into their OMOP CDM

By the end of the tutorial, attendees will be able to:

  • Build an imaging-enhanced OMOP dataset from DICOM archives
  • Integrate imaging features with clinical data for downstream analyses
  • Apply reproducible, FAIR-aligned workflows for imaging research within OHDSI

Complex Phenotyping at Scale with and without LLMs Using PhenotypeR

Faculty: Ed Burn (University of Oxford), Dani Prieto-Alhambra (University of Oxford), Nuria Mercade Besora (University of Oxford), Isabella Kaczmarczyk (IQVIA), Atif Adam (IQVIA)

This tutorial will teach students how to run and understand study-specific diagnostics using the OHDSI R package PhenotypeR (https://github.com/OHDSI/PhenotypeR), which enables complex phenotype diagnostics, including drug and measurement diagnostics, matched control sampling, and survival analysis. This can be done with or without support from large language models (LLMs) to help you interpret your results. 

Part 1: Theory

  • We will start with an introduction to clinical phenotyping and its central importance to generating reliable evidence. Different steps in the process will be explained including identifying relevant codes, characterising cohorts and their matched controls as benchmark, drug and measurement specific diagnostics, the use of survival estimates, and population diagnostics including incidence rates and prevalence.
  • The role of diagnostics review by clinical experts and the support provided by LLMs will be introduced and their merits discussed. LLMs are used to contextualise the findings from phenotype characterisation vs previous knowledge, e.g. on disease presentation (signs/symptoms), work-up (e.g. lab or procedures), and treatment/s. Throughout this session lessons learned will be shared from case studies. Results from different LLMs (e.g. Google Gemini, OpenAI, and Mistral) are used to provide examples of the support available.

Part 2: Practical

  • We will do an interactive session using the PhenotypeR package in R (no prior coding experience necessary!). Participants will be given access to an environment with all required packages installed so that they can code along. Using synthetic data, we will first show you how to create cohorts and then how to run diagnostics against these cohorts. Lastly, we will show how to incorporate clinician or LLM based expectations which diagnostics can be checked against.
  •  
Preliminary Agenda:

8.00 – 8.15

Welcome and introductions (Ed)

8.15 – 8.30

Introduction to standardised and reproducibly phenotyping (Dani)

8.30 – 9.00

Designing phenotypes (Atif)

9.00 – 9.30

Introduction to the PhenotypeR R package (Isabella)

9.30 – 10.00

Break

10.00 – 11.00

Running PhenotypeR (Nuria)

11.00 – 11.30

Adding LLM-based expectations to PhenotypeR (Ed)

11.30 – 11.45

Take-home messages (Dani)

11.45 – 12.00

Q&A (All)

OHDSI Change Leaders & Leadership: Helping Community Members Grow Adoption, Engagement and Impact

Faculty: Liesbet Peeters (Hasselt University), Chris Baldwin (Unison), Christian Hogberg (Passion 2 Improve Sweden AB), J. Swetha (Global Value Web)

This interactive tutorial is designed for OHDSI community members who want to strengthen their ability to drive change, engage stakeholders and grow the adoption of OHDSI and the OMOP Common Data Model within their own environment. Participants will work on a real-world use case and leave with practical insights, tools and a concrete action plan. 

08:00 – 09:00 | Part I – Start with the End in Mind

Focus: Connecting personal motivation with meaningful action.

We begin by exploring why participants care about the OHDSI mission and how they hope to contribute to it. Through individual reflection and peer conversations, participants identify a concrete stakeholder engagement opportunity, meeting or initiative that they will use as a personal case throughout the workshop. 

Outcome: Participants leave with a clear personal goal and a real-world use case that serves as the foundation for the remainder of the tutorial.

09:00 – 10:00 | Part II – Building the Bridge from Us to Them

Focus: Understanding stakeholders and learning how to communicate value.

Participants explore different stakeholder perspectives, identify motivations and concerns, and learn how to build compelling narratives that connect the OHDSI mission to the needs of specific audiences. Interactive exercises help participants translate technical and scientific value into meaningful stakeholder stories. 

Outcome: Participants develop a draft stakeholder engagement narrative tailored to their selected use case.

10:00 – 10:15 | Networking Break

Opportunity to connect with peers, exchange experiences and continue informal discussions.

10:15 – 11:15 | Part III – Becoming a Change Leader

Focus: Driving adoption and leading change.

This section introduces practical concepts from change management, leadership and community growth. Participants explore the role they can play in helping others discover, adopt and support OHDSI initiatives. 

Outcome: Participants gain a better understanding of the mindset, skills and approaches that help accelerate adoption and engagement.

11:15 – 12:00 | Part IV – Bringing It All Together

Focus: Applying everything to a real-world challenge.

Participants work collaboratively to further develop their stakeholder engagement approach and receive feedback from peers and facilitators. The emphasis is on practical application and preparing for conversations, meetings and initiatives that participants are likely to encounter shortly after the workshop. 

Outcome: Participants leave with a concrete plan, actionable next steps and feedback that can immediately be applied within their own context.

Who Should Attend?

This tutorial is intended for:

  • National Node Leaders
  • Study leads and project leaders
  • Researchers and data scientists
  • Healthcare professionals and innovators
  • Industry representatives
  • Community members who want to support OHDSI growth and adoption within their organisation, region or network

Anyone interested in becoming a stronger advocate, connector, facilitator or change leader within the OHDSI ecosystem is welcome.

Mastering OMOP: Transforming EHR Data with Practical Strategies, Best Practices, and OHDSI Integration

Faculty: Melanie Philofsky (EPAM Systems), Rakesh Babu (Atlantic Health System)

Join us for an immersive 4-hour workshop tailored to professionals and researchers working with electronic health record (EHR) data, whether you are new to the OHDSI community or have been actively involved for years. Led by an experienced OHDSI instructor alongside seasoned veterans from top-tier academic medical centers, this workshop will strengthen your understanding of the OMOP Common Data Model (CDM) while offering practical strategies for its adoption, implementation, and use within health systems.

The first hour of the session will explore the foundation and evolution of the OMOP CDM, addressing critical questions about why this model is essential and the broader mission of OHDSI. We will discuss the challenges of working with EHR data and how OMOP provides a unique framework for driving cross-institutional, reproducible research. You will gain insights on how OMOP compares to other data models conceptually, preparing you for the practical considerations ahead.

The second section delves into the real-world intricacies of creating and maintaining a pragmatic OMOP CDM. Through engaging lectures and relatable best practices, you will learn how to tailor the OMOP CDM to your organization’s specific needs, align vocabulary mappings to international standards, and maintain high-quality data through customized quality checks. This segment will demystify some of the most technical aspects of OMOP adoption, ensuring attendees leave with a concrete understanding of how to start or refine their processes.

In the final hour, we will transition into broader discussions focused on joining and leveraging OHDSI’s collaborative networks. You will discover opportunities to contribute to global research initiatives, build partnerships within disease-based or location-based working groups, and collaborate with others. A facilitated networking session will allow attendees to interact with OHDSI experts, share experiences, and expand their professional connections within the community.

This interactive workshop blends lectures, small group discussions, and networking opportunities to ensure attendees receive practical insights and actionable strategies for their work. Whether you are embarking on your OHDSI journey or are a seasoned contributor, this session promises to provide value, tools, and connections to empower your work with EHR data.

Afternoon Sessions (1 pm - 5 pm ET)

Building and Using the OHDSI Evidence Network: From Data Partner to Federated Study Execution

Faculty: Clair Blacketer (Johnson & Johnson), Paul Nagy (Johns Hopkins University, Scott Duvall (Purple Lab), Hanieh Razzaghi (CHOP)
 
The OHDSI Evidence Network is a global, open, federated research infrastructure designed to support large-scale real-world evidence generation while preserving local data governance and institutional autonomy. As participation in the network grows, there is increasing demand for clear, practical guidance on how to engage with the network—both from the perspective of data partners contributing data and study leads conducting multi-site research.  This tutorial is structured as two complementary 2-hour sessions, designed to align expectations across the network and reduce friction in federated research.
 
Session 1: Participating in the Evidence Network as a Data Partner (2 hours)
This session is intended for data custodians, analysts, and institutional stakeholders at organizations with OMOP CDM–mapped data. Topics include:
– An overview of the Evidence Network’s federated, opt-in operating model
– How to communicate the value and return on investment of participation to institutional leadership
– Common governance and oversight considerations (e.g., IRB review, data sharing boundaries)
– What participation in a network study entails, including roles, responsibilities, and timelines
– How data partners engage with feasibility, execution, and iterative study runs
– The goal of this session is to equip data partners with a clear understanding of how and why to participate, and what to expect when engaging in network studies.
 
Session 2: Conducting Studies Through the Evidence Network as a Study Lead (2 hours)
This session is intended for investigators and methodologists interested in leading or coordinating federated studies. Topics include:
– Designing studies suitable for federated execution
– The Evidence Network study lifecycle, from idea through synthesis
– Roles and responsibilities of study leads, data partners, and coordinating teams
– Execution, QA/QC, and managing cross-site variability
– Synthesizing and publishing results from multi-site studies
 
Learning Outcomes
Across both sessions, participants will gain a shared mental model of how the Evidence Network operates, practical guidance for participation or leadership, and tools to support high-quality federated research within the OHDSI community. 

Symposium Registration Process

From Multi-Modal Data to Real-World Evidence: Hands-on with the Data2Evidence Platform for OMOP Data Curation and Analytics

 
Faculty: Peter Hoffmann (Data4Life), Mukkesh Kumar (OptiExacta Labs), Tim Walz (Date4Life)
 
Hands-on activities will walk participants through an example study from specification to execution and review. Attendees will (1) extract raw test data (e.g., CSV, database, EHR export) and transform it into the OMOP CDM format, (2) load the transformed data into an OMOP CDM database and validate it with DQD to ensure data quality, (3) configure a Data2Evidence study from a template, (4) run the Data2Evidence workflow on the provided dataset or local OMOP CDM connection, (5) review key diagnostics and sensitivity checks for the study, and (6) produce an evidence “packet” suitable for internal review or multi-site collaboration. Participants can also discuss with the developers on how to translate their own research questions into Data2Evidence-ready specifications for scalable evidence generation.
 
Prerequisites: basic familiarity with OMOP CDM and cohort concepts is not required but helpful. Participants should bring a laptop and will receive pre-tutorial setup instructions (software, credentials, and test data options). By the end, attendees will be able to author an OHDSI network study specification, execute it reproducibly, interpret diagnostics, and share a standardized evidence artifact for collaboration and decision-making

Integrating Geospatial Data Into OMOP CDM

Faculty: Jared Houghtaling (Johnson & Johnson), Robert Miller (Miller Data Solutions), Polina Talapova (SciForce), Timothy Norris (University of Miami), Jacob Zelko (Georgia Tech), Jay Greenfield (CoData)
 
This hands-on tutorial introduces the OHDSI GIS toolchain for integrating environmental exposures and social determinants into health research. Participants will discover, process, and analyze place-based health determinants using OMOP extensions.
 
Session 1: Cataloging and Data Discovery (1 hour)
Discover geospatial datasets using gaiaCatalog’s Schema.org-compliant interface. Participants will search datasets, understand metadata documentation, and author functional metadata for automated data retrieval. Exercises cover environmental, social, and demographic data sources relevant to health research.
 
Session 2: The Gaia Pipeline (1 hour)
Learn the workflow from raw geospatial data to OMOP-standardized tables. Deploy gaiaDocker locally, ingest public datasets (EPA air quality), perform spatial transformations using gaiaDb/PostGIS, and populate the external_exposure table. Includes geocoding with DeGauss, spatial joins, and privacy-preserving aggregation.
 
Session 3: OMOP Integration (1 hour)
Explore external_exposure and location_history tables that extend OMOP CDM for geospatial analytics. Understand vocabulary integration (OMOP GIS, Exposome, SDoH) and query exposure data alongside clinical observations. Calculate temporal exposure metrics (e.g., PM2.5 during pregnancy) and link exposures to cohort definitions.
 
Session 4: Analytical Applications (1 hour)
Integrate HADES tools for geospatially-informed research. Use FeatureExtraction for spatial covariates and PatientLevelPrediction with environmental features. Explore prototype extensions (GeoFeatureExtraction, SpatialCohortMethod), privacy-preserving visualization, and federated network studies.
Target Audience: Researchers and informaticians interested in environmental epidemiology or social determinants. Basic OMOP familiarity helpful.

Introduction to OHDSI Phenotype Development & Evaluation

Faculty: Azza Shoaibi (Johnson & Johnson), Gowtham Rao (CoReason.AI)
 
Accurate phenotyping is a cornerstone of reliable observational research, yet it remains one of the greatest methodological challenges in real‑world data analytics. Observational data are inherently prone to misclassification and these errors can meaningfully influence study validity. This mid‑level OHDSI tutorial is designed for participants who are familiar with OHDSI standardized vocabularies and already comfortable developing cohorts in ATLAS, and who now seek to deepen their scientific understanding and strengthen their applied phenotype development skills. The tutorial will begin with an overview of the science of misclassification error, covering key concepts such as sensitivity, specificity, predictive value, and index event misclassification. Participants will learn how these errors arise within observational data sources and how they directly affect effect estimation, transportability, and downstream decision‑making.  Building on this foundation, the session will transition to OHDSI’s established best practices for phenotype development and evaluation. Using real examples, we will walk through tools and methods that support transparent, reproducible, and scalable phenotype definitions—including cohort diagnostics, standard vocabulary exploration, benchmarking against population characteristics.  The latter portion of the tutorial introduces an emerging opportunity for improving phenotype development: an AI‑aided iterative workflow. We will demonstrate how AI/LLM can support iterative refinement as part of phenotype development.
 
By the end of this tutorial, participants will have a richer understanding of the scientific foundations of phenotype misclassification, practical experience applying OHDSI best practices, and early exposure to how AI‑enabled workflows can enhance rigor and reproducibility. This session is ideal for researchers ready to advance from simply using ATLAS to mastering phenotype design as a scientific discipline. 

OHDSI Standardized Vocabularies on FHIR: A Deep Dive Using the Echidna Terminology Server

Faculty: Davera Gabriel (Evidentli), Guy Tsafnat (Echidna Systems), Jean Duteau (Dogwood Health Consulting)
 
This hands-on tutorial teaches participants to use Echidna (https://echidna.fhir.org), the authoritative source for OHDSI Standardized Vocabularies on a FHIR Terminology Server. Beginning with the conceptual foundations of terminology management across FHIR, OMOP, and openEHR, participants will progress to practical skills executing FHIR API calls against a production terminology server –  including concept lookup, value set expansion for phenotype development, and source code translation to OMOP Standard Concepts. The tutorial concludes with an overview of the HL7 Vulcan FHIR-to-OMOP Implementation Guide and patterns for integrating Echidna into ETL pipelines and OHDSI workflows.

Using OMOP Model in Registry Context & Clinical Trials Standardization Context: Conventions, Past Use Cases, SDTM & Regulatory Consideration, Challenges

Faculty: Vojtech Huser (EPAM Systems), Cynthia Sung (Duke-NUS Medical System), Mike Hamidi (Walmart Health & Wellness), Darya Zhukova (EPAM Systems)
 
Registry and clinical trial (CT) data often need to fill the gap not covered by traditional RWD. If RWD follows OMOP, using the same model for harmonization of registry and CT data has advantages.
Outline:
1. Special considerations found in registry and CT data (e.g., converting relative dates to OMOP compliant dates, CRF data, adverse drug event (seriousness, severity, and relatedness between drug and AE))
2. Use of 2B concepts (custom vocabularies) vs OMOP vocabularies
3. Use cases published in literature (e.g., UK Biobank, OHDSI 2024 poster: Application of OMOP Common Data Model to Disease Registry Data, AllOfUs)
4. Other use cases (AACR GENIE)
5. OMOP and SDTM comparison (+ FDA regulatory considerations)
6. Combining trial and RWD data (external control arm) considerations (data granularity mismatch, RWD limited to routine healthcare, lack of advanced COAs, data collection not regulatory grade)