Collaborator Spotlight: Masha Khitrun

Maryia (aka Masha) Khitrun is the Technical Lead of the OHDSI Vocabulary team at Odysseus Data Services (an EPAM company), based in Vilnius, Lithuania. A medical doctor by training, she brings hands-on clinical experience to her work in health informatics. Masha started her career in OHDSI as a junior data analyst performing concept mappings and has grown into her expertise in vocabulary development and support, step by step. 

Masha is an active member of the Vocabulary Working Group, and she has led vocabulary workshops at OHDSI Symposia, as well as present at Community Calls. She is the person who brings your expectations for vocabularies to life.

In the latest edition of the Collaborator Spotlight, Masha discusses her career journey, the challenges and importance of maintaining standardized vocabularies, new processes to enhance community contribution, and plenty more.

Can you discuss your background and career journey?

I am a medical doctor by training. My first specialty was General Medicine, and I gained experience across a range of settings — from rural primary care to the ER and inpatient hospital work. But I always dreamed of becoming a neurologist; the human brain has always fascinated me as the greatest mystery in medicine. A few years later, my dream came true: I completed a post-graduate certification in neurology and worked in both General Neurology and Neuroinfectology. Those were fascinating times! I had the chance to treat patients with neuro-COVID and, together with my colleagues, navigate the early days of the epidemic.

During my second maternity leave, I was offered an opportunity to try my hand at data analysis, and I joined Odysseus Data Services. After the war broke out in Ukraine, my family and I decided to leave Belarus. With my clinical practice on hold, I channeled all my energy into the Vocabulary team — and that’s how I ended up where I am today. 

As a member of the vocabulary team with EPAM, your work supports every study in the OHDSI network. What is the biggest hidden challenge in managing global vocabularies that the average analyst might take for granted?

Perhaps the most challenging aspect for me is predicting the downstream impact of any given vocabulary change on ETL pipelines and phenotyping. In our presentations, we often depict the OHDSI ontology as a set of building blocks, with each concept in its place. I would describe it differently: more like a liquid crystal — a structure that is ordered and highly dynamic at the same time, where even the smallest change can have a butterfly effect.

I remember a time when, at the beginning of my career in OHDSI, I mapped a COVID-19-associated HCPCS code incorrectly, which had a dramatic effect on US-based studies and required an urgent hotfix release. That experience taught me just how much the vocabulary work matters to the community.

It’s a common assumption that the vocabulary support represents an effortless, fully automated pipeline. In reality, it’s anything but, and our community contributors can confirm. Every release carries enormous responsibility, involving brainstorming sessions and many rounds of review and testing before the final result is available for download from Athena.

Many people associate OHDSI with OMOP, but can you explain why maintaining updated standardized vocabularies is so critical to the community mission?

OMOP CDM is often what people see first — the beautifully organized structure. But the vocabularies are what give that structure meaning — a universal language that enables collaborative research across various countries.

I like to think of it this way: if OMOP is the grammar of a shared scientific language, then vocabularies are the words. And just like any living language, they must evolve. New drugs come to market, new diseases are described, clinical practice evolves, and coding systems are updated. If vocabularies don’t keep pace, the research built on top of them quietly drifts out of sync with clinical reality — and that’s a problem nobody immediately notices, but everyone eventually feels.

This is especially important for a global OHDSI network, where a study might simultaneously draw on data from the US, the UK, Japan, and the Nordic countries. For that research to be meaningful, a concept must mean the same thing everywhere. Standardized, updated vocabularies are what make that possible.

How does your experience as a physician impact your work as a clinical data analyst, and how important is it to bring wide-ranging experience into health data research?

I would highlight two aspects. First, and most obviously, the medical knowledge in my head allows me to see the data in its entirety. I understand how a doctor thinks when making entries in a patient’s medical record, and I can envision the best way to convert those notes into a standardized format so that doctors and researchers around the world can speak the same language. I also recognize the challenges that arise when a single concept carries multiple meanings depending on perspective.

Furthermore, medical practice has taught me that every action has consequences. Just as in treating a patient, every step and every change in vocabulary development must be well thought out. In difficult cases, we convene a consilium — a collective medical consultation — within our Vocabulary Working Group and discuss possible solutions together.

The result of our work may not be as immediate or tangible as discharging a recovered patient, but it gives me great satisfaction to know that my efforts are also contributing to improving people’s quality of life.

You were a key presenter during a recent OHDSI Vocabulary Workgroup spotlight where the team showcased enhancements to the community contribution process. How do these new processes make it easier for collaborators worldwide to help improve the standard vocabularies?

I’ve always believed that the best improvements come from the people closest to the work, which is exactly why community contributions matter so much to us. I deeply appreciate every contributor who comes to us with ideas for improving the ontology of OHDSI Vocabularies.

However, we noticed that many things that had long been obvious to us could cause real difficulties for contributors. As a rule, we collect contributions right up until the release deadline to cover as many requests as possible. And then, just before the release, it turns out that a seemingly simple contribution isn’t so simple after all — it’s not enough to share your expertise with the Vocabulary team; the submission also needs to be structured correctly. We would start an email thread pointing out mistakes and asking contributors to correct them. With a looming deadline, this caused a lot of stress for everyone involved. And few people enjoy having their mistakes pointed out — it can be discouraging, and next time a contributor might hesitate to reach out again.

That’s why we started thinking about how we could help contributors format their submissions correctly from the start. The automatic verification system, developed by the Vocabulary team together with community members, is designed to help contributors feel confident bringing their ideas forward.

I would like to express my admiration and gratitude to Jared, Konstantin, Anna, and all the colleagues who participated in its development. I’m sure this will make the contribution process more enjoyable and seamless for everyone.

Looking ahead to the rest of 2026 and beyond, what are some primary milestones or major vocabulary updates that you are most excited about for the community?

I’m looking forward to the complex community contribution by the Oncology WG. I know they are doing an amazing job transforming cancer-related vocabularies and improving research in oncology. It also seems to be a fascinating challenge for us to merge these changes into the vocabulary ecosystem.

On a broader scale, I find the potential for applying AI solutions to vocabulary development and QA truly inspiring. As many in the community know, we’ve accumulated a significant amount of technical debt that needs to be resolved, and I hope that with the help of AI, we’ll be able to iron out the kinks and accelerate the process of improving our pipeline.

What are some of your hobbies, and what is one interesting thing that most community members might not know about you?

I can honestly say that life itself is my hobby. I can’t imagine it without books, cooking (I see the kitchen as my own chemical lab), collecting minerals with my kids, and traveling whenever I can. But if I had to name one passion that truly captivates me, it would be the history of language evolution.

Where do words come from, and what historical events are they linked to? How have languages evolved over time, and how have the branches of language families intricately intertwined to shape our culture? Why is the Basque language unique? And why do Finnish and Hungarian belong to the same group, even though they are so different from each other?

At the 2026 OHDSI Europe Symposium, we had a fascinating informal discussion about how we refer to a dog in different languages — Belarusian, Lithuanian, Spanish, Greek, and English. And you know what? It turned out that in Belarusian, the words for “dog” and “puppy” have completely different origins, while in English and Greek, the origin of the word “dog” is simply unknown. Etymology allows us to travel through time, to rediscover how people communicated in the past, and how our culture has transformed to become what we know today.