
Most drug discovery genomic data comes from a thin slice of the world, and that bias follows every decision downstream. Your team can run a Mendelian randomization study on 35,000 patients and still walk away with a single signal that doesn't even apply to the population you care about. If your phenotype definitions are fuzzy, more data won't save you. Erika Kvikstad is a computational biologist who led precision medicine for cardiovascular disease at Bristol-Myers Squibb, working on therapies including Camzyos for hypertrophic cardiomyopathy. She now works independently on genomic data equity, focused on how reference populations shape everything from target discovery to clinical trial recruitment. You'll get a practical look at how to evaluate real-world data vendors, why heart failure is nearly impossible to define cleanly from billing codes, and where statistical power breaks down even with tens of thousands of patients. Erika also explains how her team used AI to reconstruct missing imaging data and validate cardiomyopathy diagnoses at scale. This episode covers GWAS studies, Mendelian randomization, UK Biobank, proteome-wide analysis, and the practical gap between biobank-scale data and disease-specific cohorts. It's built for data and analytics leaders working in life sciences who need to understand where genomic bias enters their pipeline, not just that it exists. Clarification Around 57:58–58:24, in discussing the proteome-wide Mendelian randomization study, Erika moved quickly between two related findings. BTN3A2 was identified as a candidate associated with ischemic stroke and potential immune-modulatory biology. Separately, single-cell expression data helped contextualize other candidate signals, including some with enriched expression in cardiomyocyte populations. Cardiomyocyte-enriched expression was not a specific finding for BTN3A2. Chapter Markers 00:00 Whose genome are we designing drugs for 01:34 Erika's path from academic genomics to BMS 03:48 Building the precision medicine strategy at BMS 06:37 Ross shares his own HCM diagnosis 07:09 Why heart failure resists clean definition 11:11 How medication use reclassifies patients 14:35 Imaging as a biomarker, and its data gaps 20:23 Data infrastructure gaps across regions 22:44 What to look for when evaluating a data vendor 27:35 Consortia and biobanked specimens for rare mutations 29:52 Cardiovascular data infrastructure versus oncology 32:29 Where statistical power breaks down 37:07 UK Biobank's strengths and its limits 40:01 Bridging broad biobanks with disease-specific cohorts 44:32 How reference population bias propagates downstream 48:53 Where genomic bias hits hardest in the pipeline 53:18 Inside a proteome-wide Mendelian randomization study 59:42 Choosing the right computational tool for the question 1:06:38 Building globally representative genomic infrastructure 1:08:04 Ross's takeaways on bias and statistical power Useful Links & Resources - Erika on LinkedIn: https://www.linkedin.com/in/erikakvikstad - UK Biobank: https://www.ukbiobank.ac.uk - Alliance for Genomic Discovery: https://alliancegenomicdiscovery.org - SHaRe Registry (DCM Foundation): https://dcmfoundation.org Connect With the Show - Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/ - (Ross Katz on X: https://x.com/brosskatz - CorrDyn LinkedIn: https://www.linkedin.com/company/corrdyn/ Have you run into genomic reference bias in your own work? Tell us what it looked like and how your team caught it. Visit corrdyn.com to learn how CorrDyn can help your organization extract value from data. Subscribe to Data in Biotech so you don't miss the next conversation.
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

DrugBank CEO: Why Your AI Model Is Only Giving You Half the Answer

How to Turn Single-Cell Data Into a New Class of Cell-Depleting Therapies

Why Biotech Talks About AI But Won't Pay for the Data It Needs

Beyond Language: Why Drug Discovery Needs Physical AI, Not Just Large Language Models
Free AI-powered recaps of Data in Biotech and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.