
Ask a general AI model how many approved drugs hit a target, and it might tell you three when the real answer is six, sounding just as confident either way. If your team is grounding drug discovery decisions in AI output with no way to trace where the answer came from, you're one regulator's question away from a very expensive problem. Lisa Downey is CEO of DrugBank, a structured biomedical intelligence platform cited in more than 60,000 papers and used by nine of the top 20 global pharma companies. She previously built Clarivate's genomic and rare disease data business from the ground up and held leadership roles at GlobalData, giving her almost 20 years across healthcare and life sciences data. Lisa breaks down why the bottleneck in AI-driven drug discovery has shifted from data scarcity to trustworthy grounding, and what that means for teams making target identification and go/no-go calls. You'll hear how DrugBank's knowledge graph separates causation from correlation, why reproducibility matters more than speed, and what questions to ask before building a reference data layer in-house. This episode covers deterministic versus probabilistic data, human-in-the-loop versus human-over-the-loop curation, and how biopharma teams connect grounding layers to their AI agents through MCP. It's built for data and analytics leaders, R&D teams, and anyone deciding whether to build or buy their biomedical data infrastructure. Key Takeaways - A general model asked how many approved drugs hit PD-L1 will answer with total confidence, and total inaccuracy, missing half the real number without any signal that it's wrong. - Anthropic's own benchmarks found frontier models pulling public genomic data got it right as little as 17% of the time, until a deterministic tool pushed accuracy past 99%. - DrugBank moved from human-in-the-loop curation to human-over-the-loop oversight once its data was connected enough that one expert validating one relationship could cascade trust across dozens of related facts. - Before building or buying a reference data layer, Lisa lays out four questions that separate real infrastructure from marketing, starting with whether every fact traces back to a source and a date. Chapter Markers 00:00 Why data scarcity isn't the real bottleneck anymore 01:22 What drew Lisa to DrugBank's mission 03:04 What DrugBank is and who relies on it 05:03 The grounding layer: completeness and reproducibility 07:28 Anthropic's benchmark on data infrastructure 09:21 The high-stakes decisions DrugBank data informs 12:23 Where lost cycle time actually comes from 14:13 DrugBank versus homegrown knowledge graphs 19:43 Human-in-the-loop versus human-over-the-loop curation 24:12 How DrugBank checks its own data quality 25:37 Deterministic versus probabilistic data explained 28:52 The J&J case: separating causation from correlation 33:06 Connecting DrugBank to your AI stack via MCP 37:28 Four questions to ask before you build or buy 42:15 Where DrugBank fits, and where it doesn't 44:44 AI as an amplifier of both good and bad decisions Useful Links & Resources - Connect with Lisa Downey on LinkedIn (https://www.linkedin.com/in/lisaldowney/) Connect With the Show - Ross Katz on LinkedIn (https://www.linkedin.com/in/b-ross-katz/) - CorrDyn on LinkedIn (https://www.linkedin.com/company/corrdyn/) Have you run into an AI model giving you a confident, wrong answer in your own R&D work? Tell us about it in the comments, we're always looking for real examples for future episodes. Visit corrdyn.com to learn how CorrDyn can help your organization extract value from data. #DataInBiotech #BiotechAI #DrugDiscovery #DataScience #LifeSciences
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

How to Identify the Blind Spots in Your Biotech's Genomic Data Before They Cost You a Drug Target

How to Turn Single-Cell Data Into a New Class of Cell-Depleting Therapies

Why Biotech Talks About AI But Won't Pay for the Data It Needs

Beyond Language: Why Drug Discovery Needs Physical AI, Not Just Large Language Models
Free AI-powered recaps of Data in Biotech and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.