
Experimentation and validation of LLM performance is critical when building LLM-driven systems that must reliably deliver a service, from customer service chat bots to intelligence analysis tools. To help teams meet the need for rigorous evaluation methods, a research team in the SEI's AI Division led by Violet Turri has developed the Evaluating Large Language Models (ELM) library, which is built on best practices for LLM evaluation and benchmarking. In the latest episode from the Carnegie Mellon University Software Engineering Institute, Turri sits down with Katie Robinson, a design researcher also in the SEI's AI division, to discuss the ELM library, which turns evaluation from an ad-hoc process into a repeatable, extensible framework.
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

Data-Driven Defense: Cyber Resilience in the Age of AI

Software-Defined Warfare: Expanding the Frontier

From Coordination Chaos to Mission Focus: The Waypoints Framework

Protecting AI Systems Against Data Poisoning
Free AI-powered recaps of Software Engineering Institute (SEI) Podcast Series and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.