In The Loop

How StackOne built an auto-research loop to continually improve their product

August 21, 2026·57 min
Episode Description from the Publisher

Once a week, an agent at StackOne goes hunting for new prompt injection attacks. It reads papers, trawls Reddit, tries what it finds against real models, keeps the attacks that land, and retrains StackOne's defence model on them. A person still approves every deployment. Guillaume Lebedel, their CTO, puts the whole thing on screen, including the five experiments out of six that failed.In this episode of In The Loop, Guillaume talks about how and why they have built an auto-research loop. What one is in plain words, how you pick a goal ai agents can measure, how you sample 100,000 test cases down to something you can afford, and why he runs evals on cheap models before trusting anything. We also spoke about token leaderboards and the impact its had internally. ⏭️ Episode highlights(05:41) – What an auto-research loop is(11:53) – Why an agent is a folder(14:47) – Screen-share: the attack-hunting agent(18:00) – Six experiments, one promoted(27:29) – Sampling 100,000 test cases down(31:48) – Ninety per cent of tokens are wasted(39:33) – Turning off extra usage the same day(41:13) – Screen-share: the token derby

Podzilla Summary coming soon

Sign up to get notified when the full AI-powered summary is ready.

Get Free Summaries →

Free forever for up to 3 podcasts. No credit card required.

Listen to This Episode

Get summaries like this every morning.

Free AI-powered recaps of In The Loop and your other favorite podcasts, delivered to your inbox.

Get Free Summaries →

Free forever for up to 3 podcasts. No credit card required.