Intermediate

Leave with tests that block a bad model release.

Day
Day 1 · Tue 18 May
Starts
14:00
Length
3 hours
Price
€180 per person
Build an eval suite in an afternoon

What you'll build

A working evaluation suite for a prompt or agent you already run: fifty cases, a pass mark and a script that fails a release when the mark is missed.

Who it's for

Engineers and product people who ship AI features and have no reliable way to tell whether a change made things better or worse. You should be comfortable running a script from a terminal.

Agenda

  • 14:00 What a good eval case looks like, with examples that caught real regressions
  • 14:40 Write your first fifty cases from your own inputs
  • 15:40 Break
  • 16:00 Wire the suite into a pull request check
  • 16:40 Review in pairs and plan the next two hundred cases
You leave with tests that run, not slides about tests.

What to bring

  • Laptop and charger
  • A prompt or agent you already run
  • Twenty real inputs from your product

Intermediate

Leave with tests that block a bad model release.

Day
Day 1 · Tue 18 May
Starts
14:00
Length
3 hours
Price
€180 per person
Build an eval suite in an afternoon

What you'll build

A working evaluation suite for a prompt or agent you already run: fifty cases, a pass mark and a script that fails a release when the mark is missed.

Who it's for

Engineers and product people who ship AI features and have no reliable way to tell whether a change made things better or worse. You should be comfortable running a script from a terminal.

Agenda

  • 14:00 What a good eval case looks like, with examples that caught real regressions
  • 14:40 Write your first fifty cases from your own inputs
  • 15:40 Break
  • 16:00 Wire the suite into a pull request check
  • 16:40 Review in pairs and plan the next two hundred cases
You leave with tests that run, not slides about tests.

What to bring

  • Laptop and charger
  • A prompt or agent you already run
  • Twenty real inputs from your product

Intermediate

Leave with tests that block a bad model release.

Day
Day 1 · Tue 18 May
Starts
14:00
Length
3 hours
Price
€180 per person
Build an eval suite in an afternoon

What you'll build

A working evaluation suite for a prompt or agent you already run: fifty cases, a pass mark and a script that fails a release when the mark is missed.

Who it's for

Engineers and product people who ship AI features and have no reliable way to tell whether a change made things better or worse. You should be comfortable running a script from a terminal.

Agenda

  • 14:00 What a good eval case looks like, with examples that caught real regressions
  • 14:40 Write your first fifty cases from your own inputs
  • 15:40 Break
  • 16:00 Wire the suite into a pull request check
  • 16:40 Review in pairs and plan the next two hundred cases
You leave with tests that run, not slides about tests.

What to bring

  • Laptop and charger
  • A prompt or agent you already run
  • Twenty real inputs from your product
Confetti falling over the crowd at the closing keynote

11 · See you there

Doors open · Tue, 18 May 2027 at 09:00 · Pavilhão do Rio

See the lineup

64 speakers · 1,800 seats · 3 days

Confetti falling over the crowd at the closing keynote

11 · See you there

Doors open · Tue, 18 May 2027 at 09:00 · Pavilhão do Rio

See the lineup

64 speakers · 1,800 seats · 3 days

Confetti falling over the crowd at the closing keynote

11 · See you there

Doors open · Tue, 18 May 2027 at 09:00 · Pavilhão do Rio

See the lineup

64 speakers · 1,800 seats · 3 days

The summit

18–20 May 2027
Lisbon

Pavilhão do Rio
Cais do Rio 14, 1200-450 Lisbon, Portugal

hello@mainstage.events

Get speaker announcements.

Two attendees talking in the hallway
The audience during a keynote
A speaker in front of the main stage screen
© 2026 Mainstage. All rights reserved.Made by SoloFoundry

Create a free website with Framer, the website builder loved by startups, designers and agencies.