Doors08:30→08:30
MAINSTAGE
Next up
Home
Intermediate
Build an eval suite in an afternoon
Leave with tests that block a bad model release.
- Day
- Day 1 · Tue 18 May
- Starts
- 14:00
- Length
- 3 hours
- Price
- €180 per person

What you'll build
A working evaluation suite for a prompt or agent you already run: fifty cases, a pass mark and a script that fails a release when the mark is missed.
Who it's for
Engineers and product people who ship AI features and have no reliable way to tell whether a change made things better or worse. You should be comfortable running a script from a terminal.
Agenda
- 14:00 What a good eval case looks like, with examples that caught real regressions
- 14:40 Write your first fifty cases from your own inputs
- 15:40 Break
- 16:00 Wire the suite into a pull request check
- 16:40 Review in pairs and plan the next two hundred cases
You leave with tests that run, not slides about tests.
What to bring
- Laptop and charger
- A prompt or agent you already run
- Twenty real inputs from your product
Intermediate
Build an eval suite in an afternoon
Leave with tests that block a bad model release.
- Day
- Day 1 · Tue 18 May
- Starts
- 14:00
- Length
- 3 hours
- Price
- €180 per person

What you'll build
A working evaluation suite for a prompt or agent you already run: fifty cases, a pass mark and a script that fails a release when the mark is missed.
Who it's for
Engineers and product people who ship AI features and have no reliable way to tell whether a change made things better or worse. You should be comfortable running a script from a terminal.
Agenda
- 14:00 What a good eval case looks like, with examples that caught real regressions
- 14:40 Write your first fifty cases from your own inputs
- 15:40 Break
- 16:00 Wire the suite into a pull request check
- 16:40 Review in pairs and plan the next two hundred cases
You leave with tests that run, not slides about tests.
What to bring
- Laptop and charger
- A prompt or agent you already run
- Twenty real inputs from your product
Intermediate
Build an eval suite in an afternoon
Leave with tests that block a bad model release.
- Day
- Day 1 · Tue 18 May
- Starts
- 14:00
- Length
- 3 hours
- Price
- €180 per person

What you'll build
A working evaluation suite for a prompt or agent you already run: fifty cases, a pass mark and a script that fails a release when the mark is missed.
Who it's for
Engineers and product people who ship AI features and have no reliable way to tell whether a change made things better or worse. You should be comfortable running a script from a terminal.
Agenda
- 14:00 What a good eval case looks like, with examples that caught real regressions
- 14:40 Write your first fifty cases from your own inputs
- 15:40 Break
- 16:00 Wire the suite into a pull request check
- 16:40 Review in pairs and plan the next two hundred cases
You leave with tests that run, not slides about tests.
What to bring
- Laptop and charger
- A prompt or agent you already run
- Twenty real inputs from your product

11 · See you there
See you in Lisbon.
Doors open · Tue, 18 May 2027 at 09:00 · Pavilhão do Rio

11 · See you there
See you in Lisbon.
Doors open · Tue, 18 May 2027 at 09:00 · Pavilhão do Rio

11 · See you there
See you in Lisbon.
Doors open · Tue, 18 May 2027 at 09:00 · Pavilhão do Rio


