Young Founders Network
Programm
Vortrag Tech & KI Expert:innen English

Evaluating LLM output without a labelled dataset

Practical evaluation for teams that have users but no annotation budget.

Worum es geht

Nour Haddad presents the setup Klarsicht Labs uses in production: a small golden set, pairwise comparison and a weekly regression run that costs less than a coffee. Includes the failure cases it does not catch.

Speakerin oder Speaker