Programm
Vortrag
Tech & KI
Expert:innen
English
Evaluating LLM output without a labelled dataset
Practical evaluation for teams that have users but no annotation budget.
Worum es geht
Nour Haddad presents the setup Klarsicht Labs uses in production: a small golden set, pairwise comparison and a weekly regression run that costs less than a coffee. Includes the failure cases it does not catch.