Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
How evals make or break products
Learn how to turn business KPIs into practical evaluation pipelines, from logging and failing tests to guardrails, cost‑aware metrics, and CI/CD integration.
Bosky:
LLM features don’t fail because the model is “bad.” They fail because teams ship vibe checks instead of measurable evaluation. In this talk I’ll share a practical playbook I use with product teams to move from demos to durable systems.
We’ll break down:
• Mapping business KPIs to evals so PMs, engineering, and SMEs align on “good.”
• Logging first: traces/spans as the backbone for system-level and step-level evals.
• Designing evals that are meant to fail (at first) to expose blind spots instead of overfitting.
• Guardrails vs evals: when to block, when to monitor, and how to decide.
• Small gold sets, smart synthetic data, and programmatic prompt optimization when real data is scarce.
• Cost-aware strategies: target failure slices with user analytics instead of blanket judging everything.
• Multi-turn/agent workflows and tool-selection evals.
• RAG reality: fix chunking and vector dimensions before chasing fancy retrieval metrics.
• Shipping gates and CI/CD: make evals part of the release, not an afterthought.
Format: 35–40 minutes + 10 minutes Q&A. Attendees leave with checklists, sample metrics, and templates they can use the same week.
Compose Email
Loading recent emails...