๐ถ Get started
Learn LLM evals by fixing an app that's quietly lying to its customers.
A dog daycare's AI writes report cards for owners. Rufus ate nothing, never played and hid under a bench โ his report says he had a lovely, sociable day. Nine steps take you from spotting that to measuring it, with Braintrust.
curl -fsSL https://destinio.github.io/evals-by-example/install.sh | bash
Needs Bun 1.4+ and git. It clones the repo, installs
dependencies and asks for two API keys. Read the script first,
or do it by hand: git clone https://github.com/destinio/evals-by-example.git && cd evals-by-example && bun run setup
๐ถ The whole course. 1 of 9 steps written โ the rest are planned below and get written as the course is worked through, so they can quote real runs.
Learning Braintrust, on a real app
You've been hired by a dog daycare to make their AI better. This is the path.
Each step is one idea, one markdown file, and one commit. You write the code; the step file tells you what and why. Nothing here is a finished solution you paste blindly โ the point is to leave with the habits, not the repo.
Before step 1
- The app runs (
bun run app) and you've read a couple of report cards. .envhasNOUS_API_KEY. Get a free Braintrust account and addBRAINTRUST_API_KEYtoo โ step 1 needs it.- You've seen the problem: click Rufus, compare his report to the timeline underneath.
Stuck? Ask the tutor
In a Claude Code session in this repo:
/tutor
The tutor figures out where you are, explains the next idea, and checks your work when you say a step is done. Good moments to call it: "what's next", "check my work", "why isn't this showing up in Braintrust", or any time a word like span or scorer stops making sense.
It won't write the step for you. You'll get the explanation and the code to apply yourself, one idea at a time.
How to work through it
Do it all in your copy of the repo, on your own branch, so main stays clean:
git switch -c my-course
(The one-line installer already did this for you.)
Then work through the steps in order, in the same folder. Commit whenever you like โ it's your branch.
Getting course updates
New steps land on main. To pull them into your branch:
git switch main && git pull
git switch my-course && git merge main
Starting over
git switch main # the untouched app
rm app/data/happytails.db # forget generated reports; reseeds on next launch
Your my-course branch is still there if you want to go back to it.
The steps
Nine steps, each one leaving you able to show something, not just know it. The last column is the moment you'd put on a screen for your team.
| # | Step | What you build | Braintrust | The demo beat |
|---|---|---|---|---|
| 1 | Observe โ | wrapOpenAI + traced in the app | Tracing, Instrumentation | "Every report we've ever written, with its cost and the data behind it" |
| 2 | Score one thing | A scorer in plain code, and your first Eval() | Evaluation | "67%. And here's the exact row that failed" |
| 3 | Keep the hard cases | A dataset: the four dogs plus deliberately nasty days | Datasets | "The awkward days are permanent now โ nothing can quietly break them again" |
| 4 | Fix the prompt, prove it | A second prompt version and a second experiment | Experiment comparison | "v1 against v2, row by row, including what got worse" |
| 5 | Judge what code can't | An LLM-as-judge scorer, checked against your own opinion | Scorers, Loop | "Code catches invented numbers. A judge catches buried medication" |
| 6 | Ask the humans | Thumbs and staff corrections in the app, flowing back | Annotation, human review | "A complaint on Tuesday becomes a test case on Wednesday" |
| 7 | Iterate without deploying | The prompt pulled from Braintrust instead of SQLite | Playground, Prompts | "Change the prompt, try it on 20 real days, ship it โ no deploy" |
| 8 | Watch production | Scorers running on live traffic, and a dashboard | Observation, Deployment | "Quality and cost per day, and an alert when either moves" |
| 9 | Ship it honestly | A pre-ship check you'd actually run | bt CLI, CI | "What runs before anyone changes the prompt" |
Steps 2 onward are written as you reach them, so each one describes what actually happened in your runs โ real scores, real regressions, the false alarm your first scorer raises โ rather than a script written in advance.
Where it's going
Steps 1โ4 are the core loop and stand alone: see it, score it, keep the hard cases, prove a change helped. Steps 5โ6 handle what code can't judge and where human opinion enters. Steps 7โ8 are what your colleagues will care about most โ prompt changes that aren't deploys, and production that scores itself. Step 9 makes it routine, and ends by mapping all of it back to whatever prompt you're really responsible for.
The words
You'll hit these constantly, and they're easy to confuse.
| Term | What it is |
|---|---|
| Span | One recorded step โ a function call, a model call. |
| Trace | A span plus everything nested inside it. One report card is one trace. |
| Log | A trace from real use. Continuous, unpredictable, no right answer attached. |
| Dataset | A fixed set of test cases, usually an input and some notion of what good looks like. |
| Task | The code under test. Here: writeReport(). |
| Scorer | A function that grades one output, 0 to 1. Plain code, or another model. |
| Experiment | One run of the task over the dataset with the scorers applied. Has an average score, and can be compared to any earlier run. |
Logs are what happened. Experiments are what would happen if. They're different screens in Braintrust, and looking in the wrong one is the most common early confusion.
The demo at the end
The point of the course is a working system; the point of the repo is that you can show it. Each step file ends with what you can now demonstrate, and step 9 assembles those into learn/demo.md โ a run of show with timings, the commands to type, and the lines to say.
Two things make the demo land, and both are decisions made early:
- Rufus. A quiet, sad little day that the model turns into a social triumph. Everyone in the room sees the problem in five seconds without knowing anything about evals.
- The numbers move. The scores in your demo are real runs from this repo, not slides. The regressions are real too โ including the one in step 4 where the fix breaks something else.
The loop you're building toward
- Log what production does.
- Find the bad ones.
- Turn them into a dataset with a notion of correct.
- Run an experiment โ that number is your baseline.
- Change the prompt, run again, compare.
- Ship if it went up, and keep the dataset forever, so the next change can't quietly break what you just fixed.
Step 5 is the payoff. Without it, every prompt edit is a coin flip โ and the classic failure is fixing one complaint while breaking three things nobody retested.