the repo โ†’

๐Ÿถ Get started

Learn LLM evals by fixing an app that's quietly lying to its customers.

A dog daycare's AI writes report cards for owners. Rufus ate nothing, never played and hid under a bench โ€” his report says he had a lovely, sociable day. Nine steps take you from spotting that to measuring it, with Braintrust.

curl -fsSL https://destinio.github.io/evals-by-example/install.sh | bash

Needs Bun 1.4+ and git. It clones the repo, installs dependencies and asks for two API keys. Read the script first, or do it by hand: git clone https://github.com/destinio/evals-by-example.git && cd evals-by-example && bun run setup

๐Ÿถ The whole course. 1 of 9 steps written โ€” the rest are planned below and get written as the course is worked through, so they can quote real runs.

Learning Braintrust, on a real app

You've been hired by a dog daycare to make their AI better. This is the path.

Each step is one idea, one markdown file, and one commit. You write the code; the step file tells you what and why. Nothing here is a finished solution you paste blindly โ€” the point is to leave with the habits, not the repo.

Before step 1

  1. The app runs (bun run app) and you've read a couple of report cards.
  2. .env has NOUS_API_KEY. Get a free Braintrust account and add BRAINTRUST_API_KEY too โ€” step 1 needs it.
  3. You've seen the problem: click Rufus, compare his report to the timeline underneath.

Stuck? Ask the tutor

In a Claude Code session in this repo:

/tutor

The tutor figures out where you are, explains the next idea, and checks your work when you say a step is done. Good moments to call it: "what's next", "check my work", "why isn't this showing up in Braintrust", or any time a word like span or scorer stops making sense.

It won't write the step for you. You'll get the explanation and the code to apply yourself, one idea at a time.

How to work through it

Do it all in your copy of the repo, on your own branch, so main stays clean:

git switch -c my-course

(The one-line installer already did this for you.)

Then work through the steps in order, in the same folder. Commit whenever you like โ€” it's your branch.

Getting course updates

New steps land on main. To pull them into your branch:

git switch main && git pull
git switch my-course && git merge main

Starting over

git switch main                    # the untouched app
rm app/data/happytails.db          # forget generated reports; reseeds on next launch

Your my-course branch is still there if you want to go back to it.

The steps

Nine steps, each one leaving you able to show something, not just know it. The last column is the moment you'd put on a screen for your team.

#StepWhat you buildBraintrustThe demo beat
1Observe โœ…wrapOpenAI + traced in the appTracing, Instrumentation"Every report we've ever written, with its cost and the data behind it"
2Score one thingA scorer in plain code, and your first Eval()Evaluation"67%. And here's the exact row that failed"
3Keep the hard casesA dataset: the four dogs plus deliberately nasty daysDatasets"The awkward days are permanent now โ€” nothing can quietly break them again"
4Fix the prompt, prove itA second prompt version and a second experimentExperiment comparison"v1 against v2, row by row, including what got worse"
5Judge what code can'tAn LLM-as-judge scorer, checked against your own opinionScorers, Loop"Code catches invented numbers. A judge catches buried medication"
6Ask the humansThumbs and staff corrections in the app, flowing backAnnotation, human review"A complaint on Tuesday becomes a test case on Wednesday"
7Iterate without deployingThe prompt pulled from Braintrust instead of SQLitePlayground, Prompts"Change the prompt, try it on 20 real days, ship it โ€” no deploy"
8Watch productionScorers running on live traffic, and a dashboardObservation, Deployment"Quality and cost per day, and an alert when either moves"
9Ship it honestlyA pre-ship check you'd actually runbt CLI, CI"What runs before anyone changes the prompt"

Steps 2 onward are written as you reach them, so each one describes what actually happened in your runs โ€” real scores, real regressions, the false alarm your first scorer raises โ€” rather than a script written in advance.

Where it's going

Steps 1โ€“4 are the core loop and stand alone: see it, score it, keep the hard cases, prove a change helped. Steps 5โ€“6 handle what code can't judge and where human opinion enters. Steps 7โ€“8 are what your colleagues will care about most โ€” prompt changes that aren't deploys, and production that scores itself. Step 9 makes it routine, and ends by mapping all of it back to whatever prompt you're really responsible for.

The words

You'll hit these constantly, and they're easy to confuse.

TermWhat it is
SpanOne recorded step โ€” a function call, a model call.
TraceA span plus everything nested inside it. One report card is one trace.
LogA trace from real use. Continuous, unpredictable, no right answer attached.
DatasetA fixed set of test cases, usually an input and some notion of what good looks like.
TaskThe code under test. Here: writeReport().
ScorerA function that grades one output, 0 to 1. Plain code, or another model.
ExperimentOne run of the task over the dataset with the scorers applied. Has an average score, and can be compared to any earlier run.

Logs are what happened. Experiments are what would happen if. They're different screens in Braintrust, and looking in the wrong one is the most common early confusion.

The demo at the end

The point of the course is a working system; the point of the repo is that you can show it. Each step file ends with what you can now demonstrate, and step 9 assembles those into learn/demo.md โ€” a run of show with timings, the commands to type, and the lines to say.

Two things make the demo land, and both are decisions made early:

The loop you're building toward

  1. Log what production does.
  2. Find the bad ones.
  3. Turn them into a dataset with a notion of correct.
  4. Run an experiment โ€” that number is your baseline.
  5. Change the prompt, run again, compare.
  6. Ship if it went up, and keep the dataset forever, so the next change can't quietly break what you just fixed.

Step 5 is the payoff. Without it, every prompt edit is a coin flip โ€” and the classic failure is fixing one complaint while breaking three things nobody retested.