the repo โ†’

Step 1 โ€” Observe

The idea: you can't improve what you can't see. Before any prompt work, make every report the app writes visible and reviewable.

Working with Claude Code? Run /tutor and it'll walk this step with you, then check it.

Make sure you're on your own branch (the installer does this), and the app is running:

git switch -c my-course     # skip if you're already on it
bun run app

Why this first

The staff say the reports are "sometimes weird." Try answering any of these today:

You can't answer one of them. The model call happens, the text lands on a screen, and it's gone. Every later step โ€” scoring, datasets, comparing prompts โ€” needs a record, and no record exists.

What you'll add

Two wraps, doing different jobs:

The model call nests inside the operation, so one report card becomes one trace with two levels:

daily report                    input: Rufus's day log ยท output: the report card
โ””โ”€โ”€ chat.completions.create     system + user messages ยท 412 tokens ยท $0.0004 ยท 2.1s

Wrap only the client and you'll know a call happened but not who it was for. Wrap only the operation and you lose cost and tokens. You want both.

Do it

1. Install the SDK

bun add braintrust

2. app/llm.ts โ€” wrap the client

import OpenAI from 'openai'
import { initLogger, wrapOpenAI } from 'braintrust'

const apiKey = process.env.NOUS_API_KEY
if (!apiKey) {
  throw new Error('Set NOUS_API_KEY in .env โ€” run `bun run setup` if you have not yet.')
}

// Where traces go. The project is created on first use.
export const logger = initLogger({ projectName: 'Happy Tails' })

export const llm = wrapOpenAI(
  new OpenAI({
    baseURL: process.env.NOUS_BASE_URL ?? 'https://inference-api.nousresearch.com/v1',
    apiKey,
  }),
)

export const MODEL = process.env.MODEL ?? 'anthropic/claude-haiku-4.5'

3. app/report.ts โ€” wrap the operation

Add import { traced } from 'braintrust' at the top, then wrap the existing body of writeReport:

  return traced(
    async (span) => {
      const res = await llm.chat.completions.create({
        model: MODEL,
        max_tokens: 500,
        messages: [
          { role: 'system', content: active.body },
          { role: 'user', content: JSON.stringify(dog, null, 2) },
        ],
      })

      const body = res.choices[0]?.message?.content ?? ''

      span.log({
        input: dog,
        output: body,
        metadata: { dog: dog.name, prompt_version: active.version, model: MODEL },
      })

      return { body, promptVersion: active.version, model: MODEL }
    },
    { name: 'daily report' },
  )

4. Generate some traffic

Restart the server and click all four dogs, using Regenerate for any that already have a report.

Check yourself

Open braintrust.dev โ†’ Happy Tails โ†’ Logs. You should have four traces named daily report. Expand one and find:

Now find Rufus's trace and read the output next to the input. Same fiction as on the website, except now it's a row you can point at, link to, and count.

Then: metadata is free now and impossible later

Ask Braintrust: show me only reports for dogs that were on medication. You can't. Nothing recorded it.

Add facts about the situation to that span.log call:

        metadata: {
          dog: dog.name,
          prompt_version: active.version,
          model: MODEL,
          has_meds: dog.meds.length > 0,
          has_incidents: dog.incidents.length > 0,
          play_minutes: dog.play.reduce((n, p) => n + p.minutes, 0),
          ate_nothing: dog.meals.every((m) => m.eaten_g === 0),
        },

Regenerate the four reports, then filter Logs on metadata.has_meds = true. Nacho, and only Nacho.

The habit worth keeping: log the situation, not just input and output. The question you'll want to answer six months from now is "did quality drop on the hard cases?" โ€” and it's only answerable if something back here marked which cases were hard.

Commit

git add -A && git commit -m "step 1: log every report to Braintrust"

What you can do now that you couldn't before

Read any report the app has ever written, with its cost, its latency, and the exact data it came from โ€” and filter down to the situations you care about.

What you still can't do: say whether any of it is good. Four traces, zero judgements. That's step 2: scoring one specific thing, in plain code, with no AI involved.