Skip to content
Gauge

ax-check.com

Braintrust

braintrust.dev

ax-check.com

Checked 7 October 2026

B

84/100

Agents need your login

Braintrust's quickstart, CLI, and pricing all work for agents out of the box.

21 of 23 checklist items passed: install commands, code samples, API/MCP docs, and openly stated pricing (Starter $0, Pro $249/mo) all came through clean.

2 fixes to start with

Concrete changes, backed by what we found.

Onboarding

Give agents a clear login handoff and a way to resume setup.

Agents cannot progress through onboarding without human help. Provide an agent-accessible path through authentication and setup.

Docs

Return Markdown when clients request text/markdown.

Homepage returned text/html for a text/markdown request; no Markdown representation offered.

Recorded coding sessions

Same task, 3 models, no braintrust.dev account. Recorded 7 October 2026 on Pi.

3 recorded · 2m 46s

2m 46s · recorded

Session findings

Select a finding to see it in context.

piv0.84.2
modelQwen 3.8 Maxcwd/sandbox/repo
Help me build a simple example using Braintrust. Tell me how pricing works, and briefly tell me whether this product will be easy for you to manage. Let me know if you get blocked. If this product has no developer workflow you can act on, say so plainly and stop. Stay light: use the hosted product through its SDK or API. Do not start local service stacks or wait for long-running commands; if the quickstart requires either, say so plainly and stop.
Thinking…

Let me look at the environment to see what's available.

The user asks: build a simple example using Braintrust, explain pricing, and say whether the product is easy for me to manage. Also: if there's no developer workflow I can act on, say so plainly and stop.

Let me check the sandbox for credentials, repo contents, etc.

ls -la /sandbox/repo 2>/dev/null | head -50; echo "---"; env | grep -iE "braintrust|api_key|token" | sed 's/=.*/=<set>/'
total 16
total 16
drwxr-xr-x 3 root root 4096 Sep 18 08:00 .
drwxr-xr-x 7 root root 4096 Oct  7 17:06 ..
drwxr-xr-x 8 root root 4096 Sep 24 20:05 .git
-rw-r--r-- 1 root root  408 Sep 18 08:00 .gitignore
---
PI_GATEWAY_API_KEY=<set>
cd /sandbox/repo 2>/dev/null && git log --oneline -5 2>/dev/null; echo "---"; python3 -c "import braintrust; print(braintrust.__version__)" 2>&1 | head -3; node -e "console.log(require.resolve('braintrust'))" 2>&1 | head -3; echo "---net---"; curl -s -o /dev/null -w "%{http_code}\n" --max-time 8 https://api.braintrust.dev/ 2>&1
9ebf5fb Initial workspace
9ebf5fb Initial workspace
---
Traceback (most recent call last):
  File "<string>", line 1, in <module>
ModuleNotFoundError: No module named 'braintrust'
node:internal/modules/cjs/loader:1433
  throw err;
  ^
---net---
200
Thinking…

Network access works (api.braintrust.dev returned 200). But there's no BRAINTRUST_API_KEY. So I cannot actually log data to Braintrust hosted product — I'd be blocked on credentials.

Let me verify: check if there's any config file, and test whether network to pypi/npm works. Also check whether the API key exists anywhere.

The key question: can I act on a developer workflow? Braintrust is a hosted product with SDK/API. The workflow is: pip install braintrust, write an eval with Eval() and @task, run braintrust eval. But running requires BRAINTRUST_API_KEY. Without it, the eval run fails to log.

Let me check network install ability and whether key exists.

ls -la ~ 2>/dev/null; cat ~/.braintrust* 2>/dev/null; ls /sandbox 2>/dev/null; env | sort | grep -viE "^(LS_COLORS|PATH|HOME|SHELL|TERM|PWD|SHLVL|_)" | sed 's/=.\{0,6\}.*/=<redacted>/' | head -30
total 28
total 28
drwxr-xr-x  7 root root 4096 Oct  7 17:06 .
drwxr-xr-x 22 root root 4096 Sep 18 07:59 ..
drwxr-xr-x  3 root root 4096 Oct  7 17:06 .cache
drwxr-xr-x  3 root root 4096 Oct  7 17:06 .gauge
drwxr-xr-x  3 root root 4096 Oct  7 17:06 .pi
drwxrwxrwt  2 root root 4096 Sep 18 08:00 .tmp
drwxr-xr-x  3 root root 4096 Sep 18 08:00 repo
repo
AI_AGENT=<redacted>
ALG_INFERENCE_MODEL=<redacted>
ALG_INFERENCE_PROVIDER=<redacted>
ALG_SANDBOX_PROVIDER=<redacted>
AWS_CA_BUNDLE=<redacted>
CARGO_HOME=<redacted>
CARGO_HTTP_CAINFO=<redacted>
CURL_CA_BUNDLE=<redacted>
DENO_CERT=<redacted>
GIT_CONFIG_COUNT=<redacted>
GIT_CONFIG_KEY_0=<redacted>
GIT_CONFIG_VALUE_0=<redacted>
GIT_SSL_CAINFO=<redacted>
GRPC_DEFAULT_SSL_ROOTS_FILE_PATH=<redacted>
INVOCATION_ID=<redacted>
JOURNAL_STREAM=<redacted>
LANG=<redacted>
LOGNAME=<redacted>
NODE_COMPILE_CACHE=<redacted>
NODE_EXTRA_CA_CERTS=<redacted>
OLDPWD=<redacted>
PIP_CERT=<redacted>
PI_CODING_AGENT=<redacted>
PI_GATEWAY_API_KEY=<redacted>
PI_MODEL=<redacted>
PI_PROVIDER=<redacted>
PI_REASONING_LEVEL=<redacted>
PI_SESSION_FILE=<redacted>
PI_SESSION_ID=<redacted>
REQUESTS_CA_BUNDLE=<redacted>
cd /tmp && timeout 60 pip download braintrust==0.0.* -d /tmp/btdl --no-deps -q 2>&1 | tail -3; ls /tmp/btdl 2>/dev/null; echo "--- curl auth probe (no key) ---"; curl -s --max-time 8 -w "\nHTTP %{http_code}\n" https://api.braintrust.dev/v1/ping
braintrust-0.0.203-py3-none-any.whl
braintrust-0.0.203-py3-none-any.whl
--- curl auth probe (no key) ---
<!DOCTYPE html><html class="btsans_2ea6ee4b-module__estiUG__variable inter_55581303-module__j1eaMG__variable suissemono_c6b80dcc-module__hr7OOa__variable h-full antialiased" lang="en"><head><meta charSet="utf-8"/><meta name="viewport" content="width=device-width, initial-scale=1"/><link rel="stylesheet" href="/_next/static/chunks/3vp37yl1n37l1.css" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi" data-precedence="next"/><link rel="stylesheet" href="/_next/static/chunks/1n05e8tr62xip.css" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi" data-precedence="next"/><link rel="preload" as="script" fetchPriority="low" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi" href="/_next/static/chunks/39uyso5_rcuh3.js"/><script src="/_next/static/chunks/2dxbpasd181jo.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/3y5im_vduyaso.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/2oonhe4-ijz8n.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/2iupwzj-vr1cu.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/3-2w3t6a6vz42.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/2d6m-ar4t0hhi.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/3__ctr6ufwmpu.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/28wyskek06jf4.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/0amhsd9s_lfy6.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/turbopack-01ha54usn9or2.js" async="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/17qqmm2_l7fif.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/0qg88v7idv8l4.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/2eo0w-i-5mhzl.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/3kh14c-t_09xz.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/3er5s5e399z8l.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/2omx0z-4mx4ce.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/3sz2rzd8pl_xl.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/29q_1jxfooupk.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/1tkr8-iwxrxo5.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/2o6-su7mxtpc1.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/2gfgvjko02s7b.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/0n4jjbk_sggjh.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/0eummq6hg4jr_.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/15f8l08r6z715.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/41_zcgdadwyq1.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/09-mgoufhk-zt.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="/_next/static/chunks/1jjfyrama_cbg.js" async="" crossorigin="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script src="https://clerk.braintrust.dev/npm/@clerk/clerk-js@6/dist/clerk.browser.js" data-clerk-js-script="true" async="" crossorigin="anonymous" data-clerk-publishable-key="pk_live_Y2xlcmsuYnJhaW50cnVzdC5kZXYk" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><link rel="preload" href="https://clerk.braintrust.dev/npm/@clerk/ui@1/dist/ui.browser.js" as="script" crossorigin="anonymous" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"/><meta name="robots" content="noindex"/><meta name="theme-color" content="#ddd" media="(prefers-color-scheme: light)"/><meta name="theme-color" content="#222" media="(prefers-color-scheme: dark)"/><title>Braintrust - The active observability platform for agents</title><meta name="keywords" content="AI evaluation,LLM evaluation,AI testing,LLM testing,AI observability,LLM observability,AI monitoring,LLM monitoring,AI debugging,LLM debugging,AI development platform,LLM development platform,AI evaluation framework,LLM evaluation framework,AI performance testing,LLM performance testing,braintrust,openai,anthropic,claude,gpt,llm ops,mlops,ai ops,prompt engineering,prompt testing,ai metrics,llm metrics,ai analytics,llm analytics"/><meta name="robots" content="index, follow"/><meta name="googlebot" content="index, follow, max-video-preview:-1, max-image-preview:large, max-snippet:-1"/><link rel="canonical" href="https://www.braintrust.dev"/><meta property="og:title" content="Braintrust - The active observability platform for agents"/><meta property="og:url" content="https://www.braintrust.dev"/><meta property="og:site_name" content="Braintrust"/><meta property="og:locale" content="en_US"/><meta property="og:image" content="https://www.braintrust.dev/og?title=Braintrust+-+The+active+observability+platform+for+agents&amp;v=3"/><meta property="og:type" content="website"/><meta name="twitter:card" content="summary_large_image"/><meta name="twitter:site" content="@braintrustdata"/><meta name="twitter:creator" content="@braintrustdata"/><meta name="twitter:title" content="Braintrust - The active observability platform for agents"/><meta name="twitter:image" content="https://www.braintrust.dev/og?title=Braintrust+-+The+active+observability+platform+for+agents&amp;v=3"/><link rel="icon" href="/icon.png?v=2" media="(prefers-color-scheme: light)" type="image/x-icon"/><link rel="icon" href="/icon-dark.png?v=2" media="(prefers-color-scheme: dark)" type="image/x-icon"/><link rel="apple-touch-icon" href="/icon180.png?v=2" media="(prefers-color-scheme: light)"/><link rel="apple-touch-icon" href="/icon-dark180.png?v=2" media="(prefers-color-scheme: dark)"/><meta name="next-size-adjust" content=""/><link rel="preload" href="https://clerk.braintrust.dev/npm/@clerk/ui@1/dist/ui.browser.js" as="script" crossorigin="anonymous" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"/><script nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi">
  (() => {
    const featureFlags = (() => {
      try {
        return JSON.parse(
          localStorage.getItem("app::featureFlags") ?? "{}",
        );
      } catch {
        return {};
      }
    })();
    if (featureFlags?.enableProductionReactDevTools === true) {
      return;
    }

    const key = "__REACT_DEVTOOLS_GLOBAL_HOOK__";
    const hook = window[key];
    if (hook) {
      hook.isDisabled = true;
      return;
    }

    Object.defineProperty(window, key, {
      value: { isDisabled: true },
      configurable: false,
    });
  })();
</script><script src="/_next/static/chunks/0cz1d0mv5g_q7.js" noModule="" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi"></script><script nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi">(function(){try{var c={"attribute":"class","storageKey":"theme","defaultTheme":"system","forcedTheme":null,"enableSystem":true,"themes":["light","dark"]};var d=document.documentElement;var t=c.forcedTheme;if(!t){try{t=localStorage.getItem(c.storageKey);}catch(e){}}if(!t)t=c.defaultTheme;var r=t;if(t==="system"){if(c.enableSystem){r=window.matchMedia("(prefers-color-scheme: dark)").matches?"dark":"light";}else{r=c.defaultTheme==="system"?"light":c.defaultTheme;}}if(c.attribute==="class"){d.classList.remove.apply(d.classList,c.themes);d.classList.add(r);}else{d.setAttribute(c.attribute,r);}if(r==="light"||r==="dark"){d.style.colorScheme=r;}}catch(e){}})();</script></head><body class="h-full"><div hidden=""><!--$--><!--/$--></div><!--$--><!--/$--><!--$--><!--/$--><main class="relative isolate flex min-h-[100svh] w-full items-center justify-center overflow-hidden bg-background px-4 font-display text-primary-950"><!--$!--><template data-dgst="BAILOUT_TO_CLIENT_SIDE_RENDERING"></template><!--/$--><section class="pointer-events-none relative z-10 flex flex-col items-center gap-8 text-center"><h1 class="text-7xl leading-none font-semibold sm:text-8xl"><span class="box-decoration-clone bg-background">404</span></h1><p class="max-w-md text-pretty text-xl leading-relaxed"><span class="box-decoration-clone bg-background px-2 py-1">The page you’re looking for doesn’t exist.</span></p><a class="inline-flex flex-none items-center justify-center cursor-pointer font-medium disabled:pointer-events-none disabled:opacity-60 relative gap-2 disabled:bg-primary-50 border border-primary-200 h-11 min-w-11 px-8 group pointer-events-auto z-10 rounded-none border-none bg-black dark:bg-white text-lg transition-[border-radius] duration-200 hover:rounded-lg hover:bg-black hover:dark:bg-white! text-white hover:text-white! dark:text-black dark:hover:text-black! select-none" href="/"><svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-arrow-left flex-none size-4 transition-transform duration-200 group-hover:-translate-x-0.5" aria-hidden="true"><path d="m12 19-7-7 7-7"></path><path d="M19 12H5"></path></svg>Go back</a></section></main><!--$--><!--/$--><script src="/_next/static/chunks/39uyso5_rcuh3.js" nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi" id="_R_" async=""></script><script nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi">(self.__next_f=self.__next_f||[]).push([0])</script><script nonce="N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi">self.__next_f.push([1,"1:\"$Sreact.fragment\"\n3:I[288662,[\"/_next/static/chunks/17qqmm2_l7fif.js\",\"/_next/static/chunks/0qg88v7idv8l4.js\",\"/_next/static/chunks/2eo0w-i-5mhzl.js\",\"/_next/static/chunks/3kh14c-t_09xz.js\",\"/_next/static/chunks/3er5s5e399z8l.js\",\"/_next/static/chunks/2omx0z-4mx4ce.js\",\"/_next/static/chunks/3sz2rzd8pl_xl.js\",\"/_next/static/chunks/29q_1jxfooupk.js\",\"/_next/static/chunks/1tkr8-iwxrxo5.js\",\"/_next/static/chunks/2o6-su7mxtpc1.js\",\"/_next/static/chunks/2gfgvjko02s7b.js\",\"/_next/static/chunks/0n4jjbk_sggjh.js\",\"/_next/static/chunks/0eummq6hg4jr_.js\",\"/_next/static/chunks/15f8l08r6z715.js\",\"/_next/static/chunks/41_zcgdadwyq1.js\"],\"default\"]\n4:I[722773,[\"/_next/static/chunks/17qqmm2_l7fif.js\",\"/_next/static/chunks/0qg88v7idv8l4.js\",\"/_next/static/chunks/2eo0w-i-5mhzl.js\",\"/_next/static/chunks/3kh14c-t_09xz.js\",\"/_next/static/chunks/3er5s5e399z8l.js\",\"/_next/static/chunks/2omx0z-4mx4ce.js\",\"/_next/static/chunks/3sz2rzd8pl_xl.js\",\"/_next/static/chunks/29q_1jxfooupk.js\",\"/_next/static/chunks/1tkr8-iwxrxo5.js\",\"/_next/static/chunks/2o6-su7mxtpc1.js\",\"/_next/static/chunks/2gfgvjko02s7b.js\",\"/_next/static/chunks/0n4jjbk_sggjh.js\",\"/_next/static/chunks/0eummq6hg4jr_.js\",\"/_next/static/chunks/15f8l08r6z715.js\",\"/_next/static/chunks/41_zcgdadwyq1.js\"],\"default\"]\n6:I[80121,[\"/_next/static/chunks/17qqmm2_l7fif.js\",\"/_next/static/chunks/0qg88v7idv8l4.js\",\"/_next/static/chunks/2eo0w-i-5mhzl.js\",\"/_next/static/chunks/3kh14c-t_09xz.js\",\"/_next/static/chunks/3er5s5e399z8l.js\",\"/_next/static/chunks/2omx0z-4mx4ce.js\",\"/_next/static/chunks/3sz2rzd8pl_xl.js\",\"/_next/static/chunks/29q_1jxfooupk.js\",\"/_next/static/chunks/1tkr8-iwxrxo5.js\",\"/_next/static/chunks/2o6-su7mxtpc1.js\",\"/_next/static/chunks/2gfgvjko02s7b.js\",\"/_next/static/chunks/0n4jjbk_sggjh.js\",\"/_next/static/chunks/0eummq6hg4jr_.js\",\"/_next/static/chunks/15f8l08r6z715.js\",\"/_next/static/chunks/41_zcgdadwyq1.js\"],\"OutletBoundary\"]\n7:\"$Sreact.suspense\"\na:I[80121,[\"/_next/static/chunks/17qqmm2_l7fif.js\",\"/_next/static/chunks/0qg88v7idv8l4.js\",\"/_next/static/chunks/2eo0w-i-5mhzl.js\",\"/_next/static/chunks/3kh14c-t_09xz.js\",\"/_next/static/chunks/3er5s5e399z8l.js\",\"/_next/static/chunks/2omx0z-4mx4ce.js\",\"/_next/static/chunks/3sz2rzd8pl_xl.js\",\"/_next/static/chunks/29q_1jxfooupk.js\",\"/_next/static/chunks/1tkr8-iwxrxo5.js\",\"/_next/static/chunks/2o6-su7mxtpc1.js\",\"/_next/static/chunks/2gfgvjko02s7b.js\",\"/_next/static/chunks/0n4jjbk_sggjh.js\",\"/_next/static/chunks/0eummq6hg4jr_.js\",\"/_next/static/chunks/15f8l08r6z715.js\",\"/_next/static/chunks/41_zcgdadwyq1.js\"],\"ViewportBoundary\"]\nc:I[80121,[\"/_next/static/chunks/17qqmm2_l7fif.js\",\"/_next/static/chunks/0qg88v7idv8l4.js\",\"/_next/static/chunks/2eo0w-i-5mhzl.js\",\"/_next/static/chunks/3kh14c-t_09xz.js\",\"/_next/static/chunks/3er5s5e399z8l.js\",\"/_next/static/chunks/2omx0z-4mx4ce.js\",\"/_next/static/chunks/3sz2rzd8pl_xl.js\",\"/_next/static/chunks/29q_1jxfooupk.js\",\"/_next/static/chunks/1tkr8-iwxrxo5.js\",\"/_next/static/chunks/2o6-su7mxtpc1.js\",\"/_next/static/chunks/2gfgvjko02s7b.js\",\"/_next/static/chunks/0n4jjbk_sggjh.js\",\"/_next/static/chunks/0eummq6hg4jr_.js\",\"/_next/static/chunks/15f8l08r6z715.js\",\"/_next/static/chunks/41_zcgdadwyq1.js\"],\"MetadataBoundary\"]\ne:I[185143,[\"/_next/static/chunks/17qqmm2_l7fif.js\",\"/_next/static/chunks/0qg88v7idv8l4.js\",\"/_next/static/chunks/2eo0w-i-5mhzl.js\",\"/_next/static/chunks/3kh14c-t_09xz.js\",\"/_next/static/chunks/3er5s5e399z8l.js\",\"/_next/static/chunks/2omx0z-4mx4ce.js\",\"/_next/static/chunks/3sz2rzd8pl_xl.js\",\"/_next/static/chunks/29q_1jxfooupk.js\",\"/_next/static/chunks/1tkr8-iwxrxo5.js\",\"/_next/static/chunks/2o6-su7mxtpc1.js\",\"/_next/static/chunks/2gfgvjko02s7b.js\",\"/_next/static/chunks/0n4jjbk_sggjh.js\",\"/_next/static/chunks/0eummq6hg4jr_.js\",\"/_next/static/chunks/15f8l08r6z715.js\",\"/_next/static/chunks/41_zcgdadwyq1.js\",\"/_next/static/chunks/09-mgoufhk-zt.js\"],\"default\"]\n:HL[\"/_next/static/chunks/3vp37yl1n37l1.css\",\"style\",{\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}]\n:HL[\"/_next/static/chunks/1n05e8tr62xip.css\",\"style\",{\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}]\n9:X\n0:{\"P\":null,\"c\":[\"\",\"api\",\"ping\",\"get\"],\"q\":\"\",\"i\":false,\"f\":[[[\"\",{\"children\":[\"/_not-found\",{\"children\":[\"__PAGE__\",{},\"$undefined\",\"$undefined\",4096]},\"$undefined\",\"$undefined\",4096]},\"$undefined\",\"$undefined\",4112],[[\"$\",\"$1\",\"c\",{\"children\":[[[\"$\",\"link\",\"0\",{\"rel\":\"stylesheet\",\"href\":\"/_next/static/chunks/3vp37yl1n37l1.css\",\"precedence\":\"next\",\"crossOrigin\":\"$undefined\",\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"link\",\"1\",{\"rel\":\"stylesheet\",\"href\":\"/_next/static/chunks/1n05e8tr62xip.css\",\"precedence\":\"next\",\"crossOrigin\":\"$undefined\",\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-0\",{\"src\":\"/_next/static/chunks/17qqmm2_l7fif.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-1\",{\"src\":\"/_next/static/chunks/0qg88v7idv8l4.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-2\",{\"src\":\"/_next/static/chunks/2eo0w-i-5mhzl.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-3\",{\"src\":\"/_next/static/chunks/3kh14c-t_09xz.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-4\",{\"src\":\"/_next/static/chunks/3er5s5e399z8l.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-5\",{\"src\":\"/_next/static/chunks/2omx0z-4mx4ce.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-6\",{\"src\":\"/_next/static/chunks/3sz2rzd8pl_xl.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-7\",{\"src\":\"/_next/static/chunks/29q_1jxfooupk.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-8\",{\"src\":\"/_next/static/chunks/1tkr8-iwxrxo5.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-9\",{\"src\":\"/_next/static/chunks/2o6-su7mxtpc1.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-10\",{\"src\":\"/_next/static/chunks/2gfgvjko02s7b.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-11\",{\"src\":\"/_next/static/chunks/0n4jjbk_sggjh.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-12\",{\"src\":\"/_next/static/chunks/0eummq6hg4jr_.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-13\",{\"src\":\"/_next/static/chunks/15f8l08r6z715.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"script\",\"script-14\",{\"src\":\"/_next/static/chunks/41_zcgdadwyq1.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}]],\"$L2\"]}],{\"children\":[[\"$\",\"$1\",\"c\",{\"children\":[null,[\"$\",\"$L3\",null,{\"parallelRouterKey\":\"children\",\"error\":\"$undefined\",\"errorStyles\":\"$undefined\",\"errorScripts\":\"$undefined\",\"template\":[\"$\",\"$L4\",null,{}],\"templateStyles\":\"$undefined\",\"templateScripts\":\"$undefined\",\"notFound\":\"$undefined\",\"forbidden\":\"$undefined\",\"unauthorized\":\"$undefined\"}]]}],{\"children\":[[\"$\",\"$1\",\"c\",{\"children\":[\"$L5\",[[\"$\",\"script\",\"script-0\",{\"src\":\"/_next/static/chunks/1jjfyrama_cbg.js\",\"async\":true,\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}]],[\"$\",\"$L6\",null,{\"children\":[\"$\",\"$7\",null,{\"name\":\"Next.MetadataOutlet\",\"children\":\"$@8\"}]}]]}],{},null,false,null]},null,false,\"$9\"]},null,false,null],[\"$\",\"$1\",\"h\",{\"children\":[[\"$\",\"meta\",null,{\"name\":\"robots\",\"content\":\"noindex\"}],[\"$\",\"$La\",null,{\"children\":\"$Lb\"}],[\"$\",\"div\",null,{\"hidden\":true,\"children\":[\"$\",\"$Lc\",null,{\"children\":[\"$\",\"$7\",null,{\"name\":\"Next.Metadata\",\"children\":\"$Ld\"}]}]}],[\"$\",\"meta\",null,{\"name\":\"next-size-adjust\",\"content\":\"\"}]]}],false]],\"m\":\"$undefined\",\"G\":[\"$e\",[[\"$\",\"link\",\"0\",{\"rel\":\"stylesheet\",\"href\":\"/_next/static/chunks/3vp37yl1n37l1.css\",\"precedence\":\"next\",\"crossOrigin\":\"$undefined\",\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}],[\"$\",\"link\",\"1\",{\"rel\":\"stylesheet\",\"href\":\"/_next/static/chunks/1n05e8tr62xip.css\",\"precedence\":\"next\",\"crossOrigin\":\"$undefined\",\"nonce\":\"N2U5ZTgyZWYtNWYzNi00MGUxLWExMGQtNzA1OWVhNWJlYzhi\"}]]],\"S\":false,\"h\":n

Showing the first 20,000 characters. Full output is available in raw details.

Thinking…

So: no BRAINTRUST_API_KEY available. Network works, pip download works. Braintrust's hosted product requires an API key from a Braintrust account (created via web UI, Clerk auth, requires email signup). I can't create an account.

The question: "If this product has no developer workflow you can act on, say so plainly and stop."

Braintrust DOES have a developer workflow (SDK evals, API). But acting on it end-to-end requires an account + API key, which I cannot obtain (requires interactive web signup with email). However, I could still write the example code and verify it locally... Actually the braintrust SDK supports running evals locally without an API key? Let me check. There is BRAINTRUST_API_KEY required for logging. But there's also "local" mode? Historically braintrust required login. There is an option --no-send-logs / send_logs=False and BTS_*? Let me check the SDK's behavior: Eval(..., send_logs=False) will run locally and print results without uploading. Actually yes, braintrust supports running evals without an API key if send_logs=False — it prints a summary in the terminal. Let me verify by installing and trying.

But wait — the task also says "use the hosted product through its SDK or API". Without an API key I can't send anything to the hosted product. So I'd be blocked on the hosted part.

The honest approach:

  1. Build the example (install SDK, write eval file, run it locally with send_logs=False to prove it works).
  2. Report blocked: cannot log to hosted Braintrust without BRAINTRUST_API_KEY, which requires creating an account through the web UI (Clerk email signup) — I can't do that.
  3. Explain pricing (from knowledge; maybe fetch the pricing page to be accurate).
  4. Manageability assessment.

Let me be efficient. First, let me check pricing page content via curl, and install braintrust SDK.

Let me install braintrust and check whether offline local run works.

Actually the LLM calls: the eval needs a model. There's PI_GATEWAY_API_KEY... but not an OpenAI key. I can write a deterministic task (no LLM) so the example runs without model keys. Braintrust evals support that.

Let's do it.

cd /tmp && timeout 180 pip install -q braintrust 2>&1 | tail -3; python3 -c "import braintrust,os;print(braintrust.__version__ if hasattr(braintrust,'__version__') else 'ok')"; pip show braintrust 2>/dev/null | head -3
ok
ok
Name: braintrust
Version: 0.44.1
Summary: SDK for integrating Braintrust
curl -s --max-time 15 https://www.braintrust.dev/pricing | python3 -c "
import sys,re,html
t=sys.stdin.read()
t=re.sub(r'<script.*?</script>','',t,flags=re.S)
t=re.sub(r'<style.*?</style>','',t,flags=re.S)
t=re.sub(r'<[^>]+>',' ',t)
t=html.unescape(t)
t=re.sub(r'\s+',' ',t)
print(t[:4000])
"
Pricing - Braintrust Product Resources Customers Pricing Contact us Sign in Sign up Observe Trace everything Evaluate Test what ships Discover Find behaviors D
 Pricing - Braintrust Product Resources Customers Pricing Contact us Sign in Sign up Observe Trace everything Evaluate Test what ships Discover Find behaviors Docs How Braintrust works Blog Latest announcements Workshops Online events Eval Research Open-source studies Academy Learn to eval Encyclopedia Eval dictionary Product Observe Evaluate Discover Resources Docs Blog Workshops Eval Research Academy Encyclopedia Customers Pricing Contact us Sign in Sign up Predictable pricing. Designed to scale. Start now. Instrument when you’re ready. No credit card required. Free usage is included for traces, evals, and your whole team. Sign up Start building for free Model credits Processed data Scores Data retention Features Starter For everyone $0 / month $10 credits + tok rates 1 GB processed data + $4/GB 10k scores + $2.50/1k 14-day retention Unlimited users, projects, datasets, playgrounds, and experiments Sign up No credit card required Pro For AI native teams $249 / month $100 credits + tok rates 5 GB processed data + $3/GB 50k scores + $1.50/1k 30-day retention + $0.50/GB/mo Custom charts, environments, priority support, RBAC, and more Get Pro 6–12 months free for qualifying startups Enterprise For teams at scale Custom pricing Custom data retention and export, RBAC, and premium support with on-prem or hosted deployment for high volume or privacy-sensitive data. Contact us Trusted by the best AI teams View detailed pricing, including model and token rates Starter for everyone Sign up No credit card required Pro for AI native teams Get Pro 6–12 months F ree for qualifying startups Enterprise for teams at scale Contact us Core platform Tracing, evaluation, and storage infrastructure for your AI stack. Platform fee $0 / month $249 / month Custom Model credits Your monthly model credit covers AI features like Topics, as well as any built-in models you use without your own API key. $10 credits / month included then token rates $100 credits / month included then token rates Custom then token rates Processed data 1 GB / month included then +$4/GB 5 GB / month included then +$3/GB Custom Scores Scored outputs from LLM-as-a-judge, autoevals, or custom code scorers. 10K / month included then $2.50 per 1,000 50K / month included then $1.50 per 1,000 Custom Users Unlimited Unlimited Unlimited Projects Unlimited Unlimited Unlimited Playgrounds Unlimited Unlimited Unlimited Experiments Unlimited Unlimited Unlimited Datasets Unlimited Unlimited Unlimited Human review scores Configurable human annotation scores that your team fills in manually when reviewing traces. 1 per project Unlimited Unlimited Loop agent Braintrust's built-in AI agent that can run evaluations, generate test cases, and iterate on prompts autonomously. Playground annotations Provide feedback on playground outputs to quickly optimize prompts. ﹣ Customization Dashboards, custom charts, and views to understand your system in production. Saved table views Unlimited Unlimited Unlimited Custom columns Unlimited Unlimited Unlimited Custom trace views Unlimited Unlimited Unlimited Custom charts Build charts and dashboards using custom metrics, score aggregations, and usage data from your projects ﹣ Environments Tag object versions with production, staging, development, or a custom tag to make it easy to pull for specific environments ﹣ Data retention and export Control how long your data is stored and how it leaves Braintrust. Included data retention Up to 180-day data retention, billed at $0.50 per GB per month after the included 30-day period 14-day retention 30-day retention then $0.50 per GB per month Custom Custom policies Define per-project retention rules — keep certain data longer or purge it sooner based on your compliance requirements ﹣ ﹣ S3 data export Automatically export traces to your own S3 bucket for long-term storage or downstream analysis ﹣ ﹣ Security and compliance Enterprise-grade access controls, authentication, and compliance agreements. OAuth sign-in Multi-fac
Thinking…

Good, I have pricing. Now build the example. Check whether eval can run without API key (local / send_logs=False).

Let me write the example in /sandbox/repo.

File edits

1 file

Completed
Writtenbraintrust_demo/evals/support_bot.pyView changes

/sandbox/repo/braintrust_demo/evals/support_bot.py

Written content · previous contents were not recorded

"""Minimal Braintrust eval example.
Run:
python evals/support_bot.py # local run, no upload
BRAINTRUST_API_KEY=... python evals/support_bot.py # upload to Braintrust cloud
The task below is deterministic (no LLM call) so the example runs with zero
model credentials. Swap in a real model call inside `classify()` when ready.
"""
import os
from braintrust import Eval, init_logger, task
from autoevals import Levenshtein
# Set send_logs=False and the SDK runs the eval locally and prints a summary
# instead of uploading spans to the hosted project.
SEND_LOGS = bool(os.environ.get("BRAINTRUST_API_KEY"))
@task
def classify(message: str) -> str:
"""Stand-in for an LLM classifier: keyword rules -> intent label."""
m = message.lower()
if "refund" in m:
return "refund"
if "password" in m or "log in" in m or "login" in m:
return "account_access"
if "invoice" in m or "charged" in m:
return "billing"
return "other"
def dataset():
return [
{
"input": "I want a refund for order #1234",
"expected": "refund",
"metadata": {"channel": "email"},
},
{
"input": "I can't log in after resetting my password",
"expected": "account_access",
"metadata": {"channel": "chat"},
},
{
"input": "Why was I charged twice on my invoice?",
"expected": "billing",
"metadata": {"channel": "email"},
},
{
"input": "What colors does the jacket come in?",
"expected": "other",
"metadata": {"channel": "chat"},
},
]
def score(output: str, expected: str) -> float:
# Exact match first; fall back to string distance so partial credit is visible.
return 1.0 if output == expected else round(Levenshtein()(output, expected), 3)
if __name__ == "__main__":
Eval(
"support-bot-demo", # project name
data=dataset, # list of {input, expected, metadata}
task=classify, # function(input) -> output
scores=[score], # function(output, expected) -> 0..1
experiment_name=None, # auto-generated timestamp name
send_logs=SEND_LOGS,
max_concurrency=4,
)
if not SEND_LOGS:
print(
"\n(no BRAINTRUST_API_KEY set -> ran locally; "
"set it to upload this experiment to Braintrust)"
)
cd /sandbox/repo && ls; python3 -c "import autoevals; print('autoevals ok')" 2>&1 | tail -2
braintrust_demo
braintrust_demo
  File "<string>", line 1, in <module>
ModuleNotFoundError: No module named 'autoevals'
Thinking…

autoevals not installed — it's a dependency? Apparently not auto-installed. Install it. Then run.

Also remove unused init_logger import.

cd /tmp && timeout 180 pip install -q autoevals 2>&1 | tail -2; cd /sandbox/repo && timeout 180 python braintrust_demo/evals/support_bot.py 2>&1 | tail -30
Traceback (most recent call last):
Traceback (most recent call last):
  File "/sandbox/repo/braintrust_demo/evals/support_bot.py", line 13, in <module>
    from braintrust import Eval, task
ImportError: cannot import name 'task' from 'braintrust' (/opt/freestyle/python/lib/python3.12/site-packages/braintrust/__init__.py)

File edits

1 file

Completed
Modifiedbraintrust_demo/evals/support_bot.py+1−1View changes

/sandbox/repo/braintrust_demo/evals/support_bot.py

Replacement excerpt

-from braintrust import Eval, init_logger, task
+from braintrust import Eval, task
Thinking…

In braintrust 0.44.1, the API changed: Eval is imported from braintrust but task decorator may be gone — the new API uses @braintrust.traced and just plain functions. Let me inspect the package exports.

python3 - <<'EOF'
import braintrust
names=[n for n in dir(braintrust) if not n.startswith('_')]
print(len(names))
print([n for n in names if any(k in n.lower() for k in ('eval','task','trace','span','log','dataset','init'))])
EOF
466
466
['BACKGROUND_LOGGER_BASE_SLEEP_TIME_S', 'BraintrustLogHandler', 'DEFAULT_IS_LEGACY_DATASET', 'Dataset', 'DatasetEvent', 'DatasetPipeline', 'DatasetPipelineDefinition', 'DatasetPipelineRow', 'DatasetPipelineScope', 'DatasetPipelineSource', 'DatasetPipelineTarget', 'DatasetPipelineTransform', 'DatasetPipelineTransformArgs', 'DatasetPipelineTransformResult', 'DatasetRef', 'DatasetSnapshot', 'DatasetSummary', 'Eval', 'EvalAsync', 'EvalCase', 'EvalHooks', 'EvalParameters', 'Evaluator', 'ExperimentDatasetEvent', 'ExperimentDatasetIterator', 'GitMetadataSettings', 'LOGS3_OVERFLOW_REFERENCE_TYPE', 'LogItemWithMeta', 'LogLevel', 'LogQueue', 'Logger', 'LoginResult', 'Logs3OverflowInputRow', 'NOOP_SPAN', 'NOOP_SPAN_PERMALINK', 'ParentSpanIds', 'ProjectDatasetMetadata', 'ProjectLogsEvent', 'RemoteEvalParameters', 'RunEval', 'Span', 'SpanAttributes', 'SpanComponentsV3', 'SpanComponentsV4', 'SpanCustomizer', 'SpanExportData', 'SpanIFrame', 'SpanIds', 'SpanImpl', 'SpanInternalOptions', 'SpanObjectTypeV3', 'SpanOriginEnvironment', 'SpanScope', 'SpanType', 'SpanTypeAttribute', 'TRACEPARENT_HEADER', 'TRACESTATE_HEADER', 'ToolFunctionDefinition', 'TraceScope', 'TracebackType', 'TracedThreadPoolExecutor', 'construct_logs3_data', 'construct_logs3_overflow_request', 'current_logger', 'current_span', 'dataset_pipeline', 'ensure_dataset_record', 'extract_trace_context', 'format_traceparent', 'get_span_parent_object', 'init', 'init_dataset', 'init_experiment', 'init_function', 'init_logger', 'inject_trace_context', 'log', 'log_exc_info_to_span', 'logger', 'logging', 'login', 'login_lock', 'login_to_state', 'logs', 'merge_span_origin_context', 'parse_traceparent', 'pick_logs3_overflow_object_ids', 'set_span_customizers', 'span_cache', 'span_components_to_object_id', 'span_customizer', 'span_identifier_v1', 'span_identifier_v2', 'span_identifier_v3', 'span_identifier_v4', 'span_origin', 'span_types', 'split_logging_data', 'start_span', 'trace', 'traceback', 'traced', 'update_span']
Thinking…

So task no longer exists in 0.44.1; use plain functions or @traced. Let me check Eval signature (EvalParameters) for valid kwargs: data, task, scores, experiment_name, send_logs, max_concurrency.

Let me check EvalParameters fields.

python3 - <<'EOF'
from braintrust import EvalParameters
import dataclasses
if dataclasses.is_dataclass(EvalParameters):
    for f in dataclasses.fields(EvalParameters):
        print(f.name)
else:
    print(EvalParameters.__annotations__ if hasattr(EvalParameters,'__annotations__') else 'n/a')
EOF
{}
{}
Thinking…

No fields exposed. Let's check the Eval class init signature via inspect on the actual source. It's likely a compiled/bundled file. Let me grep the installed source for "send_logs" and "max_concurrency".

P=/opt/freestyle/python/lib/python3.12/site-packages/braintrust; ls $P | head; grep -rn "send_logs" $P --include=*.py -l | head; echo ---; grep -rn "def Eval" $P --include=*.py | head
__init__.py
__init__.py
__pycache__
_generated_types.py
_test_cassette_usage.py
api
audit.py
auto.py
aws.py
bt_json.py
btx
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/type_tests/test_eval_generics.py
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/type_tests/test_workflow_eval.py
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/type_tests/test_autoevals_scorers.py
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/type_tests/test_dict_scorers.py
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/test_framework.py
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/cli/eval.py
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/framework.py
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/workflow_eval.py
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/test_workflow_eval.py
---
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/framework.py:818:async def EvalAsync(
/opt/freestyle/python/lib/python3.12/site-packages/braintrust/framework.py:947:def Eval(
Thinking…

I need to look at the Eval signature and how the type tests are being used.

P=/opt/freestyle/python/lib/python3.12/site-packages/braintrust; sed -n '947,1010p' $P/framework.py; echo "=== example usage ==="; sed -n '1,60p' $P/type_tests/test_eval_generics.py
def Eval(
def Eval(
    name: str,
    data: EvalData[Input, Expected],
    task: EvalTask[Input, Output, Expected],
    scores: Sequence[EvalScorer[Input, Output, Expected]] | None = None,
    classifiers: Sequence[EvalClassifier[Input, Output, Expected]] | None = None,
    experiment_name: str | None = None,
    trial_count: int = 1,
    metadata: Metadata | None = None,
    tags: Sequence[str] | None = None,
    is_public: bool = False,
    update: bool = False,
    reporter: ReporterDef[Input, Output, Expected, EvalReport] | None = None,
    timeout: float | None = None,
    max_concurrency: int | None = None,
    project_id: str | None = None,
    base_experiment_name: str | None = None,
    base_experiment_id: str | None = None,
    git_metadata_settings: GitMetadataSettings | None = None,
    repo_info: RepoInfo | None = None,
    error_score_handler: ErrorScoreHandler[Input, Expected] | None = None,
    description: str | None = None,
    summarize_scores: bool = True,
    no_send_logs: bool = False,
    parameters: EvalParameters | RemoteEvalParameters | None = None,
    on_start: Callable[[ExperimentSummary], None] | None = None,
    stream: Callable[[SSEProgressEvent], None] | None = None,
    parent: str | None = None,
    state: BraintrustState | None = None,
    enable_cache: bool = True,
) -> EvalResultWithSummary[Input, Output, Expected]:
    """
    A function you can use to define an evaluator. This is a convenience wrapper around the `Evaluator` class.

    For callers running in an async context, use `EvalAsync()` instead.

    Example:
    ```python
    Eval(
        name="my-evaluator",
        data=lambda: [
            EvalCase(input=1, expected=2),
            EvalCase(input=2, expected=4),
        ],
        task=lambda input, hooks: input * 2,
        scores=[
            NumericDiff,
        ],
    )
    ```

    :param name: The name of the evaluator. This corresponds to a project name in Braintrust.
    :param data: Returns an iterator over the evaluation dataset. Each element of the iterator should be a `EvalCase`.
    :param task: Runs the evaluation task on a single input. The `hooks` object can be used to add metadata to the evaluation.
    :param scores: A list of scorers to evaluate the results of the task. Each scorer can be a Scorer object or a function
    that returns a number, a `Score` object, a `ScoreDict`, or a sequence of named scores.
    A single dict can omit `name` to use the scorer's name; each score in a sequence must include a name.
    :param experiment_name: (Optional) Experiment name. If not specified, a name will be generated automatically.
    :param trial_count: The number of times to run the evaluator per input. This is useful for evaluating applications that
    have non-deterministic behavior and gives you both a stronger aggregate measure and a sense of the variance in the results.
    :param metadata: (Optional) A dictionary with additional data about the test example, model outputs, or just about
    anything else that's relevant, that you can use to help find and analyze examples later. For example, you could log
    the `prompt`, example's `id`, or anything else that would be useful to slice/dice later. The values in `metadata`
    can be any JSON-serializable type, but its keys must be strings.
=== example usage ===
"""Type-check tests for the Eval framework generic parameters.

These tests verify that pyright/mypy accept valid usage patterns
and that the runtime behavior is correct.

Run as type checks:
    nox -s test_types

Run as pytest:
    pytest src/braintrust/type_tests/test_eval_generics.py
"""

from typing import TypedDict

import pytest
from braintrust.framework import EvalAsync, EvalCase, EvalResultWithSummary
from braintrust.generated_types import ObjectReference
from braintrust.score import Score
from braintrust.types._eval import EvalCaseDict, EvalCaseDictNoOutput


# --- Domain types for testing ---
class ModelOutput(TypedDict):
    answer: str
    confidence: float


class AssertionSpec:
    """Assertion specification — not the same type as the model output."""

    def __init__(self, field: str, expected_value: str):
        self.field = field
        self.expected_value = expected_value


# ============================================================
# Case 1: Same-type Output and Expected (should always work)
# ============================================================


def same_type_data():
    return iter([EvalCase(input="query", expected="golden answer")])


async def same_type_task(input: str) -> str:
    return "model answer"


async def same_type_scorer(input: str, output: str, expected: str | None = None) -> Score:
    return Score(name="match", score=1.0 if output == expected else 0.0)


# ============================================================
# Case 2: Divergent Output and Expected (the bug from #240)
# ============================================================


def divergent_data():
    return iter(
        [
Thinking…

The parameter is no_send_logs, not send_logs. Let me fix the example accordingly and use EvalCase dicts. Also scorer signature (input, output, expected). Let me update.

File edits

1 file

Completed
Writtenbraintrust_demo/evals/support_bot.pyView changes

/sandbox/repo/braintrust_demo/evals/support_bot.py

Written content · previous contents were not recorded

"""Minimal Braintrust eval example (braintrust SDK >= 0.44).
Run locally (no upload, no credentials needed):
python evals/support_bot.py
Run against the hosted Braintrust project:
export BRAINTRUST_API_KEY=... # from https://www.braintrust.dev/app -> API keys
python evals/support_bot.py
The task below is deterministic (no LLM call) so the example runs with zero
model credentials. Swap in a real model call inside `classify()` when ready.
"""
import os
from braintrust import Eval, EvalCase
from autoevals import Levenshtein
# Without an API key the SDK would fail to upload, so run in local/offline mode
# and just print the summary. With a key, logs go to the hosted project.
NO_SEND_LOGS = not os.environ.get("BRAINTRUST_API_KEY")
def classify(input: str) -> str:
"""Stand-in for an LLM classifier: keyword rules -> intent label."""
m = input.lower()
if "refund" in m:
return "refund"
if "password" in m or "log in" in m or "login" in m:
return "account_access"
if "invoice" in m or "charged" in m:
return "billing"
return "other"
def dataset():
return iter(
[
EvalCase(
input="I want a refund for order #1234",
expected="refund",
metadata={"channel": "email"},
),
EvalCase(
input="I can't log in after resetting my password",
expected="account_access",
metadata={"channel": "chat"},
),
EvalCase(
input="Why was I charged twice on my invoice?",
expected="billing",
metadata={"channel": "email"},
),
EvalCase(
input="What colors does the jacket come in?",
expected="other",
metadata={"channel": "chat"},
),
]
)
def intent_match(input: str, output: str, expected: str | None = None) -> float:
"""Exact match, with string-distance partial credit so gaps stay visible."""
if output == expected:
return 1.0
return round(Levenshtein()(output, expected), 3)
if __name__ == "__main__":
Eval(
name="support-bot-demo", # Braintrust project name
data=dataset, # iterator of {input, expected, metadata}
task=classify, # task(input) -> output
scores=[intent_match], # scorer(input, output, expected) -> 0..1
max_concurrency=4,
no_send_logs=NO_SEND_LOGS,
)
if NO_SEND_LOGS:
print(
"\n(no BRAINTRUST_API_KEY set -> ran offline; "
"set it to upload this experiment to the hosted project)"
)
Thinking…

I need to verify that autoevals is properly installed before proceeding.

cd /tmp && timeout 240 pip install -q autoevals 2>&1 | tail -3; python3 -c "import autoevals;print('autoevals',autoevals.__name__)"
autoevals autoevals
autoevals autoevals
Thinking…

I need to run the evaluation now.

cd /sandbox/repo && timeout 240 python braintrust_demo/evals/support_bot.py 2>&1 | tail -30
support-bot-demo (data): 0it [00:00, ?it/s] support-bot-demo (data): 4it [00:00, 43129.09it/s]
support-bot-demo (data): 0it [00:00, ?it/s]
support-bot-demo (data): 4it [00:00, 43129.09it/s]

support-bot-demo (tasks):   0%|          | 0/4 [00:00<?, ?it/s]
support-bot-demo (tasks): 100%|██████████| 4/4 [00:00<00:00, 1573.11it/s]

=========================SUMMARY=========================
100.00% 'intent_match' score



(no BRAINTRUST_API_KEY set -> ran offline; set it to upload this experiment to the hosted project)
Thinking…

It works locally. Now, blocked on hosted: no BRAINTRUST_API_KEY. Creating an account requires interactive web signup (Clerk, email/password + OAuth) — I can't do that. So I'm blocked on the hosted half.

Let me also add a tracing example? Keep light. Maybe add a small traced-logging snippet file showing production tracing (braintrust.traced / @traced decorator) — one small extra file, and a README with pricing. Keep it light.

Also maybe demonstrate the HTTP API shape for logging without SDK? That requires key too. I'll include a short README.

Let me quickly verify the @traced decorator exists and works offline for the second example, so I don't ship untested code.

File edits

1 file

Completed
Writtenbraintrust_demo/trace_demo.pyView changes

/sandbox/repo/braintrust_demo/trace_demo.py

Written content · previous contents were not recorded

"""Production-style tracing with the Braintrust SDK.
python trace_demo.py # offline, spans are no-ops
BRAINTRUST_API_KEY=... python trace_demo.py # spans land in the hosted project
`@traced` turns any function into a span; nested calls nest automatically.
"""
import os
from braintrust import traced
NO_SEND_LOGS = not os.environ.get("BRAINTRUST_API_KEY")
@traced(name="retrieve_context")
def retrieve_context(ticket_id: str) -> str:
return f"order {ticket_id}: shipped 2024-05-01, refundable"
@traced(name="draft_reply")
def draft_reply(ticket_id: str) -> str:
context = retrieve_context(ticket_id) # nested span
return f"Thanks — here's what I found: {context}"
if __name__ == "__main__":
# `project` + `send_logs=False` keeps this local without a key.
with traced(
name="support_ticket",
project="support-bot-demo",
send_logs=not NO_SEND_LOGS,
) as span:
reply = draft_reply("1234")
span.log(
input={"ticket_id": "1234"},
output=reply,
metrics={"prompt_tokens": 120, "completion_tokens": 45},
scores={"helpfulness": 1.0},
)
print(reply)
if NO_SEND_LOGS:
print("(no BRAINTRUST_API_KEY -> spans were not uploaded)")
Thinking…

I need to run it.

cd /sandbox/repo && timeout 120 python braintrust_demo/trace_demo.py 2>&1 | tail -20
Traceback (most recent call last):
Traceback (most recent call last):
  File "/sandbox/repo/braintrust_demo/trace_demo.py", line 29, in <module>
    with traced(
TypeError: 'functools.partial' object does not support the context manager protocol
Thinking…

traced is a decorator factory only. The context manager is start_span or init_logger().start_span(). Let's check start_span signature.

P=/opt/freestyle/python/lib/python3.12/site-packages/braintrust; grep -n "^def start_span\|^def traced" $P/logger.py | head; python3 -c "
import inspect, braintrust
print(inspect.signature(braintrust.start_span))
" 2>&1 | head -5
2914:def traced(f: F) -> F:
2914:def traced(f: F) -> F:
2919:def traced(*span_args: Any, **span_kwargs: Any) -> Callable[[F], F]:
2923:def traced(*span_args: Any, **span_kwargs: Any) -> Callable[[F], F]:
3075:def start_span(
(name: str | None = None, type: braintrust.span_types.SpanTypeAttribute | None = None, span_attributes: braintrust._generated_types.SpanAttributes | collections.abc.Mapping[str, typing.Any] | None = None, start_time: float | None = None, set_current: bool | None = None, parent: str | dict[str, str] | None = None, propagated_event: dict[str, typing.Any] | None = None, state: braintrust.logger.BraintrustState | None = None, internal: braintrust.logger.SpanInternalOptions | None = None, **event: Any) -> braintrust.logger.Span
Thinking…

Use init_logger(project=..., api_key=...) then logger.start_span(name=..., ...). But without a key, init_logger would attempt to fetch project id → fail. Let's check: braintrust requires login for logging. In offline mode, init_logger may raise. Let me test quickly with a script that tries init_logger without key and catch the error, to see what happens.

Actually simpler: keep the trace demo as pure decorator usage plus a note that uploading requires init. Let me test init_logger without key quickly (short timeout).

cd /tmp && timeout 60 python - <<'EOF' 2>&1 | tail -20
from braintrust import init_logger
try:
    lg = init_logger(project="support-bot-demo")
    with lg.start_span(name="support_ticket") as s:
        s.log(input={"a":1}, output="hi")
    print("OK offline logging")
except Exception as e:
    print("ERR", type(e).__name__, str(e)[:300])
EOF
self._lazy_id = LazyValue(lambda: self.id, use_mutex=False)
    self._lazy_id = LazyValue(lambda: self.id, use_mutex=False)
                                      ^^^^^^^
  File "/opt/freestyle/python/lib/python3.12/site-packages/braintrust/logger.py", line 5984, in id
    return self.project.id
           ^^^^^^^^^^^^
  File "/opt/freestyle/python/lib/python3.12/site-packages/braintrust/logger.py", line 5980, in project
    return self._lazy_metadata.get().project
           ^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/freestyle/python/lib/python3.12/site-packages/braintrust/util.py", line 209, in get
    res = self.callable()
          ^^^^^^^^^^^^^^^
  File "/opt/freestyle/python/lib/python3.12/site-packages/braintrust/logger.py", line 1964, in compute_metadata
    state.login(org_name=org_name, api_key=api_key, app_url=app_url, force_login=force_login)
  File "/opt/freestyle/python/lib/python3.12/site-packages/braintrust/logger.py", line 786, in login
    state = login_to_state(
            ^^^^^^^^^^^^^^^
  File "/opt/freestyle/python/lib/python3.12/site-packages/braintrust/logger.py", line 2471, in login_to_state
    raise ValueError(
ValueError: Could not login to Braintrust. You may need to set BRAINTRUST_API_KEY in your environment or nearest .env.braintrust file.
Failed to construct log records to flush after 3 attempts. Dropping batch
Thinking…

Confirmed: any hosted logging requires BRAINTRUST_API_KEY. I'm blocked there.

Rewrite trace_demo.py to use decorators only + init_logger guarded, and document that it requires the key. Test the decorator-only path offline (spans are no-ops without login? @traced without init may attempt login lazily too — it likely creates NOOP_SPAN if not logged in... let me test).

Let me rewrite the file to: use @traced decorators, and wrap the upload block in a conditional that only calls init_logger when the key exists. Then test the offline path.

File edits

1 file

Completed
Writtenbraintrust_demo/trace_demo.pyView changes

/sandbox/repo/braintrust_demo/trace_demo.py

Written content · previous contents were not recorded

"""Production-style tracing with the Braintrust SDK.
python trace_demo.py # offline: spans are inert
BRAINTRUST_API_KEY=... python trace_demo.py # spans land in the hosted project
`@traced` turns any function into a span; nested calls nest automatically.
Uploading requires a logged-in logger (init_logger), which needs an API key.
"""
import os
from braintrust import traced
HAS_KEY = bool(os.environ.get("BRAINTRUST_API_KEY"))
@traced(name="retrieve_context")
def retrieve_context(ticket_id: str) -> str:
return f"order {ticket_id}: shipped 2024-05-01, refundable"
@traced(name="draft_reply")
def draft_reply(ticket_id: str) -> str:
context = retrieve_context(ticket_id) # nested span
return f"Thanks — here's what I found: {context}"
def main() -> None:
reply = draft_reply("1234")
print(reply)
if not HAS_KEY:
print("(no BRAINTRUST_API_KEY -> spans were not uploaded)")
return
from braintrust import init_logger
logger = init_logger(project="support-bot-demo")
with logger.start_span(name="support_ticket") as span:
span.log(
input={"ticket_id": "1234"},
output=reply,
metrics={"prompt_tokens": 120, "completion_tokens": 45},
scores={"helpfulness": 1.0},
)
logger.flush() # blocking flush so short scripts don't exit before upload
print("uploaded to project 'support-bot-demo'")
if __name__ == "__main__":
main()
Thinking…

I need to run it.

cd /sandbox/repo && timeout 120 python braintrust_demo/trace_demo.py 2>&1 | tail -20
Thanks — here's what I found: order 1234: shipped 2024-05-01, refundable
Thanks — here's what I found: order 1234: shipped 2024-05-01, refundable
(no BRAINTRUST_API_KEY -> spans were not uploaded)
Thinking…

Both work. Now write README with pricing + how to run + requirements.txt. Keep light.

Pricing details from page:

  • Starter: $0/mo platform fee; $10 model credits/mo included then token rates; 1 GB processed data included then $4/GB; 10k scores included then $2.50 per 1,000; 14-day retention; unlimited users/projects/datasets/playgrounds/experiments; 1 human review score per project; no custom charts/environments.
  • Pro: $249/mo; $100 credits/mo then token rates; 5 GB included then $3/GB; 50k scores included then $1.50/1k; 30-day retention then $0.50/GB/mo (up to 180 days); custom charts, environments, RBAC, priority support, S3 export, unlimited human review scores, Loop agent.
  • Enterprise: custom pricing, custom retention/export, on-prem or hosted.

Write requirements.txt and README.

File edits

1 file

Completed
Writtenbraintrust_demo/requirements.txtView changes

/sandbox/repo/braintrust_demo/requirements.txt

Written content · previous contents were not recorded

braintrust>=0.44
autoevals

File edits

1 file

Completed
Writtenbraintrust_demo/README.mdView changes

/sandbox/repo/braintrust_demo/README.md

Written content · previous contents were not recorded

# Braintrust demo (hosted product, via SDK)
Two small, runnable examples using the hosted Braintrust product through its
Python SDK. No local servers, no containers.
```
evals/support_bot.py # Eval(): dataset -> task -> scorer, prints a summary
trace_demo.py # @traced spans + logger.start_span for app tracing
requirements.txt # braintrust, autoevals
```
## Setup
```bash
pip install -r requirements.txt
```
## Run offline (no account, no credentials)
```bash
python evals/support_bot.py # prints "100.00% 'intent_match' score"
python trace_demo.py # prints the traced function's output
```
The eval runs locally with `no_send_logs=True`; the trace demo's spans are
inert. Nothing leaves the machine.
## Run against the hosted project
Get an API key from the Braintrust web app (Settings → API keys), then:
```bash
export BRAINTRUST_API_KEY=...
python evals/support_bot.py # creates project "support-bot-demo" + an experiment
python trace_demo.py # uploads spans to the same project
```
Both scripts auto-detect the key: present = upload, absent = offline.
## Notes on the SDK (verified against braintrust 0.44.1)
- `Eval(name, data, task, scores, ...)`; `data` yields `EvalCase(input=..., expected=..., metadata=...)`.
- Scorers are plain functions `(input, output, expected) -> float | Score`.
- The offline flag is `no_send_logs=True` (there is no `send_logs=` kwarg on `Eval`).
- There is no `task` decorator anymore; use plain functions or `@traced`.
- `traced(...)` is a decorator factory, not a context manager — use
`init_logger(project=...).start_span(...)` for a `with` block.
- Any hosted write path calls `login()` and hard-fails without
`BRAINTRUST_API_KEY` ("Could not login to Braintrust").
## Pricing (from braintrust.dev/pricing)
| | Starter | Pro | Enterprise |
|---|---|---|---|
| Platform fee | $0/mo | $249/mo | Custom |
| Model credits | $10/mo included, then token rates | $100/mo included, then token rates | Custom, then token rates |
| Processed data | 1 GB/mo included, then $4/GB | 5 GB/mo included, then $3/GB | Custom |
| Scores | 10k/mo included, then $2.50 per 1,000 | 50k/mo included, then $1.50 per 1,000 | Custom |
| Retention | 14 days | 30 days, then $0.50/GB/mo up to 180 days | Custom |
| Users / projects / datasets / experiments | Unlimited | Unlimited | Unlimited |
Pro adds custom charts, environments, RBAC, S3 export, unlimited human-review
scores, and the Loop agent. Starter allows 1 human-review score per project.
Costs here are usage-based on three meters (data ingested, scores computed,
model tokens via included credits), so a demo like this one stays inside the
free tier; volume comes from tracing every production request and from
LLM-as-a-judge scorers, which each count as a score.
Thinking…

Final verification run of both, and show tree.

cd /sandbox/repo/braintrust_demo && find . -type f | sort && echo "--- eval ---" && timeout 120 python evals/support_bot.py 2>&1 | tail -5 && echo "--- trace ---" && timeout 120 python trace_demo.py 2>&1 | tail -3
./README.md
./README.md
./evals/support_bot.py
./requirements.txt
./trace_demo.py
--- eval ---
100.00% 'intent_match' score



(no BRAINTRUST_API_KEY set -> ran offline; set it to upload this experiment to the hosted project)
--- trace ---
Thanks — here's what I found: order 1234: shipped 2024-05-01, refundable
(no BRAINTRUST_API_KEY -> spans were not uploaded)
Thinking…

Now I'm wrapping up with the final answer, covering the pricing details, how manageable this is, and where things got blocked.

Braintrust does have a real developer workflow I could act on, so I built it. One thing is blocked (see below).

What I built — /sandbox/repo/braintrust_demo/

FileWhat it does
evals/support_bot.pyEval(name, data, task, scores) over a 4-case intent-classification dataset, scored with exact-match + Levenshtein partial credit. Runs green: 100.00% 'intent_match' score
trace_demo.py@traced nested spans plus init_logger().start_span() for app tracing. Runs green
README.md, requirements.txtSetup, run commands, SDK gotchas, pricing table

Hosted product only, via pip install braintrust autoevals. No local stacks, nothing long-running. Both scripts auto-detect BRAINTRUST_API_KEY: present → upload to project support-bot-demo; absent → run offline (no_send_logs=True).

SDK notes I had to discover by reading the installed source (0.44.1), since the quickstart-era docs are stale:

  • the offline flag is no_send_logs=True, not send_logs=
  • there is no task decorator anymore — plain functions or @traced
  • traced(...) is a decorator factory, not a context manager; use init_logger(...).start_span(...)
  • autoevals is not pulled in as a dependency but the scorers live there

Blocked

I cannot exercise the hosted half. There is no BRAINTRUST_API_KEY in this environment, and every write path calls login() and hard-fails: ValueError: Could not login to Braintrust. You may need to set BRAINTRUST_API_KEY.... Minting a key requires creating an account through the web app (Clerk email/OAuth signup) — an interactive step I can't perform. So I verified the code paths that don't need credentials and left the upload path written but unverified against a live project. Give me a key and the same two commands will create the project, experiment, and spans.

Pricing

Usage-based on three meters, on top of a platform fee (from braintrust.dev/pricing):

StarterProEnterprise
Platform fee$0/mo$249/moCustom
Model credits$10/mo incl., then token rates$100/mo incl., then token ratesCustom + token rates
Processed data1 GB incl., then $4/GB5 GB incl., then $3/GBCustom
Scores10k incl., then $2.50/1k50k incl., then $1.50/1kCustom
Retention14 days30 days, then $0.50/GB/mo to 180Custom

Users, projects, datasets, playgrounds, and experiments are unlimited on every tier. Pro adds custom charts, environments, RBAC, S3 export, and the Loop agent. The billable surprise is scores: every LLM-as-a-judge call counts as one, so a judge-heavy eval suite scales cost linearly with dataset size × trial count.

Will this be easy for me to manage?

Mostly yes, with one structural caveat.

Easy: it's a plain pip install plus pure-Python functions — no infra, no daemon, no config files. Evals are just data/task/scores triples I can generate and edit mechanically, scorers are ordinary functions I can unit-test, and everything is scriptable so I can run a suite and read a numeric summary from stdout. Failures are loud and immediate.

The caveat: the value is in the hosted UI, which I can't see. Diffs between experiments, trace trees, playground prompt iteration, dashboards — all of that is visual and behind login. I can write and run the instrumentation, but I can't review the results the way a human does, and I can't do any account-level work (create the project, mint keys, set RBAC/retention) since that's web-console-only. Practical split: I own the eval/trace code and CI wiring; a human owns credentials and reads the UI.

AX Fundamentals

23 checks. Open an area for details.

This check is a quick preview. Test more tasks and models with Gauge Agents.

Test your product with Gauge Agents

AX Check by Gauge