What Is an AI Eval? A Builder’s Guide to Testing LLM Outputs Before You Ship

VTechNews Editorial Team · · 9 min read · 1,720 words

Bottom line: an eval is just a test suite for your model’s outputs instead of your code’s control flow — and if you’re shipping an LLM feature without one, the first regression you catch will be a user complaint, not a CI failure. This guide builds one real eval suite for a 3-endpoint RAG app, shows the actual grader code (both code-based checks and LLM-as-judge), and runs it before and after swapping the underlying model, because that’s the exact moment evals earn their keep: a model swap that looks fine in five manual spot-checks can still regress silently at scale.

Quick Answer

  • An AI eval is an automated test suite that scores LLM outputs against a labeled dataset, using either code-based checks (regex, exact match, JSON schema) or LLM-as-judge grading for open-ended answers.
  • Minimum viable eval suite: 50–200 labeled test cases per endpoint, a mix of code-based and judge-based graders, run automatically on every prompt or model change.
  • Tooling in 2026: OpenAI Evals (MIT, offline, YAML config) for OpenAI-only stacks; LangChain’s OpenEvals/AgentEvals for LangGraph-native agents; Braintrust for a dataset-first, model-agnostic hosted option.
  • When you need one: before any model swap, prompt rewrite, or RAG retrieval change that ships to production.

What is an AI eval, actually?

An eval is a dataset of inputs paired with a way to score the outputs — nothing more exotic than that. For a support-ticket classifier, the dataset might be 150 real tickets with the correct category attached; the grader checks whether the model’s output matches. For a RAG chatbot, the dataset is a set of questions with the facts that a correct answer must contain; the grader can be a second LLM call asking “does this response contain these facts, yes or no.” The distinction that trips builders up: evals score outputs, not code paths, so a suite that passes doesn’t mean the code works — it means the model’s answers, on this dataset, meet your bar.

Why not just eyeball a few outputs before shipping?

Because five manual checks and a 200-case eval suite answer different questions. Manual review tells you the model can produce a good answer. A suite tells you how often it does, and on which slice of inputs it doesn’t. We built a suite for a 3-endpoint RAG app — document search, summarization, and a citation-answer endpoint — and ran the same 200 test cases before and after swapping the answer model from GPT-5.5 to Gemini 3 Pro. Five manual spot-checks on the new model looked clean. The full suite told a different story on the citation-answer endpoint specifically: the kind of gap a small manual sample structurally can’t surface, because it doesn’t cover enough of the input distribution to catch a regression concentrated in one slice of traffic.

How do you build the grader code for code-based checks?

Code-based graders are the cheapest and most reliable option whenever the correct answer has a checkable structure — JSON schema validity, an exact string match, a regex, a numeric range. Here’s a minimal grader for the document-search endpoint, checking that the top result’s ID is in the expected set:

def grade_search(output, expected_ids, k=3):
    top_k_ids = [r["doc_id"] for r in output["results"][:k]]
    hit = any(doc_id in expected_ids for doc_id in top_k_ids)
    return {"pass": hit, "score": 1.0 if hit else 0.0}

Run this against every labeled case in the dataset, aggregate the pass rate, and you have a number you can compare across model or retrieval changes — not a vibe.

When do you need LLM-as-judge instead of code-based checks?

Whenever “correct” doesn’t reduce to a string match — a summary, an explanation, a citation-backed answer. LLM-as-judge uses a second model call to grade the first model’s output against a rubric. This is the pattern LangChain’s OpenEvals and AgentEvals libraries package as pre-built evaluators, and it’s also how Braintrust’s dataset-first scoring works under the hood for open-ended tasks. A grader for the citation-answer endpoint might look like this:

JUDGE_PROMPT = """
Question: {question}
Reference facts: {reference_facts}
Model answer: {answer}

Does the model answer contain all the reference facts, without
contradicting any of them? Answer only "pass" or "fail", then a
one-sentence reason.
"""

def grade_citation_answer(question, reference_facts, answer, judge_model):
    result = judge_model.complete(
        JUDGE_PROMPT.format(question=question,
                             reference_facts=reference_facts,
                             answer=answer)
    )
    verdict = result.strip().lower().startswith("pass")
    return {"pass": verdict, "raw": result}

The judge model doesn’t have to be the same model you’re testing — using a different model as judge avoids a model grading its own homework too generously, a known failure mode when the judge and the model under test share biases.

Pro Tip: Cache judge-model responses keyed on the input hash. Re-running a 200-case suite with a fresh LLM-as-judge call every time turns a five-minute CI check into a real line item on your OpenRouter bill.

Which eval tool should you actually use in 2026?

It depends on what your stack already looks like, not which tool has the most features. OpenAI Evals (MIT-licensed) fits a mostly-OpenAI workload where a YAML config file is enough — note the hosted version of the product is being retired later in 2026, so treat it as an offline/self-hosted grading library, not a dashboard product. If every agent run in your app is already a LangChain or LangGraph execution, LangSmith with OpenEvals/AgentEvals gives you trajectory-level tracing that’s native to the framework instead of bolted on. Braintrust is the pick if your team wants a polished hosted UI and dataset-first workflow with model-agnostic, sandboxed Python custom scorers — it closed an $80M Series B at an $800M valuation in February 2026, with Notion, Replit, Cloudflare, and Ramp among its customers, which says something about how many teams outgrew a homegrown eval script.

ToolBest fitModel dependencyNotable limit
OpenAI EvalsMostly-OpenAI stacks, offline gradingOpenAI-centricHosted product retiring late 2026
LangChain OpenEvals / AgentEvalsLangGraph-native agentsModel-agnosticBest value inside LangChain ecosystem
BraintrustTeams wanting a hosted UI + dataset-first workflowModel-agnosticPaid; overkill for a single small suite
Watch out: A judge model grading against a vague rubric (“is this a good answer?”) produces noisy, inconsistent scores. Anchor every judge prompt to explicit reference facts or a checklist, the same way the grade_citation_answer example does above — not an open-ended quality question.

How big does an eval dataset need to be before it’s useful?

Fewer than 30 cases and you’re back to manual-review-with-extra-steps — too small to catch a regression concentrated in one input slice. Somewhere between 50 and 200 labeled cases per endpoint is the practical range most builder-scale suites land in: enough to cover the real variety of inputs an endpoint sees (short questions, multi-part questions, edge-case phrasing, adversarial inputs) without requiring a labeling team. The 3-endpoint RAG app example above used roughly 200 cases per endpoint, sourced from real logged queries with facts hand-verified against the source documents — not synthetic questions generated by the same model being tested, which tends to produce an eval set biased toward what that model already handles well.

Pro Tip: Pull eval cases from real production logs (with PII stripped) rather than hand-writing them from scratch. Logged queries surface the phrasing and edge cases actual users produce, which is usually stranger than what a builder thinks to write by hand.

What should you do when your eval suite catches a regression?

Treat the failing cases as a debugging dataset, not just a red X. Pull the specific inputs that flipped from pass to fail after the model or prompt change, and look for a pattern — a regression concentrated in multi-hop questions, or in citations pulled from a specific document type, is a very different fix than a regression spread evenly across the dataset. This is the same before/after discipline that matters when picking the retrieval layer underneath your RAG app in the first place — see our embedding model comparison for the same before/after benchmarking method applied one layer down the stack, and our vector database pricing comparison if the retrieval side of your RAG app is the thing you’re about to change next. If the model swap itself is the variable under test, our GPT-5.5 context window explainer covers what actually changes in long-running RAG and agent workloads when you move to a larger-context model.

Key Takeaways

  • An eval is a labeled dataset plus a grader — code-based for checkable structure, LLM-as-judge for open-ended answers.
  • 50–200 labeled cases per endpoint, sourced from real logs, is the practical range for a builder-scale suite.
  • Use a different model as judge than the model under test, to avoid grading bias.
  • OpenAI Evals fits OpenAI-only offline grading; LangSmith/OpenEvals fits LangGraph-native agents; Braintrust fits teams wanting a hosted, model-agnostic UI.
  • Run the suite before every model swap or prompt change that ships — manual spot-checks can’t reliably surface a regression concentrated in one input slice.

Frequently Asked Questions

What’s the difference between an eval and a unit test?

A unit test checks deterministic code behavior against a fixed expected output. An eval scores probabilistic model outputs against a rubric or reference, usually reporting a pass rate across a dataset rather than a single pass/fail.

Do I need a paid tool like Braintrust to run evals?

No. OpenAI Evals is MIT-licensed and free to self-host, and a minimal grader like the code examples above can run in a plain Python script with no framework at all. Paid tools buy you a hosted UI, dataset versioning, and team collaboration features, not the core capability.

How often should I re-run my eval suite?

On every prompt change, model swap, or retrieval-layer change before it ships — ideally wired into CI the same way a test suite runs on every pull request.

Can the same model grade its own outputs?

It can, but it’s a known failure mode for the judge to score its own outputs more generously than an independent model would. Using a different model as judge is cheap insurance against that bias.

What’s a reasonable eval dataset size to start with?

Start with whatever you can label accurately — even 30–50 cases beats zero. Grow toward 100–200 per endpoint as you pull more real queries from production logs.

Is LLM-as-judge reliable enough to trust for shipping decisions?

It’s reliable when anchored to explicit reference facts or a checklist rubric, not an open-ended “is this good” question. Vague rubrics produce noisy scores that swing run to run.

Last updated: 2026-09-23

FREE DAILY NEWSLETTER

Get the AI News That Matters

3-minute daily digest for executives. Curated by AI, edited by humans.

Get the 1k+ ChatGPT Prompts Bible (Free)

Join 5,000+ executives getting our 3-minute daily AI digest and get instant access to the Premium Knowledge Vault.

Leave a Comment