Agent evaluation

Compare task quality, access controls, and action behavior without hiding failures in one score.

Illustrative reference architecture 3 min read

Reference diagram

Separate execution evidence from release judgment

Controlled execution

  1. Versioned casesQualified labels, permissions, and expected outcomes
    Run with bounded test authority
  2. Candidate runtimeIsolated targets and the actual authorization gateway
    Record results and control decisions
  3. Execution evidenceAnswers, retrieval, actions, traces, and failures

Assessment and release

  1. Separate evaluatorsTask quality, access, actions, cost, and latency
    Escalate disagreement and blocking failures
  2. Human adjudicationReview evidence and qualify uncertain labels
    Decide against stated release criteria
  3. Promotion recordApproved versions, regressions, and rollback target

Connections between paths

  • Execution evidence→Separate evaluators

    Evaluate the observed run against versioned criteria

  • Human adjudication→Versioned cases

    Add reviewed failures as new regression cases

  • Promotion record→Candidate runtime

    Select the approved revision as the baseline for subsequent tests

Illustrative reference architecture. Quality scores and permission checks remain separate. The promotion record governs a separate deployment process. Test fixtures contain side effects; integrated access tests cover current authority across retries, pauses, and delegation.

Start with reviewable cases

I would begin with a curated set of tasks, not a benchmark score. This reference design assumes versioned agent releases, permitted test data, and an isolated execution environment. Each case records the question, relevant evidence, caller permissions, expected outcome, and label rationale.

Include ordinary requests, ambiguous questions, missing information, and known regressions. Keep uncertain labels visible and separate from scored cases. Version the dataset and its rubrics so a comparison can identify whether the application changed or the definition of success changed.

Measure different responsibilities separately

Retrieval checks whether relevant, permitted evidence was found. Answer evaluation checks factual support, completeness, and appropriate uncertainty. Tool evaluation checks target selection, arguments, and observed effects. Authorization tests check what data or actions the caller may access.

A plausible answer cannot compensate for an unauthorized read. Report these dimensions separately and define blocking failures explicitly. Test the intermediate evidence and tool results, not only the final response. A system can refuse politely after already exposing a protected record to a model.

Exercise the real control boundaries

Use fixtures or isolated services for consequential actions, with test identities and no production response credentials. Record proposed and executed actions separately. Simulated tools support repeatability, while integration tests check the actual authorization gateway and destination behavior.

Include expired sessions, revoked permissions, changed tool arguments, denied resources, and access removed while a run is paused. Resume, retry, and subagent calls must recheck current authority. Test cancellation and uncertain action outcomes without blindly repeating effects. A successful initial login must not authorize the entire run forever.

Record enough to explain a comparison

Associate results with dataset, prompt, model, tool-contract, policy, and application versions. Capture retrieval references, tool outcomes, errors, latency, and available usage measurements with a trace identifier. Protect sensitive payloads and retain only the evidence needed for review.

Record end-to-end latency separately from model time and estimate cost using the applicable pricing basis. Caches, dependency behavior, and model variation limit repeatability. Repeat representative cases and report variation rather than presenting one run as a stable performance claim. A trace explains execution; it does not establish answer correctness.

Use judges with human calibration

Use deterministic checks for exact requirements, such as valid structure or prohibited tool calls. A model judge can assess a stated rubric against supplied evidence, but its judgment is another fallible output. Keep the judge model, prompt, and score rationale with the result.

Have people review representative cases, disagreements, and consequential failures. Calibrate judge behavior against adjudicated examples and keep evaluation references out of the agent's input. Preserve reviewer disagreement instead of forcing an uncertain case into a confident label.

Promote against explicit release criteria

Compare the candidate and current release on the same case versions and relevant slices. Investigate regressions individually; an improved average must not conceal broken access controls. Require an accountable review before promotion and retain the previous configuration for rollback.

Monitor permitted live samples after release and turn reviewed failures into regression cases. A small application can start with a spreadsheet and a controlled test runner. Add evaluation infrastructure as needed; automated grading never establishes that every future request will succeed.

References

Search the site

Search experience, studies, articles, projects, and contributions.

Try a topic