Start with reviewable cases
I would begin with a curated set of tasks, not a benchmark score. This reference design assumes versioned agent releases, permitted test data, and an isolated execution environment. Each case records the question, relevant evidence, caller permissions, expected outcome, and label rationale.
Include ordinary requests, ambiguous questions, missing information, and known regressions. Keep uncertain labels visible and separate from scored cases. Version the dataset and its rubrics so a comparison can identify whether the application changed or the definition of success changed.
Measure different responsibilities separately
Retrieval checks whether relevant, permitted evidence was found. Answer evaluation checks factual support, completeness, and appropriate uncertainty. Tool evaluation checks target selection, arguments, and observed effects. Authorization tests check what data or actions the caller may access.
A plausible answer cannot compensate for an unauthorized read. Report these dimensions separately and define blocking failures explicitly. Test the intermediate evidence and tool results, not only the final response. A system can refuse politely after already exposing a protected record to a model.
Exercise the real control boundaries
Use fixtures or isolated services for consequential actions, with test identities and no production response credentials. Record proposed and executed actions separately. Simulated tools support repeatability, while integration tests check the actual authorization gateway and destination behavior.
Include expired sessions, revoked permissions, changed tool arguments, denied resources, and access removed while a run is paused. Resume, retry, and subagent calls must recheck current authority. Test cancellation and uncertain action outcomes without blindly repeating effects. A successful initial login must not authorize the entire run forever.
Record enough to explain a comparison
Associate results with dataset, prompt, model, tool-contract, policy, and application versions. Capture retrieval references, tool outcomes, errors, latency, and available usage measurements with a trace identifier. Protect sensitive payloads and retain only the evidence needed for review.
Record end-to-end latency separately from model time and estimate cost using the applicable pricing basis. Caches, dependency behavior, and model variation limit repeatability. Repeat representative cases and report variation rather than presenting one run as a stable performance claim. A trace explains execution; it does not establish answer correctness.
Use judges with human calibration
Use deterministic checks for exact requirements, such as valid structure or prohibited tool calls. A model judge can assess a stated rubric against supplied evidence, but its judgment is another fallible output. Keep the judge model, prompt, and score rationale with the result.
Have people review representative cases, disagreements, and consequential failures. Calibrate judge behavior against adjudicated examples and keep evaluation references out of the agent's input. Preserve reviewer disagreement instead of forcing an uncertain case into a confident label.
Promote against explicit release criteria
Compare the candidate and current release on the same case versions and relevant slices. Investigate regressions individually; an improved average must not conceal broken access controls. Require an accountable review before promotion and retain the previous configuration for rollback.
Monitor permitted live samples after release and turn reviewed failures into regression cases. A small application can start with a spreadsheet and a controlled test runner. Add evaluation infrastructure as needed; automated grading never establishes that every future request will succeed.