AI Evals Should Start Before Prompt Tuning

Direct answer

Create the first evaluation set before serious prompt tuning or model selection. It does not need to be large, but it must represent the decisions, documents, edge cases, and failure costs the product will face. Without that baseline, teams optimize for memorable examples and cannot tell whether a change improved the product or merely moved the errors.

A polished demo is not an evaluation

Most AI prototypes begin with a handful of examples chosen by the people building them. Those examples prove that the workflow can work. They do not show how often it works, where it breaks, or whether a new prompt is better than the old one.

The gap appears quickly in production. A legal assistant meets scanned documents, conflicting clauses, and questions with no supporting evidence. An interview coach hears clipped audio, long pauses, and answers that are relevant but poorly structured. A regulatory assistant must distinguish federal guidance from the rule that applies in one city.

One happy path cannot represent that spread.

Define the decision before the score

Start with the outcome the AI supports. Then define what a reviewer would accept.

For extraction, this may be field-level accuracy and a strict rule against inventing missing values. For retrieval, it may be whether the correct source appears in the top results. For a generated clinical note, it may combine factual coverage, prohibited additions, format compliance, and clinician edit distance. For an agent, it may include task completion, tool selection, policy compliance, and recovery after a failed call.

A single “quality” score hides too much. Use separate dimensions when the consequences differ.

Build the set from real variation

A useful first set often contains 50 to 100 cases, not thousands. Include normal cases, boundary cases, known failures, ambiguous inputs, permission differences, empty evidence, malformed files, and adversarial instructions relevant to the product.

Production traffic should then keep feeding the set. Sample failures, escalations, user corrections, low-confidence cases, and expensive runs. Remove personal data where required, preserve the characteristics that made the case difficult, and label why the expected result is correct.

The set becomes a product asset: part test suite, part risk register, part institutional memory.

Use deterministic checks before model judges

Many acceptance rules do not require another model. Validate JSON structure, required fields, citations, identifiers, numeric ranges, permissions, tool arguments, and forbidden actions with code.

Use human review or a calibrated model grader for qualities that need judgment, such as completeness, groundedness, tone, or reasoning quality. Periodically compare automated graders with expert decisions. A grader that agrees with itself but not with the people accountable for the outcome creates false confidence.

Evals turn model choice into an engineering decision

Once the same cases run against every candidate, model selection becomes concrete. A smaller model may handle classification reliably while a larger model earns its cost on complex synthesis. A new prompt may improve format compliance but weaken citations. A retrieval change may raise answer quality while increasing latency.

Those trade-offs should appear before release, not in customer complaints.

This mattered in Zenveus products where AI feedback, document-grounded research, and regulatory answers depended on different definitions of “good.” The right evaluation unit followed the workflow: useful coaching feedback for an interview answer, cited evidence for a legal or regulatory response, and correct permissions before enterprise analytics could query data.

A practical release gate

Before shipping a prompt, model, retrieval, or tool change:

  1. Run the frozen evaluation set on the current and candidate versions.
  2. Compare each quality dimension, latency, and cost.
  3. Inspect regressions, not only averages.
  4. Require human review for high-consequence failures.
  5. Record the version and result with the release.

If the team cannot explain what improved, what regressed, and which failures remain acceptable, the change is not ready.

Related Zenveus services: Agentic AI Development and AI Prototype Hardening

Sources

Scroll to Top