Skip to content

Tuning scorecards

Auto QA is only as good as the scorecard it grades against. When the AI's answers do not match what your QA team would answer, the fix is usually not a different AI — it is a clearer scorecard. This page covers the two tuning levers scorecard owners have: the description fields, and the override feedback loop.

Use descriptions as grading guidance

The Description fields of questions and sections are sent to the AI along with the question text, as extra grading context. Use them to state what a correct answer looks like — the things a human evaluator "just knows" but the AI cannot guess:

  • Define terms. For "Did the agent greet the customer properly?", the description can spell out what "properly" means: a greeting within the first seconds, the company name, and an offer to help.
  • Set boundaries. State when the answer should be N/A — for example, "If the caller hung up before the agent could respond, answer N/A."
  • Resolve known confusions. If reviewers keep overriding a question for the same reason, encode that reason: "A transfer initiated by the customer does not count as a failure to resolve."

Keep the question text itself short and answerable; put the nuance in the description. Section descriptions work the same way and apply as context for all questions in the section.

Iterate on override feedback

Every override can carry a Feedback for AI Assistant comment explaining what the AI missed. This is your tuning signal:

  1. Periodically review overridden questions — advanced search can filter evaluations where the overridden answer differs from the original AI answer, and the reviewers' feedback comments explain why.
  2. Look for patterns: a question that is overridden often, and for the same reason, needs clearer wording or a better description.
  3. Update the question or its description in the scorecard designer.
  4. Verify the change with Test in Playground against a few conversations you know well — including ones the AI previously got wrong.

Existing evaluations are not re-scored when you edit the scorecard; the improvement shows up in new evaluations, which are scored against the updated revision.