Test in Playground
The Playground runs an Auto score scorecard against a conversation you pick and shows the scored result immediately — without creating an evaluation or touching any scores. Use it to iterate quickly when tuning a scorecard: adjust a question, test, read the result, adjust again.
Before you begin
- The scorecard must be of type Auto score and already linked to an Auto QA task by your administrator (see Enabling Auto QA); otherwise the test cannot run and the page reports "This form is not assigned to AI Task".
- You need edit permission on scorecards — the same as for the scorecard designer.
- The conversation you test against must have a transcript.
Running a test
- Open the scorecard in QA › Scorecards and click Test. Alternatively, open an existing auto-scored evaluation and click Open in Playground — this pre-selects that evaluation's conversation, which is convenient when a reviewer flagged it as scored wrong.
- If no conversation is pre-selected, pick one from the list ("Select for an experiment").
- The Playground shows the scorecard's sections and questions. Edit question and description text in place if you want to try a variation.
- Click Test. The AI scores the selected conversation against the form as shown, and the result — answers, justifications, and section scores — renders on the same page.
- Iterate: tweak the wording or descriptions and click Test again until the answers match what a human evaluator would give.
- When you are happy with the changes, click Save Evaluation Form to apply them to the scorecard. Until then, your edits exist only in the Playground.
Note
Playground runs are throwaway experiments: they do not create evaluations, do not appear in the Evaluations list, and do not affect any agent's scores. Test freely.
Tips
- Test against a handful of conversations you know well — including ones the AI previously answered incorrectly — rather than a single example.
- Test one change at a time, so you know which wording fixed (or broke) an answer.