Example: missing or weak instructions
User intent: "My agent gives vague answers in evals."
Symptoms
relevancepasses butcoherencefails.groundednessfails because the agent invents missing facts.similarityfails because the answer omits required structure or decisions.- Follow-up turns fail after a successful first turn.
Diagnosis
Check whether the agent instructions specify:
- supported scenarios and boundaries,
- source-grounding expectations,
- citation expectations,
- response format,
- behavior when data is missing,
- follow-up context handling.
Suggested instruction additions
Use only available workplace sources when answering source-backed questions. If the sources do not contain enough evidence, say what is missing instead of guessing.For project-status answers, use this structure: Summary, Evidence, Risks, Next actions. Keep the answer concise and include citations when available.For follow-up questions, preserve the project, customer, and time window from the prior turn unless the user changes them.Matching eval update
Add or keep regression prompts that test the new instruction:
{
"prompt": "Summarize the latest status for the project and include citations.",
"expected_response": "The agent summarizes only source-backed status, cites available sources, and states when evidence is missing.",
"evaluators": {
"Groundedness": {
"threshold": 4
},
"Citations": {
"threshold": 1
}
},
"evaluators_mode": "extend"
}