I’ve been experimenting with Agent Engine Optimization for Xperience by Kentico—making content accessible through both rendered pages and structured endpoints used by AI agents. Getting an agent to retrieve content once is easy to observe; proving that it consistently discovers the correct content and cites it accurately is harder.
For teams testing agent-facing Xperience content, what does a useful evaluation framework look like? I’m considering a repeatable question set, discovery rate, retrieval success, answer groundedness, citation precision, freshness after content updates, and consistency across website, API, GraphQL, and MCP delivery paths. It also seems important to distinguish “the content was retrievable” from “the model selected it and attributed it correctly.”
Are you testing with fixed prompts and models or across multiple assistants? Which metrics and failure cases have been most useful, and how do you avoid mistaking model variability for a CMS or content-structure problem?