Authored By

Lillian Zhao

Lead Product Manager
No items found.

How Cresta Grades Its Agents: Evaluators in an LLM World

Learn how Cresta evaluates AI Agents at scale using LLM judges, deterministic checks, simulations, and human calibration to catch failures before launch and build trust in high-stakes environments.

Learn more
Engineering
All
No items found.

The Data Comes First: Mining Real Conversations for Test Coverage

Learn how Cresta mines historical data, synthetic customers, knowledge-base Q&A pairs, and post-launch feedback to build test coverage that reflects how customers actually behave.

Learn more
Engineering
All
No items found.

Why AI Agent Evaluations Fail — and How the Swiss-Cheese Model Prevails

Learn about Cresta's forward-deployed team and how they approach building and iterating on AI agents.

Learn more
Engineering
All
This author has no guides
No items found.

How Cresta Grades Its Agents: Evaluators in an LLM World

Learn how Cresta evaluates AI Agents at scale using LLM judges, deterministic checks, simulations, and human calibration to catch failures before launch and build trust in high-stakes environments.

Learn more
Engineering
All
No items found.

The Data Comes First: Mining Real Conversations for Test Coverage

Learn how Cresta mines historical data, synthetic customers, knowledge-base Q&A pairs, and post-launch feedback to build test coverage that reflects how customers actually behave.

Learn more
Engineering
All
No items found.

Why AI Agent Evaluations Fail — and How the Swiss-Cheese Model Prevails

Learn about Cresta's forward-deployed team and how they approach building and iterating on AI agents.

Learn more
Engineering
All
This author has no guides