Muktesh Mishra warns accuracy metrics fail complex AI agent evaluations

AI Engineer////2 min read

Software development has entered a new era where Muktesh Mishra argues that evaluation-driven development must replace traditional testing. In non-deterministic systems like LLMs, the same input can yield wildly different results, making standard unit tests insufficient. Developers must now architect Evals that account for subjectivity and qualitative performance rather than just binary success or failure.

Data is the bedrock of evaluation

One data set is never enough. To build a robust AI application, you must curate multiple data sets tailored to specific user flows. Start small with synthetic data to validate basic outputs, but quickly transition to a continuous refinement process. This involves labeling data according to application prospects and observing system behavior in real-time to ensure the data remains relevant as the model evolves.

Muktesh Mishra warns accuracy metrics fail complex AI agent evaluations
How to run Evals at Scale: Thinking beyond Accuracy or Similarity — Muktesh Mishra, Adobe

Moving beyond generic similarity scores

A universal metric for AI does not exist. A Retrieval-Augmented Generation system requires measuring retrieval accuracy and usefulness, whereas code generation demands a focus on functional correctness and robustness. When evaluating agents, the focus shifts to trajectory evaluation. We must analyze the specific paths an agent takes to execute a task, rather than just the final output. This includes multi-turn simulations to test how agents handle complex, conversational state changes.

Balancing speed and fidelity at scale

Scaling Evals requires a strategic trade-off between automated speed and human fidelity. While automation provides rapid feedback, certain subjective nuances still require a human in the loop. The most effective strategy involves caching intermediate results to prevent redundant computation and running evaluations frequently through parallel orchestration. Success lies in a cycle of measuring, monitoring, and iterating until the system aligns with business goals.

Topic DensityMention share of the most discussed topics · 7 mentions across 6 distinct topics
Evals
29%· technology
Adobe
14%· companies
AI application
14%· technology
LLMs
14%· technology
Muktesh Mishra
14%· people
End of Article
Source video
Muktesh Mishra warns accuracy metrics fail complex AI agent evaluations

How to run Evals at Scale: Thinking beyond Accuracy or Similarity — Muktesh Mishra, Adobe

Watch

AI Engineer // 9:25

We turn high signal in-person events for the top AI engineers, founders, leaders, and researchers in the world into the best free learning opportunities for millions around the world here on YouTube. Your subscribes, likes, comments, speaking, attendance, or sponsorships goes a long way toward making our biz model sustainable indefinitely. We strongly believe this industry deserves a better class of community and that we know how to do this well; we just need your support.

Who and what they mention most
Anthropic
26.9%21
Claude
21.8%17
OpenAI
19.2%15
Cursor
15.4%12
2 min read0%
2 min read