Talha Sheikh says agent verification beats relying on smarter AI models
The Illusion of Agent Completion
We have all been there. You hand Claude a coding task, watch it spin up sub-agents, execute tasks, and proudly declare victory. But when you run the code, it breaks. This loop of constant, manual micro-adjustments reveals a fundamental flaw in the modern developer-agent relationship. Humans have become the enforcement layer. The agent says it is finished, but without a deterministic way to verify that completion, "done" is just a suggestion.
To bridge this gap, developers need to transition from giving better instructions to building hard enforcement systems. That is the core philosophy behind Vector Harness, a tool designed to hook into agent sessions and automatically run verification checks. If a test fails, the harness feeds the error back to the agent to try again until it actually works.
Capability Is Not Reliability

When frontier models improve, their capabilities expand, but their reliability does not scale linearly. Many developers assume that upcoming, smarter models will render verification obsolete. However, giving an LLM better context or instructions is not the same as verifying its output.
By investing in a robust, deterministic testing harness, you can actually reduce dependency on expensive frontier models. A well-designed harness with strict guardrails can help a smaller, cheaper model like Claude Haiku achieve the same reliable outputs as a massive, expensive model. The leverage lies in the testing infrastructure, not the raw parameter count.
The Shift to Language-Agnostic Patterns
Instead of building custom, isolated enforcement scripts, the industry is moving toward standardized verification patterns. This approach is language-agnostic and works at every stage of the lifecycle: during conversation, pre-commit, or as part of a multi-agent workflow.
Industry giants are adopting this paradigm. Anthropic recently introduced its "executor-advisor" pattern, while OpenAI relies heavily on harness engineering. The consensus is clear: the real value in software development is shifting away from the code we generate and toward the verification systems we design.
- Anthropic
- 17%· companies
- Claude
- 17%· products
- Claude Haiku
- 17%· products
- OpenAI
- 17%· companies
- Talha Sheikh
- 17%· people
- Vector Harness
- 17%· products

Your coding agent doesn't always follow your rules — Talha Sheikh, Checkout.com
WatchAI Engineer // 10:08
We turn high signal in-person events for the top AI engineers, founders, leaders, and researchers in the world into the best free learning opportunities for millions around the world here on YouTube. Your subscribes, likes, comments, speaking, attendance, or sponsorships goes a long way toward making our biz model sustainable indefinitely. We strongly believe this industry deserves a better class of community and that we know how to do this well; we just need your support.