Snorkel research reveals 5x training uplift from high-fidelity task data
Data quality dictates the ceiling of model improvement
When we talk about scaling laws in AI, the conversation usually centers on compute budgets or parameter counts. However, Kobie Crawford from Snorkel suggests we are looking at the wrong variables. In a rigorous experiment involving Reinforcement Learning (RL) training, the team discovered that task fidelity—the inherent quality and rigor of the training data—can determine whether a model sees a marginal gain or a significant leap in performance.
The results are stark. By training the same model with the same compute budget and the same number of tasks, the difference between "low-quality" and "high-quality" task sets resulted in a 5x performance difference. While low-quality tasks offered a meager 1% improvement over the base model, high-fidelity tasks yielded a 6% uplift. This suggests that simply throwing more data at a model is less effective than curating data that meets specific logic and reliability thresholds.
The four pillars of high-fidelity tasks
To quantify "quality," the researchers established a rigorous four-point criteria for what constitutes an accepted task. First, the task must be achievable; if a problem is unsolvable by design, it serves no training purpose. Second, it must be non-trivial, requiring actual reasoning rather than simple pattern matching. Third, it must be functionally correct, ensuring the internal logic behaves as expected. Finally, the environment must be reliable, providing a stable containerized space for the agent to operate.
Tasks that failed any of these checks were relegated to the "rejected" bucket. When analyzing how models interacted with accepted tasks, researchers noticed they required twice as many tool calls and more output tokens. These tasks were intrinsically harder, pushing the model to engage in deeper reasoning and more complex tool-use sequences. This increased difficulty is exactly what allows the model to "hill-climb" and improve its capabilities during the training process.
Cleaning the signal by removing failure noise
A critical finding involves the nature of failure. In agentic workflows, not all failures are created equal. High-fidelity tasks produce "clean" failures—instances where the model fails because the logic is genuinely difficult. This provides a clear signal for the model to learn from. In contrast, rejected tasks often fail due to "degenerate" cases, such as under-specified requirements or environmental bugs.
When a task is under-specified, the model might fail because the evaluation script expects a result that was never requested in the prompt. This creates noise rather than a learning opportunity. If the model is penalized for failing a task that was impossible to satisfy based on the provided context, the training signal becomes blurred, stalling the model's progress.
Scaling expertise with the human in the loop
Maintaining this level of task fidelity at scale requires more than just automated scripts. Snorkel emphasizes a methodology that keeps experts in the loop. By using human expertise to define rubrics and ground-truth data, they can better inform LLM judges. This hybrid approach ensures that as the horizon of tasks grows longer and more complex, the evaluation remains grounded in qualitative and quantitative rigor. As we move toward more agentic systems, the ability to generate high-fidelity, verifiable data will likely become the primary differentiator in model performance.
- Snorkel
- 33%· companies
- Kobie Crawford
- 17%· people
- LLM
- 17%· programming
- Reinforcement Learning
- 17%· programming
- TerminalBench
- 17%· programming

Task Fidelity Scaling Laws — Kobie Crawdord, Snorkel
WatchAI Engineer // 20:40
We turn high signal in-person events for the top AI engineers, founders, leaders, and researchers in the world into the best free learning opportunities for millions around the world here on YouTube. Your subscribes, likes, comments, speaking, attendance, or sponsorships goes a long way toward making our biz model sustainable indefinitely. We strongly believe this industry deserves a better class of community and that we know how to do this well; we just need your support.