Duolingo engineer warns human partners are blindly approving flawed AI decisions

AI Engineer////4 min read

The Dangerous Myth of Human Oversight

Duolingo engineer warns human partners are blindly approving flawed AI decisions
Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo

We love putting humans in our software loops. We call it safety. We call it ethical compliance. We treat the human-in-the-loop model as a ironclad safeguard against rogue artificial intelligence. But the truth is much uglier.

Your human partner isn't thinking. They are surviving.

As systems automate more cognitive labor, humans stop evaluating. They click yes. They move on. This dynamic turns human oversight into an expensive, ineffective rubber stamp. Angel Ortmann Lee, a software engineer on the Duolingo English Test team at Duolingo, recently highlighted this critical point. If we build systems for approval instead of discernment, we are just creating automated feedback loops of unverified machine assumptions.

The Psychology of Cognitive Surrender

Trust in technology is high. We do not memorize telephone numbers anymore. We follow GPS directions blindly. This habitual trust bleeds into professional AI tools.

Researchers at the Wharton School uncovered a striking behavioral pattern they named cognitive surrender. This happens when a human stops critical analysis entirely and adopts AI output as their own. The Wharton study showed that when an AI resource was correct, human test scores jumped by 25 percentage points. However, when the AI was wrong, human scores plummeted by 15 points.

Most damningly, 80 percent of human participants accepted incorrect AI answers without question. The machine did the thinking. The human just agreed.

Coin-Flip Security at Duolingo

This is not a theoretical academic problem. It affects high-stakes environments. The Duolingo English Test is a remotely proctored, high-stakes exam used for college admissions and visas. The company uses custom keystroke-monitoring models to catch "copy typing"—a cheating method where a candidate copies text instead of composing it naturally.

To test how well their highly skilled human proctors caught mistakes, the team ran a silent test. They took clean, legitimate test sessions and injected fake AI warning signals claiming the test-taker was cheating.

The results were shocking. Highly trained proctors, who consistently score above 90 percent on accuracy calibration tests, accepted these fake AI cheating signals 50 percent of the time. They falsely accused innocent test-takers based on a simulated warning flag. It was a coin flip.

The problem was not the model, which maintained a low 1 percent false positive rate. The problem was not the human staff. The problem was the interface.

Engineering Friction Into the Interface

To break this automation bias, the engineering team altered the proctoring guidelines. They updated the interface instructions with a crucial two-part change. First, they explicitly framed the AI signal as a preliminary alert, emphasizing that the human holds final responsibility. Second, they mandated that proctors locate independent video evidence before upholding any cheating flag.

This simple copy edit shattered the rubber-stamp habit. Rejection of false signals spiked by 21 percent. Proctors began evaluating rather than agreeing.

Product builders must design interactions that force friction. For high-stakes decisions, seamlessness is the enemy of safety. Speed bumps force active analytical thought.

Flipping the Interaction Loop

If your interface encourages mindless skimming, you only collect binary, junk data. When users blindly accept AI recommendations, the system logs those validations as ground truth. This creates a vicious cycle where a confident, flawed model trains itself on false human agreements.

Breaking this requires structured interfaces. Instead of presenting a simple "yes or no" button for a headphone detection flag, break it into distinct questions. Ask "were headphones present?" and "was this a policy violation?" Separating these concerns protects the underlying model from bad training data and forces the human to act as an investigator, not a validator.

Topic DensityMention share of the most discussed topics · 6 mentions across 6 distinct topics
Angel Ortmann Lee
17%· people
Duolingo
17%· companies
Duolingo English Test
17%· products
Gemini
17%· products
Google
17%· companies
Wharton School
17%· companies
End of Article
Source video
Duolingo engineer warns human partners are blindly approving flawed AI decisions

Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo

Watch

AI Engineer // 25:53

We turn high signal in-person events for the top AI engineers, founders, leaders, and researchers in the world into the best free learning opportunities for millions around the world here on YouTube. Your subscribes, likes, comments, speaking, attendance, or sponsorships goes a long way toward making our biz model sustainable indefinitely. We strongly believe this industry deserves a better class of community and that we know how to do this well; we just need your support.

Who and what they mention most
Anthropic
26.9%21
Claude
21.8%17
OpenAI
19.2%15
Cursor
15.4%12
4 min read0%
4 min read