In 2023, a New York lawyer asked an AI chatbot to find precedents for a client's case. It delivered six. They were detailed, well-formatted, and confidently cited. They were also entirely fictional. The court sanctioned the lawyers $5,000 [1].
Nobody checked. That is the whole story of this paper.
It was supposed to be an embarrassing opening act. Three years later, the public database that tracks court decisions involving AI-fabricated material lists more than 2,000 cases [2]. That isn't a bug report. It's a growth chart.
Meanwhile, in the boardroom, the slide deck says AI will automate more than half of all work.
Both facts are true at once: the technology is powerful enough to be sold as a replacement, and unreliable enough to get professionals sanctioned. AI never asks for a raise. It also never asks for a second opinion. The gap between "sold as a replacement" and "needs a babysitter" is the most expensive misunderstanding in business right now. We call it the 16% Problem.
Executive Summary
- The number: In October 2025, the best AI agent could complete just 2.5% of real, paid freelance projects at a quality a client would accept. By mid-2026, the best had climbed to roughly 16% [3][4].
- The other 84%: Still needs a human, whether to finish it, fix it, or take the blame for it.
- The gap: McKinsey estimates today's technology could in theory automate about 57% of US work hours [5]. Theory is not your Tuesday. Between "technically possible" and "actually delivered" sits a large pile of review, correction, and accountability.
- The common thread: The lawyer with the fake cases and the freelance agent with the corrupted files failed the same way. Nobody, and nothing, checked the work.
- The opportunity: Companies that pair AI tools with skilled, affordable human reviewers and builders capture the upside without inheriting the failures.

Source: Remote Labor Index [3][4]

Source: Mid-2026 RLI result [4]

Where the Number Comes From
Most AI benchmarks resemble exams: the question is clear, the answer key exists, and a script grades the result. Real work has none of these luxuries. It arrives as a vague brief, a missing file, and a client who "will know it when they see it."
The Remote Labor Index (RLI), built by Scale AI and the Center for AI Safety, was designed to close that gap. It takes 240 real freelance projects across 23 fields, including design, architecture, data analysis, video, and web development, where money actually changed hands. It asks one blunt question: would a reasonable client accept this deliverable? [3][6]
At launch, the answer was almost always no. The best agent succeeded on 2.5% of projects. The human professionals who originally did the work were paid $143,991 in total; the top AI agent earned $1,720 [3][6]. That's roughly the price of a used laptop, and the laptop would at least turn on.
By July 2026, newer models had pushed the best score to roughly 16% [4]. That is real, fast progress, and this study doesn't pretend otherwise. But read it the other way: even at the frontier, about five out of six projects still fail the "would a client pay for this?" test.
The Benchmark Illusion
A benchmark is a test with the answer key already written. Real work rarely comes with one.
| What people hear | What the data says |
|---|---|
| "AI could automate 57% of work hours" | McKinsey's estimate is technical potential, not a forecast. The authors say it isn't a prediction of job losses, and that adoption may take decades [5]. |
| "AI aces professional exams" | Exam questions are clear, complete, and gradable. Client briefs are none of these. Lawyers who relied on unchecked AI output have been sanctioned in court [1][2]. |
| "The model scored 90% on the benchmark" | The models tested at launch scored in the low single digits on real, paid projects [3][6]. |

Source: McKinsey and RLI [3][5][6]
The lesson isn't that AI is weak. It's that capability on a test and reliability in your business are different products, and only one of them shows up on the invoice.
Why the Gap Exists - Five Structural Reasons
Hallucination is a feature of the design, not a glitch
Language models generate the most plausible continuation, not the most verified one. In practice, that means an assistant that sounds equally confident whether it's right or inventing a Supreme Court case. It's the professional equivalent of an intern who has never once said "I don't know," and whose work nobody reviews.
Stanford researchers found that general-purpose chatbots of the 2023 era got direct, verifiable questions about federal court cases wrong between 58% and 88% of the time [7]. A follow-up study, testing in 2024, found that leading commercial legal research tools still produced hallucinations in roughly 17% to 33% of cases [8]. Models have improved since, but the duty to check has not gone anywhere.
The completion and self-checking problem
When the RLI team analyzed failed submissions, the failures weren't exotic. Among failed deliverables, about 46% showed poor quality, about 36% were incomplete, about 18% had corrupted files, and about 15% were internally inconsistent (categories overlap) [6]. The paper's authors also note that many failures stem from agents being unable to verify their own work and fix mistakes [6]. One example: a video that was 8 seconds long when the brief asked for 8 minutes.
Anyone who has managed a contractor recognizes the pattern: 90% done is 0% billable. And notice the theme. The missing ingredient isn't intelligence. It's a checker.

Source: Categories overlap. RLI analysis [6]
Reliability decays with task length
METR's research shows the length of tasks AI can complete with 50% reliability has been doubling roughly every seven months [9]. That's remarkable progress. It's also worth remembering that 50% reliability is a coin flip on your payroll run. Demand 80% reliability and the achievable task length shrinks sharply [9][10]. Most businesses want somewhat better odds than a casino.
The embodiment gap (Moravec's Paradox)
Robotics pioneer Hans Moravec observed that the things humans find hard (chess, arithmetic) are easy for machines, while the things toddlers do (reading a room, carrying a tray, noticing the client is upset) are brutally hard [11]. We built machines that can draft your contract and still can't tell when the meeting has gone sideways.
Nobody can sue a model
Accountability stays with humans. When the filing is wrong, the number is wrong, or the code ships broken, the AI doesn't appear at the hearing. Someone with a name, a license, and a salary does.
Case Files
Case 1 - The Law Firm That Trusted the Autocomplete
The 2023 Mata v. Avianca sanctions were a warning shot [1]. The database maintained by researcher Damien Charlotin now tracks more than 2,000 cases where courts responded to AI-fabricated material, most of them in the US [2]. And that count only includes the fabrications someone caught.
Lesson: AI drafted the brief in seconds. Verifying it would have taken a human an afternoon. Six fake cases, zero reviewers.
Case 2 - Even "Legal-Grade" AI Needs Adult Supervision
Stanford's evaluation of commercial legal research tools found that products marketed as reliable still produced incorrect or misleading answers at material rates, even though they beat general-purpose chatbots [8]. A vendor's confidence and a tool's accuracy are, it turns out, separate variables.
Lesson: "AI-powered" on the label is not a quality control process.
Case 3 - The Freelance Agent That Earned $1,720
On the Remote Labor Index, the best agent at launch earned about 1% of the available project value [3][6]. Rapid improvement since then shows the trend is real [4], but the 84% still unautomated is where deadlines, clients, and reputations live.
Lesson: Progress is not the same as arrival. Plan for the trajectory, but staff for today.
Fairness Check: This Is Not a Doom Story
An honest case study has to say this plainly: the gap is closing. The best RLI score climbed from 2.5% to roughly 16% in under a year [3][4]. METR's time-horizon trend keeps climbing [9]. Anyone selling "AI will never do X" is selling something.
But three things haven't changed:
- 1. Verification still costs human time. Someone has to check the output, and that someone needs domain skill.
- 2. Reliability, not peak capability, is the bottleneck. A tool that's brilliant on Monday and wrong on Tuesday is a liability with a great demo.
- 3. Accountability doesn't automate. Regulators, clients, and courts want a person on the hook.
The winning strategy isn't "AI or humans." It's AI-fast, human-verified.
The Framework AI-Fast, Human-Verified
| Layer | Best handled by | Why |
|---|---|---|
| First drafts, boilerplate, code scaffolding, summarization | AI | Fast, cheap, good enough to edit |
| Data preparation and cleaning | AI + human | AI accelerates; humans catch what the data is quietly hiding |
| Review, QA, fact-checking, testing | Skilled humans | The verification layer is where quality is actually decided |
| Client communication, negotiation, context | Humans | Judgment and trust |
| Accountability and sign-off | Humans | Someone has to own it |

Source: Framework summarized from this study
AI compresses the first stretch of the work. Your competitive advantage is how well, and how affordably, you staff the last stretch.
Where Riseup Asia Comes In
Every AI workflow has a missing employee: the one who checks.
The 16% Problem creates a hiring need most companies haven't budgeted for: people who can build with AI, verify AI, and finish what AI starts. That's a difficult profile to hire locally. AI-literate engineers, QA specialists, and data professionals are expensive, slow to recruit, and in short supply. This is the gap Riseup Asia's staff augmentation is built for.
What we provide
- AI-ready engineers and developers who use modern AI tooling to ship faster, and who review its output before it reaches your customers.
- QA, testing, and verification specialists who form the human checkpoint your AI workflow needs.
- Data and workflow support to prepare, clean, and structure the inputs AI depends on.
- Flexible scaling. Add capacity when a project demands it, without the overhead of local payroll, HR, compliance, or infrastructure [12].
Why the economics work
Hiring a comparable engineer in the US can cost over $150,000 a year. In one Riseup Asia engagement, a Python engineer delivered a working AI MVP for a US client within 90 days, and the client got comparable quality at a fraction of the local hiring cost [13]. Riseup Asia provides pre-vetted talent, timezone-flexible teams, and cost-effective engagement models designed to keep quality high and budgets sane [14].
Conclusion
Six cases. Zero of them real. One missing reviewer. That is the 16% Problem in miniature. AI is not a myth, and it is not a replacement. It is a fast, tireless, occasionally fictional colleague who needs an editor.
The businesses that win won't be the ones that believe the 57% headline, or the ones that dismiss the 16% reality. They'll be the ones that build the bridge between the two and staff it wisely. The machine can draft the brief. Someone still has to check that the cases exist, and the smart move is getting excellent people to do it at a price that makes sense.
References
- 1.[1] Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. June 22, 2023) (sanctions order). Summarized in ‘AI Hallucination Legal Cases: A Sanctions Tracker (2026),’ GC AI, Aug. 2026. [Online]. Available: https://gc.ai/blog/ai-hallucination-legal-cases
- 2.[2] D. Charlotin, ‘AI Hallucination Cases Database.’ [Online]. Available: https://www.damiencharlotin.com/hallucinations/
- 3.[3] Scale AI and Center for AI Safety, ‘The Remote Labor Index: Measuring the Automation of Work,’ Scale AI Blog, Oct. 2025. [Online]. Available: https://scale.com/blog/rli
- 4.[4] Center for AI Safety, ‘A Significant Increase in Digital Labor Automation,’ Jul. 2026. [Online]. Available: https://safe.ai/blog/significant-increase-in-digital-labor-automation
- 5.[5] McKinsey Global Institute, ‘Agents, Robots, and Us: Skill Partnerships in the Age of AI,’ Nov. 2025. [Online]. Available: McKinsey Global Institute report.
- 6.[6] M. Mazeika, A. Gatti, et al., ‘Remote Labor Index: Measuring AI Automation of Remote Work,’ arXiv:2510.26787, Oct. 2025. [Online]. Available: https://arxiv.org/abs/2510.26787
- 7.[7] M. Dahl, V. Magesh, M. Suzgun, and D. E. Ho, ‘Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models,’ Journal of Legal Analysis, vol. 16, no. 1, pp. 64–93, 2024. doi: 10.1093/jla/laae003
- 8.[8] V. Magesh, F. Surani, M. Dahl, M. Suzgun, C. D. Manning, and D. E. Ho, ‘Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools,’ Journal of Empirical Legal Studies, vol. 22, p. 216, 2025. [Online]. Available: https://arxiv.org/abs/2405.20362
- 9.[9] T. Kwa et al. (METR), ‘Measuring AI Ability to Complete Long Tasks,’ arXiv:2503.14499, Mar. 2025. [Online]. Available: https://arxiv.org/abs/2503.14499. See also METR, ‘Task-Completion Time Horizons of Frontier AI Models,’ https://metr.org/time-horizons/
- 10.[10] T. Ord, ‘Is there a half-life for the success rates of AI agents?’ arXiv:2505.05115, 2025. [Online]. Available: https://arxiv.org/pdf/2505.05115
- 11.[11] H. Moravec, Mind Children: The Future of Robot and Human Intelligence. Harvard University Press, 1988.
- 12.[12] Riseup Asia, ‘Why Staff Augmentation Is Becoming Essential for Sheridan & Wyoming Businesses in 2025,’ Nov. 2025. [Online]. Available: https://riseup-asia.com/why-staff-augmentation-is-becoming-essential-in-sheridan/
- 13.[13] Riseup Asia, ‘3 Months, 1 Engineer, $150K Saved: How Joy Redefined Client Success at Riseup Asia,’ Nov. 2025. [Online]. Available: https://riseup-asia.com/3-months-1-engineer-150k-saved-how-Joy-redefined-client-success-at-riseup-asia/
- 14.[14] Riseup Asia, ‘Staff Augmentation Services.’ [Online]. Available: https://riseup-asia.com/staff-augmentation/