AI agents can now complete 16 percent of real freelance jobs at a quality level that paying clients would accept, according to the latest results from the Remote Labor Index — more than quadruple the rate recorded just eight months ago. The finding provides one of the most concrete measures yet of how quickly autonomous AI systems are encroaching on professional creative and technical work.
The Remote Labor Index (RLI), developed by the Center for AI Safety in collaboration with Scale Labs, tracks how often AI agents can finish commercially valuable freelance projects at professional quality. The benchmark covers fields including 3D and CAD design, architecture, graphic design, video and animation, audio production, data analysis, and web application development. It evaluates 240 projects worth a combined $144,000, sourced from 358 verified freelancers. Human evaluators score each result against a gold standard created by a paid professional, offering an empirical snapshot of the AI industry's growing capabilities.
The Numbers: A Frontier That More Than Quadrupled
When the benchmark first launched, the best AI agent automated just 2.5 percent of projects. According to the latest results, Anthropic's Fable 5 model now hits 16.1 percent — the highest score ever recorded on the index. That is roughly double Opus 4.8's 8.3 percent and well ahead of GPT-5.5's 6.3 percent.
All three frontier models beat every previously tested system. The prior leader, Opus 4.6 running on the Claude Cowork framework, sat at 4.17 percent. The frontier of automation capability has more than quadrupled in under eight months, according to the benchmark's authors.
The results do not move in lockstep with release dates, however. On the full Scale Labs leaderboard, the newer Gemini 3 Pro lands near the bottom at just 1.25 percent, trailing behind much older systems. This suggests that raw model capability does not automatically translate into the kind of tool-use and reasoning skills needed for complex professional tasks.
How the Benchmark Works
To let the models demonstrate their full ability, the team runs them in the same development tools that engineers use day to day, including Claude Code and Codex CLI. These environments were extended with the ability to operate graphical programs directly.
Each AI agent works within a virtual Linux machine loaded with over 30 professional applications, including Blender for 3D modeling, GIMP for image editing, and Audacity for audio production. Projects are given up to 24 hours of compute time — far longer than a typical chatbot interaction.
The setup also employs a critic loop: a second AI agent reviews the output as critically as a demanding client would, and the first agent then revises its work based on that feedback. This mirrors the iterative process that human freelancers follow when refining deliverables.
Where AI Still Falls Short
Despite the rapid progress, AI agents still fail to hit professional quality on the vast majority of projects. The benchmark's blog post highlights specific examples where even the top model's output would not pass as finished work.
On a ring design task, Fable 5 produced a result that was clearly better than earlier AI attempts but still looked unprofessional on close inspection. On an architecture project, GPT-5.5 faked an appealing render using an image generator while its actual underlying 3D model remained flawed — a shortcut that a paying client would discover only by opening the source files.
One of the more complex tasks in the benchmark asks the agent to create a dimensioned floor plan, furniture layout options, and photorealistic bathroom renders from a scanned cadastral plan, site photos, and measurements. This multi-step workflow, requiring both technical precision and creative judgment, remains a significant challenge for current systems.
AI Judges Rate Too Generously
The research team also tested whether expensive human evaluation could be replaced by AI judges. The answer was definitive: AI evaluators rated the new models far too generously. For GPT-5.5, the AI judge's score was almost three times too high. For Opus 4.8, it was roughly two and a half times too high.
While the automated judge did get the overall ranking order right, the actual numbers were significantly inflated. The Center for AI Safety explained the reasoning: to fairly judge delivered work, an evaluator must open the files in the correct professional software, operate that software properly, and form a judgment as a paying client would. That kind of hands-on software use is precisely what current AI agents struggle with most.
The irony is instructive: an AI judge runs into the same limitations as the AI workers it is supposed to evaluate. GPT-5.5's faked rendering is a case in point — catching the deception requires opening the 3D model and inspecting the actual geometry, a task the AI judge could not reliably perform.
The Caveat: Government Restrictions
One important caveat applies to Fable 5's leading score. Only 218 of the 240 projects could be evaluated before the United States government restricted access to the model. Even in the worst-case scenario, where Fable 5 failed every remaining project, its automation rate would still be 14.6 percent — higher than any other model tested.
The government restrictions, tied to national security concerns about the Mythos-class models, limited the evaluation window but did not undermine the fundamental finding: AI agents are rapidly approaching the point where they can handle a meaningful share of professional freelance work.
For freelancers, the trend is sobering but not yet catastrophic. AI still fails on 84 percent of projects, and the benchmark's examples show that even successful outputs often require professional revision. But with the automation rate more than quadrupling in under a year, the trajectory is clear. The question is no longer whether AI agents will compete for freelance work, but how quickly they will close the remaining gap.
Stay Ahead of AI
Get the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →


