Anthropic’s Fable 5 Just Set a Benchmark Record — But It’s No Human Replacement

Anthropic’s Fable 5 is back after a brief government-mandated pause, and it’s already resetting expectations for what AI can do.

The Center for AI Safety (CAIS) tested it on the Remote Labor Index (RLI), which measures how often AI agents complete real freelance projects at a quality a paying client would accept. Fable 5 hit 16.1% — a record. For context, Anthropic’s Opus 4.8 scored 8.3% and OpenAI’s GPT-5.5 scored 6.3%. The previous published leader was 4.17%.

“The frontier has more than quadrupled in under eight months,” CAIS said.

Even under a worst-case assumption — that Fable 5 failed every project it couldn’t complete during testing — it would still land at 14.6%. Higher than any other model.

So what does that mean for freelancers? Probably not the end of the world just yet. 16% is nowhere near 100%. And AI isn’t plug-and-play for most organizations. Security concerns, integration hurdles, and the need for a network of agents to check work quality, budget, and timelines means the tradeoff isn’t one-to-one.

CAIS even tried replacing the human evaluator with an “LLM judge” to see how far they could push automation. It failed. “Evaluating an RLI deliverable is itself a demanding, agentic task,” they explained.

Here’s the catch: computer-use skills are the current limitation. And those are exactly what the industry is investing heavily in. If that roadblock disappears — and at the current rate of improvement, it very well could — the calculus changes fast.