Measuring what actually matters

Most AI benchmarks test isolated skills like answering questions, writing code snippets, and summarizing text. The RLI was built on a different premise. The research team wanted to know whether an AI agent could take a task from beginning to end the way a paid professional would, and whether the output would meet a paying client’s standard.

Tasks were sourced from digital labor platforms like Upwork and spanned 23 sectors, including video editing, logo and leaflet design, architecture, data analysis, jewelry design, and game development. Evaluators then compared AI-generated deliverables against human-produced ones with one question in mind. Would a client actually pay for this?

“If you consider creating a window and you have designs and all that, it could be the case that AI can create something very aesthetically pleasing,” Sehwag said. “But if the dimensions are incorrect, it doesn’t matter how pleasing it looks. A human is not going to actually pay for that.”

The benchmark also tracks a live leaderboard of AI agent performance scores, updated as new models are evaluated. The top performer as of mid-2026, claude-opus-4-6 via the CoWork platform, sits at 4.17%. Everything else is lower.

The reliability gap

The low automation rate isn’t simply a matter of AI agents producing bad work. Sehwag points to something more specific.