OpenAI has put its latest models through a new workplace stress test called GDPval, designed to see how they stack up against real human professionals across 44 job types in nine heavyweight industries like finance, healthcare, and manufacturing. In short: GPT-5-high, the premium version of the upcoming model, tied or beat human-generated work 40.6% of the time. Interestingly, Anthropic’s Claude Opus 4.1 edged it out with a 49% win/tie rate—though OpenAI attributes this to Claude’s knack for sexy visuals, not superior output.
It’s worth noting that GDPval only evaluates task-based work like reports or market analyses, not the full spectrum of interactive, messy, and emotional tasks that make up a real 9-to-5. Still, OpenAI sees progress. Their internal economist believes these tools are coming in hot as workplace copilots, freeing skilled humans from tedium so they can tackle “higher value” projects.
The jump from GPT-4o’s 13.7% to GPT-5-high’s 40.6% in little over a year is giving OpenAI’s team cause for cautious optimism. But senior ad execs should take note: these AI tools are less about replacing talent and more about supporting it—at least for now. As benchmarks like GDPval evolve to include more complex workflows, their influence could grow in shaping how brands approach AI in knowledge work.

Read more at Tech Crunch.
