RE: LeoThread 2024-10-22 09:10

Anthropic says that Claude outperforms other AI agents on several key benchmarks including SWE-bench, which measures an agent's software development skills and OSWorld, which gauges an agent's capacity to use a computer operating system. The claims have yet to be independently verified. Anthropic says Claude performs tasks in OSWorld correctly 14.9 percent of the time. This is well below humans, who generally score around 75 percent, but considerably higher than the current best agents, including OpenAI’s GPT-4, which succeed roughly 7.7 percent of the time.