A Counterintuitive Truth: Agents Still Have Low Success Rates in Real-World Work Scenarios
According to evaluations, current mainstream large models perform far worse in real-world e-commerce, daily life, and other scenario tasks than in programming. For example, in the RealReplicaBench e-commerce…
According to evaluations, current mainstream large models perform far worse in real-world e-commerce, daily life, and other scenario tasks than in programming. For example, in the RealReplicaBench e-commerce benchmark, only Claude Opus 5 passed with a rate just above 60%, while most models fell below 50%. In the world of AI, there remains a considerable gap between solving problems and actually working.
Original: https://wallstreetcn.com/articles/3779598
insigtX content is informational and educational, not investment advice.