Terminal-Bench 4.0: GLM-5.3 Rises to Third, Surpassing GPT-5.6 Sol
Terminal-Bench released version 4.0, recalibrating time, CPU, and memory for Agent task execution, and fixing 19 tasks while removing 8 tasks with issues of saturation, refusal responses, public…
Terminal-Bench released version 4.0, recalibrating time, CPU, and memory for Agent task execution, and fixing 19 tasks while removing 8 tasks with issues of saturation, refusal responses, public solutions, or quality problems. The maximum execution time for all tasks has been unified to 8 hours, primarily to reduce interference from timeouts and environmental issues on scores.
On the latest leaderboard, Opus 5 + Claude Code ranks first with 51.8%, and Fable 5 is second with 44.5%. GLM-5.3 + Claude Code achieved 41.8%, rising to third place, surpassing GPT-5.6 Sol + Codex's 37.3%. Among the top three, GLM-5.3 is the only model not from Anthropic.
In Terminal-Bench 3.0, GLM-5.3 was fourth with 32.4%, trailing GPT-5.6 Sol's 34.6%; in 4.0, GLM-5.3 rose to third, leading Sol by 4.5 percentage points.
[BlockBeats]
Original: https://www.theblockbeats.info/flash/364243
insigtX content is informational and educational, not investment advice.