#benchmarks

4 articles

Qwen3.8-Max Ships With 2.4T Parameters, 16-Day Autonomous Coding Demo, and Open Weights Coming Next Week

Alibaba's Qwen3.8-Max runs autonomous coding projects for 16 days straight, beats Claude Opus 4.8 on Terminal Bench, and will open…

Perplexity's Orchestrator Now Runs Grok 4.5, Outpaces Opus on Key Benchmark

Perplexity integrates Grok 4.5 into its orchestrator system and achieves top results on the WANDR benchmark, surpassing Opus in mu…

GPT-5.6 Sol Matches Claude Fable 5 on Code Arena — For 40% Less

GPT-5.6 Sol ties Claude Fable 5 on the Code Arena benchmark at 40% lower cost, shaking up the performance-per-dollar calculation f…

AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

Today's roundup: Musk pushes Grok 4.5 inside Tesla and SpaceX, Perplexity's orchestrator tops Opus on a benchmark, Meta launches a…