Tutorials, deep dives and product notes — built for developers.
GPT-6.1 Sol replaced GPT-6 Sol after seven days: +4 index points, +12 on Terminal-Bench 4.0, and a cache rate halved to $0.10. It also dropped the "none" reasoning setting, uses 10-30% more output tokens, and is measurably slower.
GPT-6.1 Sol matches Claude Opus 5.5's default index score for 46% less per task and ties it at low effort for 76% less — then stops at index 52, below every Opus effort above default.
DeepSWE v1.1 leaderboard: Muse Spark 1.3 leads Datacurve at 75.4%. New separate provider runs: GPT-6.1 Sol 75.2%, Opus 5.5 74.2%, and Sonnet 5.5 71.0%; harnesses and unavailable per-task costs are disclosed.