The AI Model Frontier

Understand Claude and every major AI model, daily

Curated news, benchmarks, and hands-on guides for Claude, GPT, Gemini and more - keep up with the AI model frontier in one place.

Latest articles

View all
A cream-background editorial cover showing a developer terminal connected by three terracotta routing lines to labeled a
GuidesJuly 23, 2026 · 25 min read

Call Qwen3.7 Plus with the OpenAI SDK via DashScope Compatible Mode

Port OpenAI SDK apps to Qwen3.7 Plus with base_url changes, thinking controls, long-context pricing, and runnable code.

A cream editorial cover image showing three abstract model columns labeled conceptually by family colors, with a terraco
BenchmarksJuly 23, 2026 · 21 min read

GPT-5.6 Sol vs Claude Fable 5 vs Gemini 3.1 Pro on SWE-Bench Pro

A developer-focused comparison of reported SWE-Bench Pro scores for Claude Fable 5, GPT-5.6 Sol, and Gemini 3.1 Pro.

Claude Code loop engineering control system illustration
EcosystemJune 20, 2026 · 8 min read

Claude Code Loop Engineering: How to Build an Agent That Actually Finishes

A practical breakdown of Claude Code loop engineering: goals, hooks, subagents, context, verification, and the limits of autonomous coding agents.

Cream-background editorial cover showing a developer terminal connected by terracotta lines to three labeled concept nod
GuidesJune 17, 2026 · 24 min read

Use Groq GPT-OSS 120B with the OpenAI SDK: Base URL, Pricing, and Caching

Swap one OpenAI SDK base URL to run GPT-OSS 120B on Groq, estimate cached token costs, and avoid tool billing surprises.

Cream-background editorial illustration of three abstract coding model cards racing across a polyglot test grid, with te
BenchmarksJune 17, 2026 · 20 min read

GPT-5 vs Gemini 2.5 Pro vs Claude Opus 4 on Aider Polyglot Coding

A data-first comparison of GPT-5, Gemini 2.5 Pro, and Claude Opus 4 on Aider Polyglot coding.

Cream-background editorial cover showing three abstract terminal windows as stacked charcoal cards, each connected to a
BenchmarksJune 16, 2026 · 21 min read

Gemini 3.1 Pro vs GPT-5.2 vs Claude Opus 4.6 on Terminal-Bench 2.0

Gemini 3.1 Pro leads the shared Terminal-Bench 2.0 harness, but harness choice changes the CLI coding story.

Cream-background editorial cover showing a developer terminal window connected by terracotta lines to a Grok model card
NewsJune 16, 2026 · 20 min read

Using Grok Build in Warp with a SuperGrok or X Premium Subscription

xAI now lets Warp users connect Grok or X Premium and run grok-build-0.1 inside terminal agent workflows.

A cream-background editorial illustration of a developer workspace inside a transparent sandbox cube, with file trees, t
EcosystemJune 16, 2026 · 25 min read

OpenAI Agents SDK Native Sandbox and Manifest Guide

How OpenAI’s Agents SDK sandbox and Manifest let developers build safer file-working agents without custom orchestration.