SWE-Bench Pro पर GPT-5.6 Sol बनाम Claude Fable 5 बनाम Gemini 3.1 Pro
23 जुलाई 2026 · 21 मिनट पढ़ें · Claude / GPT / Gemini

वह नंबर जो शॉर्टलिस्ट बदल देता है
रिपोर्टेड SWE-Bench Pro पंक्ति में Claude Fable 5 मात देने वाला मॉडल है: 80.0%। GPT-5.6 Sol 64.6% पर है। Gemini 3.1 Pro Preview 54.2% पर है। ये तीनों नंबर OpenAI की 9 जुलाई, 2026 की GPT-5.6 लॉन्च तालिका से आते हैं, Coding सेक्शन में SWE-Bench Pro के तहत (OpenAI).
यह Claude Fable 5 और GPT-5.6 Sol के बीच 15.4-पॉइंट का अंतर है, और Claude Fable 5 और Gemini 3.1 Pro Preview के बीच 25.8-पॉइंट का अंतर। किसी coding-agent बिल्डर के लिए यह इतना बड़ा है कि इसे यूँ ही नज़रअंदाज़ नहीं किया जा सकता।
लेकिन सही अर्थ हेडलाइन से ज़्यादा सीमित है। यह repository-level agentic coding है। यह Terminal-Bench नहीं है। यह Aider का leaderboard नहीं है। यह vibe-coded UI benchmark नहीं है। अगर आपका प्रोडक्ट वास्तविक repos में patches लगाता है, files खोलता है, tests चलाता है, कई modules में edits करता है, और उलझे हुए dependency graphs में टिके रहने की ज़रूरत रखता है, तो SWE-Bench Pro कई generic coding rows की तुलना में आपके workload के ज़्यादा करीब है। अगर आपका प्रोडक्ट मुख्यतः chat completions, code snippets, या CLI puzzle solving है, तो यह row सिर्फ़ एक signal है।

तुलना तालिका
यह साफ़ version है, जिसमें benchmark row को दूसरे coding-agent numbers से अलग रखा गया है।
| Model | Vendor family | SWE-Bench Pro reported score | Source for score | API price checked |
|---|---|---|---|---|
| Claude Fable 5 | Claude | 80.0% | OpenAI GPT-5.6 table (OpenAI) | $10 input / $50 output per 1M tokens (Anthropic) |
| GPT-5.6 Sol | GPT | 64.6% | OpenAI GPT-5.6 table (OpenAI) | $5 input / $30 output per 1M tokens (OpenAI) |
| Gemini 3.1 Pro Preview | Gemini | 54.2% | OpenAI GPT-5.6 table (OpenAI); cross-checked against Google’s Gemini 3.5 Flash model card for Gemini 3.1 Pro on SWE-Bench Pro Public (Google DeepMind) | $2 input / $12 output per 1M tokens for prompts up to 200k tokens, standard tier (Google AI) |
Gemini number का cross-check सबसे बेहतर है। Google का Gemini 3.5 Flash model card “as of May, 2026” results सूचीबद्ध करता है और “SWE-Bench Pro (Public)” पर Gemini 3.1 Pro को “Single attempt” setup के साथ 54.2% रिपोर्ट करता है (Google DeepMind). OpenAI की July table “Gemini 3.1 Pro Preview” और वही 54.2% number इस्तेमाल करती है। इससे यह साबित नहीं होता कि हर detail में harness settings बिल्कुल समान हैं, लेकिन इससे Gemini score किसी एक vendor-only citation की तुलना में कम संदिग्ध लगता है।
इस तुलना में Claude Fable 5 का 80.0% score, Anthropic benchmark table के बजाय OpenAI की table से लिया गया है। Anthropic का अपना Fable page release context और pricing के लिए अब भी उपयोगी है: Fable 5 को “hardest knowledge and coding work” के लिए position किया गया है, यह claude-fable-5 API name इस्तेमाल करता है, और इसकी कीमत $10 per million input tokens और $50 per million output tokens है (Anthropic). Anthropic यह भी कहता है कि Fable 5 और Mythos 5 को 9 जून को release किया गया, 12 जून को U.S. government directive के बाद suspend किया गया, और export controls हटाए जाने के बाद 1 जुलाई को globally restore किया गया (Anthropic).
SWE-Bench Pro असल में क्या मापता है
SWE-Bench Pro को पुराने SWE-Bench Verified ceiling से आगे जाने के लिए बनाया गया था। Scale इसे ज़्यादा realistic repositories में long-horizon software engineering tasks के benchmark के रूप में वर्णित करता है, जिसमें 41 repositories में कुल 1,865 tasks हैं: 731 public, 858 held-out, और 276 commercial (Scale). Public set मजबूत copyleft licenses वाली open-source repositories इस्तेमाल करता है; private subset proprietary startup codebases इस्तेमाल करता है।
Primary metric resolve rate है। कोई task तब resolved माना जाता है जब submitted patch नए fail-to-pass tests पास करता है और pass-to-pass tests को नहीं तोड़ता (Scale). यह detail मायने रखती है। किसी model को plausible patch, सुंदर diff, या मददगार explanation के लिए credit नहीं मिलता। उसे repo को ऐसे बदलना होता है कि benchmark environment उसे स्वीकार करे।
इसीलिए SWE-Bench Pro repository-agent bucket में आता है। एक realistic agent run “write a function” से कम और कुछ ऐसा ज़्यादा दिखता है:
1. inspect issue and requirements
2. search repository conventions
3. edit one or more files
4. run targeted tests
5. debug failures
6. rerun regression tests
7. return a patch
यही flow है जहाँ context handling, tool use, error recovery, और test discipline सचमुच मायने रखने लगते हैं। जो model isolated code generation में चमकता है, वह dependency tree hostile होने पर या failing test के framework-specific edge case में छिपे होने पर फिर भी बिखर सकता है।

पेच: Benchmark दबाव में है
असहज बात यह है: OpenAI ने GPT-5.6 table 9 जुलाई, 2026 को publish की, यानी 8 जुलाई को detailed SWE-Bench Pro audit publish करने के एक दिन बाद। उस audit में, OpenAI ने अनुमान लगाया कि SWE-Bench Pro tasks के लगभग 30% broken हैं और कहा कि वह benchmark adopt करने की अपनी पहले की recommendation वापस ले रहा है (OpenAI).
Audit में दो अलग-अलग estimates मिले: एक automated-plus-review pipeline ने 200 broken tasks, यानी 27.4%, flag किए, जबकि human annotation campaign ने 249, यानी 34.1%, पहचाने (OpenAI). सूचीबद्ध failure categories वही हैं जिनसे benchmark users डरते हैं: overly strict tests, underspecified prompts, low-coverage tests, और misleading prompts।
इससे table बेकार नहीं हो जाती। इसका मतलब यह है कि आपको second decimal place, या कुछ percentage points तक, engineering truth की तरह treat नहीं करना चाहिए। Claude Fable 5 और GPT-5.6 Sol के बीच 15.4-point gap अब भी जाँच के लायक पर्याप्त meaningful है। July 2026 में इस benchmark पर 1- या 2-point gap कमजोर evidence होगा।
Practical stance यह है: SWE-Bench Pro को reported repository-level signal की तरह इस्तेमाल करें, फिर इसके आधार पर architecture खरीदने से पहले अपना harness दोबारा चलाएँ।
इसे Terminal-Bench या Aider के साथ मत मिलाइए
OpenAI की उसी GPT-5.6 table में Terminal-Bench 2.1 भी report है, जहाँ GPT-5.6 Sol 88.8% पर है, GPT-5.6 Sol Ultra 91.9% पर है, Claude Fable 5 83.1% पर है, और Gemini 3.1 Pro Preview 70.7% पर है (OpenAI). यह अलग कहानी बताता है क्योंकि task shape अलग है। Terminal-Bench terminal workflows पर ज़ोर देता है। SWE-Bench Pro repo patch resolution पर ज़ोर देता है।
इन scores को एक “best coding model” number में मिलाना वही तरीका है जिससे teams खुद को भ्रमित करती हैं। Aider-style benchmarks एक और harness layer जोड़ते हैं: prompt format, edit mode, retry policy, और model-specific adapters results को बदल सकते हैं। वे उपयोगी हैं, लेकिन वे अपने task construction और grading rules वाले repository benchmark के साथ interchangeable नहीं हैं।
Coding-agent builders के लिए, अपनी eval sheet को lanes में अलग करें:
repo_patch_resolution:
benchmark: SWE-Bench Pro
metric: resolve_rate
terminal_workflows:
benchmark: Terminal-Bench
metric: task_success
interactive_editing:
benchmark: Aider-style runs
metric: accepted_patch_or_tests_passed
internal_repos:
benchmark: private eval set
metric: merged_patch_after_review
फिर हर lane के अंदर models की तुलना करें। जब तक आप weights को स्पष्ट न करें, lanes का average न निकालें। Migration agent, test-writing bot, और terminal ops agent की failure costs अलग-अलग होती हैं।
निष्कर्ष
अगर SWE-Bench Pro ही lane है, तो इन तीन models में Claude Fable 5 reported leader है: 80.0%, जबकि GPT-5.6 Sol 64.6% और Gemini 3.1 Pro Preview 54.2% पर हैं। List API pricing पर GPT-5.6 Sol, Fable 5 से सस्ता है, $5/$30 per million tokens बनाम $10/$50, इसलिए अगर आपका agent lower reported resolve rate सह सकता है तो यह cost-adjusted throughput पर अब भी जीत सकता है। Gemini 3.1 Pro Preview फिर और सस्ता है, 200k-token prompts तक $2/$12 पर, लेकिन reported SWE-Bench Pro gap इतना बड़ा है कि इसे economical choice कहने से पहले आपको इसे अपनी repositories पर validate करना चाहिए।
सबसे तीखा takeaway “Claude सब कुछ जीतता है” नहीं है। बात यह है: repo-level agents के लिए repo-level benchmarks इस्तेमाल करें, harness को cite करें, और benchmark health पर नज़र रखें। SWE-Bench Pro अभी कहता है कि इस comparison में Fable 5 आगे है, लेकिन benchmark में documented quality problems इतनी हैं कि serious teams को इसे deployment decision नहीं, बल्कि shortlist filter की तरह treat करना चाहिए।
जो readers इन models को hands-on try करना चाहते हैं, वे onehop के ज़रिए OpenAI-compatible API से Claude और अन्य models call कर सकते हैं: एक base_url बदलें, runs compare करें, और अपना harness stable रखें (onehop पर Claude और अन्य models call करें). New accounts को बिना card के $10 free credit मिलता है, और pricing first-party access से सस्ती है ( $10 free credit के लिए sign up करें).
संबंधित लेख

Aider Polyglot Coding पर GPT-5 बनाम Gemini 2.5 Pro बनाम Claude Opus 4
Aider Polyglot coding पर GPT-5, Gemini 2.5 Pro और Claude Opus 4 की डेटा-आधारित तुलना।
17 जून 2026 · 20 मिनट पढ़ें

Terminal-Bench 2.0 पर Gemini 3.1 Pro बनाम GPT-5.2 बनाम Claude Opus 4.6
साझा Terminal-Bench 2.0 harness में Gemini 3.1 Pro आगे है, पर harness का चुनाव CLI coding की कहानी बदल देता है।
16 जून 2026 · 21 मिनट पढ़ें

DashScope Compatible Mode के ज़रिए OpenAI SDK से Qwen3.7 Plus कॉल करें
base_url बदलाव, thinking controls, long-context pricing और चलने योग्य code के साथ OpenAI SDK ऐप्स को Qwen3.7 Plus पर पोर्ट करें।
23 जुलाई 2026 · 25 मिनट पढ़ें