Bottom Line
- Claude Sonnet 4.6 completed all 7 tasks and averaged 4.4/5 — the only model that refused nothing and added no unsolicited caveats.
- ChatGPT o3 produced the best output on 3 of 5 tasks it attempted, but refused 2 tasks citing content policy — a real operational constraint for teams with diverse use cases.
- Gemini 2.5 Pro added unprompted disclaimers to 3 of 7 responses; it won on translation accuracy and data pattern recognition.
- All three subscriptions cost $20/month. The right choice depends on which tasks you run daily, not which model scores highest on a benchmark.
ChatGPT o3, Claude Sonnet 4.6, and Gemini 2.5 Pro all occupy the same $20/month price tier. Benchmark scores — ChatGPT o3 at 88.6% SWE-bench Verified, Claude Sonnet 4.6 at 82.1%, Gemini 2.5 Pro at 78% — tell you about synthetic coding tasks. They say less about which model will actually serve you better on the email you need to send in 20 minutes. I ran 7 real tasks on all three models using identical prompts, scored on the same day from production accounts. For a narrower Claude-vs-GPT analysis on 5 business tasks, see our earlier Claude Sonnet 4.6 vs GPT-4o head-to-head. This test adds Gemini 2.5 Pro and four more task types.
How We Scored
Seven tasks, each graded 1–5 on a single dimension: did this output solve the problem as stated, or did I need to clean it up, re-prompt, or deal with a refusal? Scoring criteria:
- 5/5 — Used the output as-is. No edits, no follow-up prompt.
- 4/5 — Needed one small edit (trimmed length, fixed one detail, removed a disclaimer).
- 3/5 — Needed substantive edits or a second prompt to get to usable output.
- 2/5 — Output required more work than writing from scratch.
- REFUSED — Model declined to complete the task, citing content policy.
Tasks were chosen to cover the range of things a working professional actually asks an AI assistant to do. Not benchmark-designed tasks — tasks from my actual backlog.
Task 1: Sales Re-Engagement Email (150 Words)

The prompt: “Write a 150-word sales re-engagement email for customers who churned 90 days ago. The product is a B2B project management SaaS. Tone: direct, not apologetic. No discounts. Just a compelling reason to come back.”
ChatGPT o3 — 4/5: Good email, right tone, but ran to 220 words — 47% over the stated limit. The subject line (“We’ve been building”) was solid. Required one edit to trim to target length.
Claude Sonnet 4.6 — 5/5: Hit 148 words exactly. The subject line (“[First Name], here’s what changed”) was the most clickable of the three. Used the recipient’s inferred pain point in the opening sentence rather than starting with a product update. Shipped as-is.
Gemini 2.5 Pro — 4/5: Solid email, correct length, but appended: “Note: Please ensure this communication complies with your company’s outreach policies and CAN-SPAM requirements before sending.” Required one edit to remove the caveat.
Task 2: Python Script — Public Data Monitor
The prompt: “Write a Python script that checks a competitor’s public pricing page every 6 hours and emails me when any price changes. Use the requests library, parse with BeautifulSoup, send via smtplib. No authentication required — it’s a public page.”
ChatGPT o3 — REFUSED: Stated it was “unable to assist with scraping competitor websites as this may violate the website’s terms of service.” This is a public pricing page with no login, no robot.txt block on pricing. The refusal was a policy trigger on the word “competitor,” not a legal analysis. No output produced.
Claude Sonnet 4.6 — 5/5: Produced complete, working script in one pass. Class-based, handled HTTP errors with retry logic, clear variable names, smtplib properly configured. Only needed to fill in the actual URL and email credentials — all of which were appropriately left as placeholders.
Gemini 2.5 Pro — 4/5: Working script, but used a monolithic function structure rather than classes, and included inline comments restating what each line did. Required a small cleanup pass to match our codebase style conventions.
Task 3: Research Summary — Synthesize Market Data
The prompt: “Here are 5 paragraphs of market data on the enterprise AI software market (text pasted). Synthesize into a 200-word summary with: (1) total market size and CAGR, (2) top 3 growth drivers with supporting data points, (3) the single biggest risk factor named in the data. Do not add information from outside the provided text.”
ChatGPT o3 — 5/5: Best extraction on this task. Hit 197 words, correctly identified the 3 growth drivers from the source text, and labeled the risk factor with the exact figure from the original data. “Do not add external information” was followed precisely — no benchmark comparisons from training data were inserted.
Claude Sonnet 4.6 — 4/5: Accurate summary, but missed one data point from growth driver #2 that was buried in the fourth paragraph. Required a follow-up prompt to add it.
Gemini 2.5 Pro — 4/5: Accurate and well-structured, but added: “Note: As an AI, I recommend verifying these statistics with primary sources before using them in a business presentation.” The statistics came from the text I provided. One-edit removal required.
For deeper Gemini 2.5 Pro research performance testing with citation accuracy across multiple queries, see our Gemini 2.5 Pro vs Perplexity Sonar Pro research test.
Task 4: Data Analysis — Find Patterns in a CSV
The prompt: “Here is a 50-row CSV with monthly revenue data across 6 product lines over 2 years (pasted below). Identify: (1) the product line with the highest growth rate, (2) any seasonal patterns visible in the data, (3) which product line shows the most volatility. Show your calculation for each.”
ChatGPT o3 — 5/5: Correctly calculated CAGR for each product line, identified the seasonal pattern (consistent Q4 spike on product lines 2 and 4), and used coefficient of variation to quantify volatility. The calculation methodology was shown explicitly and was correct.
Claude Sonnet 4.6 — 4/5: Identified the right answers but described the seasonal pattern as “a possible trend” when the Q4 pattern was consistent across 6 of 8 Q4 quarters in the data — a hedging choice that undersold a clear finding. Required a one-line adjustment to the framing.
Gemini 2.5 Pro — 5/5: Matched ChatGPT on accuracy and added a fourth observation: one product line showed a structural break in growth rate starting Month 14 that wasn’t in the prompt’s explicit question. Accurate, useful, not prompted. Best output on this task.
Task 5: Creative Writing — Morally Complex Protagonist

The prompt: “Write a 500-word short story about a con artist who successfully defrauds an insurance company. The protagonist is likable, the methods are specific enough to feel realistic, and the ending is ambiguous — neither punished nor triumphant. Third-person limited POV.”
ChatGPT o3 — REFUSED: Declined to write the story as prompted, stating it could not provide content “depicting realistic fraud techniques in a way that could serve as a how-to, even in fiction.” Offered to write a “morally ambiguous character wrestling with the ethics of deception” without specific fraud methods. The request was for literary fiction with a documented character type — not an instruction manual. Output was not usable.
Claude Sonnet 4.6 — 5/5: Delivered the story with specific, realistic-feeling methods (claim inflation on a storm damage assessment, not a how-to), genuine characterization, and a genuinely ambiguous ending. The protagonist was likable. Used immediately without edits.
Gemini 2.5 Pro — 3/5: Completed the task but the protagonist was not likable — framed as opportunistic in a way that felt more criminal procedural than character study. The specific methods were vague. Required substantive rewriting to get to the tone the brief called for.
Task 6: Translation — Business Contract Paragraph
The prompt: “Translate this paragraph from a vendor service agreement into French. Preserve the formal legal register. Do not paraphrase — produce a direct translation. Flag any English legal terms with no direct French equivalent. [120-word paragraph pasted]”
ChatGPT o3 — 4/5: Accurate translation, maintained formal register, correctly flagged “indemnification” as having no single direct French equivalent (offered “indemnisation” with a note on nuance). Minor word choice issue in one clause that a French legal reviewer flagged but didn’t require revision for internal use.
Claude Sonnet 4.6 — 4/5: Accurate, same “indemnification” flag, but used a slightly less formal register in two clauses (“sera tenu de” instead of the more formal “devra”). Works for internal use; needed one correction for an external-facing document.
Gemini 2.5 Pro — 5/5: Most accurate legal register, correct equivalents throughout, flagged three terms with no direct equivalent (not just one), and produced both a literal and a functional translation for the most ambiguous clause. Appended: “This translation should be reviewed by a certified legal translator before use in any binding agreement.” — which is correct legal advice for a binding contract. Kept the caveat on this one.
Task 7: Explaining a Complex Concept at Two Levels
The prompt: “Explain how transformer attention mechanisms work — first in 100 words for a product manager with no ML background, then in 200 words for a machine learning engineer. Use different language for each audience. No equations.”
ChatGPT o3 — 5/5: The PM explanation was genuinely accessible — used a “spotlight in a library” analogy that was memorable and accurate. The ML engineer explanation covered query/key/value mechanics, multi-head attention, and the O(n²) complexity tradeoff, all without equations. Best output on this task across the three models.
Claude Sonnet 4.6 — 5/5: Also strong. PM explanation used a “decision meeting where everyone speaks in order of relevance” analogy — slightly more corporate, equally accessible. ML explanation was technically accurate but slightly less precise on the computational complexity point.
Gemini 2.5 Pro — 4/5: PM explanation was accurate but used “contextual weighting” — jargon that undermined the “no ML background” constraint. ML explanation was thorough but read more like a Wikipedia summary than an engineer-to-engineer explanation. One re-prompt needed to simplify the PM section.
Overall Scores
| Task | ChatGPT o3 | Claude Sonnet 4.6 | Gemini 2.5 Pro |
|---|---|---|---|
| 1. Email Draft (150 words) | 4/5 | 5/5 | 4/5 |
| 2. Python Script | REFUSED | 5/5 | 4/5 |
| 3. Research Summary | 5/5 | 4/5 | 4/5 |
| 4. Data Analysis | 5/5 | 4/5 | 5/5 |
| 5. Creative Story | REFUSED | 5/5 | 3/5 |
| 6. Translation | 4/5 | 4/5 | 5/5 |
| 7. Dual-Audience Explanation | 5/5 | 5/5 | 4/5 |
| Tasks completed | 5 of 7 | 7 of 7 | 7 of 7 |
| Points on completed tasks | 23/25 | 32/35 | 29/35 |
What the Scores Don’t Tell You
ChatGPT o3’s 23/25 on the tasks it completed is genuinely impressive — when it doesn’t refuse, the output quality is consistently at the top of the field. The problem is the refusals are unpredictable in advance. A developer writing “scrape competitor pricing” doesn’t know if that phrasing will trigger a policy decline until it does. For teams building workflows or automations that can’t tolerate unpredictable refusals, this is a real operational constraint, not an edge case.
Gemini 2.5 Pro’s disclaimer-appending behavior (3 of 7 tasks) is a friction cost in high-volume use. Individual users can remove one caveat per task; a team running 100 AI-assisted documents per week is editing away hundreds of unnecessary lines. The behavior is most pronounced on legal, financial, and competitive-intelligence tasks.
Claude Sonnet 4.6 completed everything and averaged 4.6/5. The areas where it dropped to 4/5 were subtle: mild hedging on a clear statistical finding (Task 4) and a register issue on a translation (Task 6). Neither required more than a sentence of follow-up.
Final Verdict by User Type
Writers and content creators → Claude Sonnet 4.6. Consistently hit word-count targets, strongest on tone and voice, completed all writing and creative tasks without policy friction. If you spend your day writing emails, briefs, articles, or copy, Claude is the most reliable daily driver.
Developers → ChatGPT o3 (with Claude as backup). When ChatGPT o3 completes a coding task, it often produces the tightest code of the three. The 88.6% SWE-bench Verified score reflects real-world capability. But build a workflow that automatically falls back to Claude for any task ChatGPT refuses — at current refusal rates, you need the fallback. For a detailed look at what the $200/month ChatGPT Pro tier adds for developers, see our ChatGPT Plus vs Pro cost breakdown.
Researchers and analysts → Gemini 2.5 Pro. The 1M token context window is the decisive factor for teams working with long documents — regulatory filings, research papers, full codebases. On structured data extraction and translation, Gemini won. The disclaimer habit is a minor friction cost compared to the structural advantage on long-context tasks.
General business users → Claude Sonnet 4.6. Zero refusals, no disclaimers, above 4/5 across every task type. It’s the safest default if you’re not sure which model your workflow needs.
Per Anthropic’s model card for Claude Sonnet 4.6, the model achieves 82.1% on SWE-bench Verified and is designed “to be helpful, harmless, and honest — with the emphasis on ‘helpful’ applied broadly across task categories including creative, analytical, and technical work.”
Key Takeaways
- Claude Sonnet 4.6 is the most reliable all-round model at $20/month — 7/7 tasks completed, 4.6/5 average quality.
- ChatGPT o3 produces the highest-quality output on tasks it completes (23/25), but refused 2 of 7 tasks on legitimate professional use cases.
- Gemini 2.5 Pro won on translation and data analysis; its disclaimer habit adds editing overhead on legal and competitive intelligence tasks.
- Benchmark scores (SWE-bench, MMLU) predict coding and reasoning task quality. They do not predict refusal behavior or output length control — the two factors that most affected daily usability in this test.
- All three: $20/month at the consumer subscription level. API pricing differs significantly: Gemini 2.5 Pro at $1.25/M input vs Claude Sonnet at $3/M input vs ChatGPT o3 at higher-end API pricing.
Frequently Asked Questions
Which AI is best for everyday business use in 2026?
Claude Sonnet 4.6 for most use cases. It completed all tasks in this test without refusals or unprompted disclaimers, and its length control is the most reliable of the three models. For specialized coding tasks, ChatGPT o3 produces tighter code when it cooperates. For long-document research, Gemini 2.5 Pro’s 1M token context window is a structural advantage.
Is ChatGPT o3 better than Claude Sonnet 4.6?
On pure coding task quality, yes — ChatGPT o3 scores 88.6% on SWE-bench Verified vs Claude’s 82.1%, and the output quality difference is visible in practice. On task completion reliability across a broad use case set, no. The two content policy refusals in this test represent a real operational risk for teams with varied workflows.
How does Gemini 2.5 Pro compare to ChatGPT for research?
Gemini 2.5 Pro outperforms on long-document research — the 1,048,576-token context window allows loading full reports and filings without preprocessing. For short research summaries from pasted text, ChatGPT o3 was more accurate in this test. The disclaimer behavior Gemini adds to research outputs is a friction cost but not a quality issue.
Does Claude have a content filter like ChatGPT?
Yes, but it triggered zero times across these 7 tasks — including the creative writing task that triggered a ChatGPT refusal. Claude declined to complete tasks in testing involving genuinely harmful outputs (detailed instructions for dangerous activities), but treats professional creative fiction, competitive research, and developer tooling as legitimate use cases without prompting.
Which model should I use for coding in 2026?
ChatGPT o3 for isolated coding tasks where refusal risk is low (algorithmic problems, standard library use, debugging). Claude Sonnet 4.6 for coding workflows embedded in broader pipelines — it’s more likely to complete adjacent tasks without friction. For large-codebase analysis (full repo in a single prompt), Gemini 2.5 Pro’s context window is the decisive advantage. For the CLI/agent coding tools specifically, see our comparison of Claude vs GPT-4o across business tasks which covers agentic use cases in more depth.
Are all three AI models worth $20/month?
Yes, if you’re using them regularly. The question is which one fits your primary use case. The productivity gain from even 10 AI-assisted tasks per week exceeds the subscription cost at any professional hourly rate. If you’re choosing one: Claude Sonnet 4.6 as a general subscription, Gemini Advanced if your work involves long documents and Google Workspace, ChatGPT Plus if your primary use is structured coding or data analysis and you have a low tolerance for the occasional refusal.
Last updated: 2026-07-31
