Quick Answer
Claude vs ChatGPT for writing comes down to polish versus range. Claude writes more natural prose, holds tone across long pieces, and follows detailed style guides more faithfully. ChatGPT is faster for quick variations, short copy, and multi-format work. For voice-driven long-form, Claude wins. For speed and versatility, ChatGPT does.
Claude vs ChatGPT for writing in 2026 comes down to a clear trade-off: Claude is the better stylist, and ChatGPT is the better all-rounder. If you care most about tone, rhythm, and prose that reads like a person wrote it, Claude usually needs less editing to get there. If you want speed, quick variations, and one tool that handles every format you throw at it, ChatGPT is the more flexible partner. Neither wins every writing task, which is exactly why so many writers keep both open and route each job to whichever is stronger. Here's the honest breakdown.
We'll go through how they differ, where each one wins, what the benchmarks say, and how to pick for your own work.
Which Writes Better, Claude or ChatGPT?
For pure writing quality, Claude has the edge in 2026, and most writers who use both say the same. Its prose tends to have more rhythm, better paragraph transitions, and a wider vocabulary, and it executes a specific tone more reliably. Ask both to write a warm but professional apology email, and Claude usually nails the register on the first try while ChatGPT's version reads a touch more generic.
But better is doing a lot of work in that sentence. The gap is real, and it is also narrow, and it flips depending on the task. ChatGPT is no weak writer. It is faster, more versatile, and often the smarter pick when you need many things quickly rather than one thing perfectly. So the honest answer isn't a knockout. It's a points decision that depends on what you're writing.
Which Versions Are You Actually Comparing?
Worth pinning down before you trust any comparison, including this one. Both lineups moved during 2026, and a test run against last year's models tells you very little about what you'd get today.
On Anthropic's side. The newest model is Claude Opus 5.5, released on 22 September 2026, and it's the one Anthropic's model documentation now tells you to start with for most workloads. Above it sits Claude Fable 5.1, released on 1 September 2026 with a 1M token context window according to Anthropic's own model page, and Anthropic suggests reaching for it only when your evaluations on Opus 5.5 at higher effort still fall short. For writing, that's a nudge worth taking. The most expensive model in a lineup isn't automatically the best prose stylist.
On OpenAI's side. The lineup moved in September too. OpenAI's model documentation now leads with the GPT-6 family: GPT-6 Astra as the flagship for complex reasoning, GPT-6 Sol as the balance of capability and cost, and GPT-6 Luna for cheap, high-volume work. The GPT-5.6 models they replaced, Sol, Terra and Luna, are still documented with a 1,050,000 token context window. Free ChatGPT has been running GPT-5.6 Luna, and OpenAI announced on 22 September that free users can now try GPT-6 Luna in the desktop app.
Two practical notes that change what you should expect.
- The chat window isn't the API, and on Claude the context window depends on the model. Those million-token figures used to belong to the developer platform only. That changed. Anthropic's help centre on context windows on paid plans now documents 1M tokens in chat for Opus 5.5, Fable 5.1, Opus 5 and Sonnet 5, with older models such as Opus 4.6 through 4.8 and Sonnet 4.6 at 500K. The pricing page lists the context window as up to 1M, varying by model, on every plan including Free. So what a subscription mostly buys is which models you can use and how much, not raw room.
- None of which changes much for writing. Even a 200K window is roughly 150,000 words. That's a full manuscript, and it's more than almost any writing job needs in one conversation. The 1M window matters if you're feeding in a whole corpus, not if you're drafting a chapter. Don't let the bigger number talk you into a subscription you don't need.
- Free is no longer the throttled experience it was. OpenAI removed text chat caps for free users in August 2026, so free ChatGPT now gives you unlimited everyday text chats. Claude free is narrower than it's often described. Anthropic's pricing page lists Sonnet and Haiku on the free plan, with Opus, Fable and Claude Code reserved for paying users, and up to 5 Projects. Free usage resets on a rolling five-hour window, and Pro gives at least five times as much per session. For a writing test that's usually enough. You can still run a serious comparison yourself without paying either company.
If you want the current specs side by side rather than in prose, our Claude review and ChatGPT review keep the numbers updated.
How Do Claude and ChatGPT Differ as Writers?
Think of it as two different writing personalities. Claude is the careful stylist. It leans toward flowing prose, holds a consistent voice, and tends to avoid the bullet-point-everything reflex. When you give it a detailed brief on tone and audience, it follows the fine print closely.
ChatGPT is the quick, adaptable generalist. It's fast, it brainstorms loosely and happily, it switches between formats without fuss, and it generates images natively, which Claude doesn't do. (Both read images you upload perfectly well. It's making them that only ChatGPT does.) Its default style is a little more structured and list-friendly, which is great for scannable web content and less ideal for a personal essay. That difference in personality runs through every category below. For the wider view across all the big models, see our guide to the best AI for writing in 2026.
Which Is Better for Long-Form Writing?
Claude, and this is where the gap is widest. For anything over about a thousand words, Claude holds a consistent argument and tone from start to finish, where ChatGPT is more likely to drift, repeat a point, or slip back into a listy structure halfway down. The tell is structural: check whether the thesis you set out in paragraph two is still the one being argued in the conclusion, or whether it quietly became a slightly different claim somewhere in the middle. That drift is what people notice in practice.
If you write essays, reports, long articles, or book chapters, that consistency is the whole game. You spend less time stitching a long piece back together after the fact.
And this drift isn't just something writers grumble about. It's measurable, and researchers have built a benchmark specifically for it. Wu, Hee, Hu and Lee at the Singapore University of Technology and Design published LongGenBench, which tests whether a model can keep following instructions across a long piece of writing rather than just answering a short question well. They ran four scenarios, three instruction types, and two output lengths, 16,000 and 32,000 tokens. The pattern was consistent: adherence to instructions drops off noticeably past roughly 4,000 tokens, and it keeps falling as the piece gets longer. Two of their examples show the scale of it. Llama3.1-8B completed 93.5 percent of subtasks at the shorter length and 77.6 percent at the longer one. Qwen2-72B fell further, from 94.3 percent to 66.2 percent.
One finding there is worth flagging because it corrects something people get wrong constantly, and it's a claim you'll see in plenty of comparison articles. A big context window does not mean good long-form writing. The paper found models that score well on long-input retrieval, meaning they can find a fact buried in a huge document, still struggled to hold instructions across a long output, and concluded the two abilities are "not entirely equivalent." So a model advertising a giant context window has told you it can read a lot. It hasn't told you the voice will survive to your final paragraph. Those are different skills, and only the second one matters for writing.
Which Is Better for Short Copy and Speed?
Here ChatGPT often pulls ahead. When the job is ten subject lines, three caption options, a short bio, and a product blurb, all in the next five minutes, ChatGPT's speed and range make it the more practical tool. It's happy to churn out variations, and it switches format on a dime.
ChatGPT is also the better brainstorming partner for many people, since it riffs more freely and can generate images when a draft needs them. So for high-volume, fast-turnaround, multi-format content work, ChatGPT frequently wins on sheer usefulness even if any single sentence is marginally less polished than Claude's. The right tool really does depend on the shape of the task, not just a quality score.
What Do the Benchmarks Actually Say?
Here's the context that keeps writing comparisons honest: the top models are converging fast, so any single-day winner is a snapshot, not a law. According to the Stanford HAI 2025 AI Index Report, the score difference between the top-ranked and tenth-ranked models fell from 11.9 percent to 5.4 percent over the course of 2024, and open-weight models closed the gap with closed ones from 8 percent to just 1.7 percent on some benchmarks in the same year. The frontier is crowded, the pack is bunching up, and the leaders trade places often.
On human-preference rankings, the two sit neck and neck. The best-known measure here is Chatbot Arena, built by Chiang and colleagues at UC Berkeley: people submit a prompt, get two anonymous answers side by side, and vote for the better one without knowing which model wrote either. The 2024 paper describing the platform reported over 240,000 votes at that point, and the count has grown enormously since.
Two things about that leaderboard are worth understanding before you treat any ranking as an answer. First, the scores at the top carry confidence intervals that overlap, so a model sitting a few points above another is often a statistical tie rather than a real lead. Second, the order changes. Every significant release reshuffles it, sometimes within days. Any specific snapshot of who is number one has a shelf life measured in weeks, which is exactly why this guide doesn't quote one.
And there's a third thing, which is less widely known and more uncomfortable. A 2025 paper called The Leaderboard Illusion, by Singh, Kapoor, Koyejo, Longpre, Hooker and colleagues, audited the arena itself and found the playing field isn't level. Big labs get to test many private variants and publish only the best result: the authors identified 27 private model variants tested by Meta in the run-up to one release. Attention is lopsided too, with Google and OpenAI estimated to have received 19.2 percent and 20.4 percent of all arena data respectively, while 83 open-weight models shared an estimated 29.7 percent between them. And that data compounds, because the paper estimates that even limited extra arena data can produce relative performance gains of up to 112 percent on the arena's own distribution.
None of that makes the leaderboard useless. It's still the best large-scale read on what people actually prefer, and it's genuinely valuable for that. But it does mean a gap of a few points between two frontier models tells you less than the number implies, and that a model's position partly reflects how much practice it got on the test. Read it as a rough signal, not a verdict. Benchmarks tell you the frontier models are close. Your own use case tells you which close-but-different style fits your writing.
| Writing task | Winner | Why |
|---|---|---|
| Long-form essays and reports | Claude | Holds tone and argument across thousands of words |
| Natural, human-sounding prose | Claude | Better rhythm, fewer generic filler phrases |
| Following detailed style briefs | Claude | Sticks to the fine print on tone and voice |
| Fast variations and short copy | ChatGPT | Quicker, happy to churn out many options |
| Range of formats and brainstorming | ChatGPT | More versatile, generates images natively |
| Free access with no account | ChatGPT | Offers a logged-out mode |
Can You Trust an AI to Judge Which AI Writes Better?
Not on its own. And this matters more than it sounds, because a lot of comparison articles that say "we tested both" quietly used a third model to grade the results. Once you know how those graders behave, you read every such piece differently, including this one.
The foundational study here is Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena by Lianmin Zheng, Wei-Lin Chiang, Ying Sheng and colleagues, published at NeurIPS 2023. They collected 3,000 expert votes and 30,000 conversations with human preferences, then checked how well a model-as-grader tracked those humans.
Start with the good news, because it's real. A strong model used as a judge hit over 80 percent agreement with human preferences, which is the same level of agreement humans reach with each other. Model graders aren't noise. They're roughly as reliable as another person.
Then the three ways they go wrong, all measured in the same paper:
- Position bias. Swap which answer appears first and the verdict often flips. GPT-4 stayed consistent 65.0 percent of the time; GPT-3.5 managed 46.2 percent, and Claude-v1 just 23.8 percent.
- Verbosity bias. Against an attack that padded an answer into a repetitive list without adding substance, Claude-v1 and GPT-3.5 were fooled 91.3 percent of the time. GPT-4 was fooled 8.7 percent of the time.
- Self-enhancement bias. GPT-4 gave its own answers about a 10 percent higher win rate, and Claude-v1 about 25 percent higher. But the authors were careful here: both also favoured other models, GPT-3.5 didn't favour itself, and they said their data couldn't establish a self-enhancement bias.
Sit with that last row for a second, because it points straight at this article. The claim running through this page is that Claude is the better stylist, and if models do prefer their own writing, a Claude-based grader is the one you'd most want to double-check. The 2023 data didn't prove that, but it's reason enough to ask. If you've read a "we ran a blind test" comparison that landed on the same verdict, the first question worth asking is who or what did the blind judging.
Two honest caveats. Those were 2023-era models, so treat the specific percentages as history rather than a current scorecard; self-preference in model judges has kept turning up in later work, but the exact size for today's versions isn't what's written above. And the numbers say nothing about which model writes better. They're about which model can be trusted to score writing, which is a different job.
What survives is a reading habit. When a comparison tells you one model writes better, check whether a human read the drafts or a model scored them, whether the same two answers were tried in both orders, and whether the winner was simply longer. If none of that is stated, you're reading a preference, not a finding. Which is the honest status of most writing comparisons on the internet, and the reason the self-test further down this page beats any of them for your particular work.
Which One Makes Up Fewer Facts?
Here's the part that cuts against everything above. Claude wins on prose. It does not automatically win on being right, and if you're writing anything that gets published, that's the risk that actually costs you.
Aisha Alansari and Hamzah Luqman at King Fahd University of Petroleum and Minerals built HalluScore, an Arabic question answering benchmark whose questions were deliberately kept because they tend to trigger hallucinations, and ran 17 models across 827 of them. On factual hallucination rate, GPT-5 came out lowest at 25.15 percent. Claude Opus followed at 33.01 percent, and Claude Sonnet 4.5 at 35.79 percent. So on this test, the family many people reach for when they want good sentences stated more things that weren't true.
Read that carefully though, because it's a narrower finding than it looks. HalluScore tests factual question answering, not writing quality. It tells you which model is likelier to state something untrue, not which one drafts a better paragraph. Those really are separate skills, and the honest read is that each tool wins a different one.
Made-up citations are their own problem, and it's worse than most people assume. Delip Rao, Eric Wong and Chris Callison-Burch published a study in April 2026 on reference hallucinations in commercial LLMs, evaluating 10 models and agents. They found 3 to 13 percent of citation URLs were hallucinated outright, and 5 to 18 percent didn't resolve at all. The paper does break results out by model: the lowest hallucinated URL rate was a search-enabled Claude 3.5 Sonnet at 3.0 percent, and the highest was Gemini 2.5 Pro Deep Research at 13.3 percent. Those are older models, so treat it as a snapshot, not a verdict on today's lineup. And one finding is worth knowing before you trust a research feature: the deep research agents produced more citations per query than ordinary search-augmented chat, but hallucinated URLs at a higher rate. More sources on the page, not more reliability.
What this means for your actual workflow is simple enough. Neither tool has earned the right to be quoted unchecked. Click every link before it goes into anything with your name on it. And notice that the better a piece of AI writing sounds, the less it prompts you to check it, which is exactly why the model with the nicest prose can be the more dangerous one to publish straight.
This does complicate the result further down this guide, where Claude scored higher on a Polish and English dental exam. Different tests, different tasks, different answers. That's not the benchmarks contradicting each other so much as a reminder that no single number tells you which model is better.
Will Either One Tell You Your Draft Is Bad?
Probably not on its own, and this is the failure mode that costs writers the most. Generating text is the easy half. The half that actually improves your work is honest feedback on a draft, and both models are trained in a way that quietly pushes against giving it to you.
The behaviour has a name. Sycophancy, meaning the tendency to tell you what you want to hear. Sharma and 18 co-authors documented it across five state-of-the-art AI assistants in Towards Understanding Sycophancy in Language Models, first published in October 2023. Their finding wasn't that one company got it wrong. It was structural: when a response matches your views, it's more likely to be preferred, and both human raters and the preference models trained on them favour convincingly-written sycophantic answers over correct ones a non-trivial share of the time. The training signal rewards agreeing with you.
The most useful measurement came later. Fanous, Goldberg, Agarwal and colleagues at Stanford built SycEval, presented at AIES 2025, testing ChatGPT-4o, Claude Sonnet and Gemini 1.5 Pro on maths and medical question sets. Three numbers matter:
- 58.19 percent of cases showed sycophantic behaviour overall. Not an edge case. The default.
- ChatGPT came out lowest at 56.71 percent, Gemini highest at 62.47 percent. Worth sitting with if you've read this far expecting Claude to sweep every category. Claude sat in between at 57.44 percent. So on this measure it didn't win, and it had the highest regressive rate of the three, 18.31 percent, meaning its cave-ins were the most likely to leave an answer worse.
- Once it caves, it stays caved. Sycophantic behaviour persisted 78.5 percent of the time across follow-up turns, with a 95 percent confidence interval of 77.2 to 79.8.
The researchers also split the behaviour in two. Progressive sycophancy, where caving to your pushback happened to land on the right answer, ran at 43.52 percent. Regressive sycophancy, where it caved into a worse answer, ran at 14.66 percent. So roughly one time in seven, arguing with the model made the output worse.
Two caveats before you take this as a verdict. Those tests ran on 2024-era models (the paper names GPT-4o from May 2024 and lists the others simply as Claude Sonnet and Gemini 1.5 Pro), not the 2026 lineup this guide compares, and both labs have since done explicit work on the problem. And maths and medical questions aren't essays. Sycophancy on a factual answer is measurable in a way that "is my opening paragraph any good" simply isn't.
But the practical lesson holds regardless of which model you're in, and it's the most useful thing on this page:
- Never ask "is this good?" You'll get a yes. Ask for the three weakest paragraphs and why they're weak.
- Don't say you wrote it. Paste the draft as something you're reviewing for someone else. The feedback changes noticeably.
- Ask it to argue the other side. "Make the strongest case that this piece doesn't work" gets you further than any request for improvement.
- Start a new chat instead of pushing back. Given that 78.5 percent persistence figure, once a conversation has drifted into flattering you, it tends to stay there. A fresh context is cheaper than an argument.
Does Using Either One Actually Make You Faster?
Yes, substantially, and there's a proper controlled experiment on it rather than just vibes. It's worth knowing before you agonise over which tool to pick, because the gap between using one and using neither is far bigger than the gap between the two.
Shakked Noy and Whitney Zhang at MIT ran a preregistered experiment on 453 college-educated professionals, published in Science in July 2023. Participants were marketers, consultants, grant writers, HR professionals, data analysts and managers, given incentivised writing tasks from their own occupations. Half got access to ChatGPT at random. Time to completion fell 40 percent. Output quality, graded blind, rose 18 percent.
Two findings underneath the headline matter more for how you use either tool.
It helped weaker writers most. The study found inequality between workers decreased, because the tool compressed the productivity distribution by lifting lower-scoring participants further than the strong ones. If you're already an excellent writer, expect a smaller gain than the 40 percent figure suggests. If writing is the part of your job you dread, expect a bigger one.
It changed what the work consisted of. Participants shifted their effort toward idea generation and editing, and away from rough drafting. That's the single most useful sentence in the paper for anyone choosing between Claude and ChatGPT: the assistant doesn't make you a faster drafter, it moves you out of drafting and into judging. Which reframes the choice. You aren't hiring a writer, you're hiring a first-draft machine whose output you'll spend your time editing. So the question isn't whose prose is prettier in isolation. It's whose drafts leave you with less to fix.
Two caveats before you take the numbers to the bank. It tested ChatGPT specifically, in 2023, so it isn't a verdict on Claude or on either tool as they stand now. And the tasks were short professional pieces, press releases and reports and analyses, not book chapters, so it says nothing about the long-form drift the section above describes. What it does establish is the shape of the benefit, and that shape has held up.
Does the Tool Change Your Own Writing?
Yes, and this is the part almost every comparison skips. The question isn't only which assistant produces better text. It's what happens to your voice when you write alongside one.
Dhruv Agarwal, Mor Naaman and Aditya Vashistha at Cornell ran a controlled cross-cultural experiment with 118 participants, roughly half in the United States and half in India, and presented it at the ACM CHI conference in April 2025 under the title AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances. Participants were randomly split, so half of each country's group wrote culturally grounded essays with AI suggestions and half wrote them without.
Two findings came out of it. The productivity gains weren't shared evenly: American participants got noticeably more efficiency from the assistant than Indian participants did. And the Indian participants' writing shifted toward American style, changing not just what they wrote but how they wrote it. Cornell reported the acceptance numbers underneath that result, and they're the interesting part. Indian participants kept 25 percent of the AI's suggestions against 19 percent for Americans, but frequently had to rewrite what they accepted, which is why they saved less time despite taking more of the help.
Two practical consequences if you're picking between Claude and ChatGPT.
First, if you don't write in a standard American register, whether that's Singlish-inflected Singaporean English, a second language, or a house style with regional idiom, then the assistant that sounds best in a generic test may also be the one sanding your voice down fastest. Judge a draft on whether it still sounds like you, not on whether it sounds polished. Those are different questions and only the first one is yours to answer.
Second, an honest caveat about what this study does and doesn't show. It tested AI writing suggestions as a category, not Claude against ChatGPT, so it is not a verdict on either one. Read it as a property of writing with an assistant at all. It's also one more argument for the test in the next-but-one section: run the comparison on your own material, in your own voice, and check what came back changed rather than just what came back better.
Why Does Either One Drift Off Your Style Over a Long Piece?
Because instruction-following decays as a conversation gets longer, and it happens in both of them. This isn't a Claude quirk or a ChatGPT quirk. It's a property of how these models read a growing context, and it's the single most common reason a draft that started in your voice ends up sounding like everything else.
Vardhan Dongre, Joseph Hsieh, Viet Dac Lai, Seunghyun Yoon, Trung Bui and Dilek Hakkani-Tür looked at the mechanism directly in a May 2026 paper. Their account is that the goal-defining tokens, the part of your prompt that says what you actually want, become harder for the model to reach through its own attention as the exchange grows. To show how sharp that cliff is, they forced the channel closed on one model in a 20-fact retention task and watched recall fall from near-perfect to 11 percent.
They tested four architectures and found the failure doesn't look the same across them. Where a model encodes goal information varied from layer 2 to layer 27, and some models kept behaving correctly even as the attention faded while others fell over despite still carrying the goal in their representations. So the drift is real everywhere, but it arrives differently depending on what you're using.
Now the finding that should change how you work. Junhyuk Choi, Yeseon Hong, Minju Kim and Bugeun Kim tested nine models for a paper on identity drift first posted in December 2024. Two results stand out, and both cut against instinct. Larger models drifted more, not less. And assigning a persona, the thing everyone does when they paste "write in the voice of a warm, plain-spoken editor" at the top, didn't reliably keep identity stable.
Sit with the second one for a second, because a lot of writing advice is built on the opposite assumption.
This also reframes the whole Claude versus ChatGPT question. Nearly every head-to-head, including the benchmark numbers further up this page, tests a single prompt and a single response. That is not how anyone writes a long piece. You work over dozens of turns, and the thing that decides whether you're happy at the end is which tool still sounds like you on turn 40. Neither company publishes drift numbers, and no public benchmark measures it well, so the honest answer is that the one-shot winner isn't guaranteed to be the long-haul winner.
What actually helps, whichever you pick:
- Repeat the style rules late, not just early. Restating your constraints in the message you're sending now beats trusting the ones you set twenty turns ago.
- Break the work into fresh conversations. One section per chat, with your style brief re-pasted at the start of each. It feels wasteful and it holds voice far better than one long thread.
- Show, don't describe. Paste 300 words of your own writing instead of adjectives about your tone. A sample is a much harder target to drift away from than "conversational but authoritative".
- Compare the end against the beginning. Read your last section next to your first, not next to your instructions. Drift is invisible turn by turn and obvious across the whole draft.
And if you want to know which one holds your voice longer, that's a question only your own writing can answer. The test further down this page is worth running over a long piece rather than a paragraph, and our guide to the best AI for writing covers how the rest of the field handles the same problem.
Can Custom Instructions Close the Gap?
Partly. But not in the way most advice tells you, and the popular trick is close to useless.
Everything above compares default behaviour. Both tools let you set standing instructions that apply to every chat, and both let you scope instructions to a single project. So the fair follow-up question is whether configuring them changes the ranking. Two pieces of research say a lot about that, and they point in opposite directions.
The persona trick does almost nothing
Start with the advice you've seen a hundred times: open with "You are an award-winning novelist." Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee and David Jurgens tested exactly that, in work published in Findings of EMNLP 2024 under the title "When 'A Helpful Assistant' Is Not Really Helpful."
They ran 162 personas, spanning 6 interpersonal relationship types and 8 expertise domains, across 4 families of models and 2,410 questions.
Adding a persona did not beat adding nothing at all.
There's a wrinkle that makes it worse rather than better. If you could somehow pick the single best persona for each individual question, accuracy would improve substantially. But automatically finding that persona performed at roughly the rate of picking one at random. So the upside exists and nobody, including the model, can reliably reach it.
One honest limit before you throw the technique out. They tested factual questions, not prose quality, and writing is harder to score than facts. This doesn't prove personas never help your essay. It should just stop you treating them as a free upgrade.
Style instructions do plenty, including things you didn't ask for
Now the opposite problem. Young-Min Cho, Yuan Yuan, Sharath Chandra Guntuku and Lyle Ungar ran what they describe as "the first systematic study of cross-feature stylistic side effects" in 2026. They surveyed 127 conversational agent papers from the ACL Anthology, pulled out the 12 style features people actually use in prompts, then measured what happens to the other eleven when you ask for one.
The headline result is in the title: a concise agent is less expert. As they put it, "prompting for conciseness significantly reduces perceived expertise."
Sit with what that means for your writing. Tell either model to be concise and you don't simply get something shorter. You get something that reads as less authoritative. If you're drafting a proposal or a technical explainer, you just traded away credibility for brevity without being asked whether you wanted to.
Their broader conclusion is that style features "are deeply entangled rather than orthogonal." You aren't turning individual dials. You're moving a cluster.
And you can't easily patch it. They tested both prompt-based and activation-steering fixes and found that while these "can partially restore suppressed traits, they often degrade the primary intended style." Adding "but sound authoritative" after "be concise" tends to cost you the concision.
Read this one with the caution the section above already earned, though. Their measurement used a pairwise LLM-as-judge setup, which is precisely the method with the position and verbosity biases described earlier on this page. Take the direction of the finding seriously and the precision less so.
What actually works
Put those two together and a fairly clear rule falls out. Costumes don't work. Constraints do.
- Write rules, not roles. "Sentences under 25 words. No bullet lists. Never open with a rhetorical question." beats any amount of "You are a brilliant essayist."
- Give it a sample instead of a description. Three paragraphs of your actual writing carries more signal than a paragraph describing your voice. A description is a persona. A sample is data.
- Name the trait you don't want to lose. If you ask for concise, say what must survive it. You won't get both perfectly, but you'll notice the trade instead of absorbing it.
- Put instructions at project level for anything long. This is the practical answer to the drift described in the section above, since a standing instruction gets reapplied rather than fading as the conversation grows.
Does any of this flip the verdict at the top of this page? No. Configuration narrows the gaps rather than reversing them, and the differences in default prose survive a good system prompt on both sides. What it does change is how much editing you do afterwards, which for most people is the number that actually matters. If you want to see how each one behaves once configured, the wider writing roundup covers the setup side in more depth.
Does Either One Train on What You Write?
Possibly, depending on a setting most people never open. If you write under NDA, draft unreleased work, or handle client material, this matters more than any prose-quality comparison, and it's the question writing roundups skip almost universally.
Both companies control it with a single toggle, and it's worth knowing where yours sits rather than assuming.
- ChatGPT. The switch is called "Improve the model for everyone," under Settings then Data Controls. OpenAI's help centre page on how your data improves model performance is the current word on what it does.
- Claude. The equivalent is the Model Improvement setting in Privacy Settings. Anthropic's privacy centre article on model training covers the consumer plans, and notes that Incognito chats stay out of training regardless of how the toggle is set.
Here's the part to hold onto, because it's the bit that survives. Both companies have changed these defaults more than once, and they differ by plan, by region, and by whether you're on a consumer or business account. Any article that tells you flatly which one trains on your text by default is telling you about the week it was written. So don't take that from a comparison guide, this one included. Open the setting in your own account and look.
Two things do hold steady across both. Turning training off is not the same as deleting anything, since past conversations stay where they are and both companies keep data for a period to monitor abuse. And a conversation flagged for safety review can still be looked at either way. So the toggle reduces what your writing gets used for. It doesn't make a chat window a private drafting environment.
Practically: for a blog post or marketing copy, none of this should change your pick. For an unpublished manuscript, a client's confidential brief, or anything under an agreement you signed, check the setting before you paste, and treat that as a separate decision from which one writes better.
Do You Actually Own What It Writes?
You own it in the sense that neither company claims it, and you may not own it in the sense that stops someone else reprinting it. Those are two different questions, and almost every writing comparison answers the easy one and skips the hard one. The good news for choosing between these two is that the answer is identical on both sides, so it's not a tiebreaker. The reason it's worth five minutes anyway is that it changes what you do with the draft after you get it.
Start with what the contracts say, because that part is simple.
- ChatGPT. OpenAI's help centre states plainly that it will not claim copyright over content its API generates for you or your end users, and its terms of use assign you whatever rights in the output it holds. Read that second half slowly. Whatever rights it holds.
- Claude. Anthropic lands in the same place. Its announcement of its commercial terms says customers retain ownership rights over any outputs they generate, and promises to defend them against copyright infringement claims brought over authorised use. OpenAI made the same kind of promise first, calling it Copyright Shield, for ChatGPT Enterprise and its developer platform. Neither indemnity covers a free consumer account. Same structure on ownership, same limit underneath it.
Here's the limit, and it's the reason the first half of this section is the easy half. A company promising not to claim your output, or assigning you the rights it holds, can only move rights that exist. On a piece of machine-generated prose there may be none to move, and no wording in a terms of service page can conjure one. Whether copyright attaches at all isn't a vendor question. It's a copyright question, and it's been answered.
The US Copyright Office published Part 2 of its report on copyright and artificial intelligence on 29 January 2025, after receiving more than 10,000 comments to its notice of inquiry, roughly half of which addressed copyrightability. Its conclusion on prompting is the line to remember: a human-authored work that's perceptible in the output can be protected, and so can a person's creative selection, arrangement or modification of what the model produced, but the mere provision of prompts is not enough. The Office leans on settled ground for that, since someone who describes to an author what a commissioned work should look like has never been a joint author under the Copyright Act.
The courts agree. On 18 March 2025 a unanimous DC Circuit panel decided Thaler v. Perlmutter, holding that a copyrightable work must be authored in the first instance by a human being, and affirming the refusal to register a wholly machine-generated image. The Supreme Court declined to take the case in 2026, which leaves that holding standing. What the panel explicitly did not decide is where the line falls, meaning how much human input turns an AI-assisted draft into a protectable work. Nobody can tell you the threshold, because nobody has drawn it yet.
So what does this mean when you're staring at a draft either tool just produced?
- Your editing isn't just quality work, it's the part you can own. Rewriting, cutting, restructuring and threading your own material through the piece are exactly the acts the Office named as protectable. Accepting a draft as it arrived is the version with the weakest claim.
- Keep the drafts. If a piece ever matters commercially, the record of what you changed is the evidence of the human authorship. That costs you nothing to do now and is impossible to reconstruct later.
- Don't assume unique output. Neither company promises that what you got is what only you got. Whatever rights either one passes you cover your output, not somebody else's, and two people prompting for the same listicle intro can land in much the same place.
- This applies equally to both. No one should pick Claude over ChatGPT, or the reverse, on ownership grounds. The prose gap is real. This gap isn't.
One more layer if you write for publication rather than for yourself. The International Committee of Medical Journal Editors says chatbots cannot be listed as authors, because they can't be responsible for the accuracy, integrity and originality of the work, and those responsibilities are what authorship means. Its recommendations, updated in January 2026, also expect authors to describe how they used the tool. The Committee on Publication Ethics reaches the same result by a different route, pointing out that AI tools aren't legal entities, so they can't declare conflicts of interest or hold copyright and licence agreements in the first place.
That's a neat summary of the whole section, actually. The tool can't be the author, which is why it can't give you an author's rights, which is why what you do to the draft is the thing that decides what you've got. Both of these models will write you something publishable. Neither of them will make you its author.
Can Anyone Tell Which One Wrote It?
Not reliably, and the way detectors fail matters more than the fact that they do. If you're picking between these two partly because you're wondering which one is less likely to get flagged, that's the wrong question, and the research on why is worth two minutes of your time.
Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou at Stanford ran seven GPT detectors over essays written by humans and published the results in Patterns, indexed on PubMed. On TOEFL essays, written by non-native English speakers, the detectors returned an average false positive rate of 61.3 percent. Every single detector unanimously flagged 19.8 percent of those human-written essays as AI. At least one detector flagged 97.8 percent of them. On essays by US students, the same detectors were accurate.
Read that again, because it isn't a story about AI writing at all. Those essays were written by people. The detectors weren't spotting machine text; they were spotting a smaller vocabulary and more predictable sentence construction, which is what writing in your second language looks like. As the paper itself put it, the unanimously flagged essays had significantly lower text perplexity, a point Stanford HAI echoed in its write-up. The tool measured plainness and called it fraud.
And the finding that should end the conversation: when the researchers ran those same TOEFL essays through ChatGPT to enrich the vocabulary, the average false positive rate fell from 61.3 percent to 11.6 percent. Using an AI to improve the writing made the detectors more confident a human wrote it. Whatever these tools measure, it isn't authorship.
So what does this mean for choosing between Claude and ChatGPT?
- Don't pick one to beat a detector. The gap between the two on any detection score is noise next to a 61 percent false positive rate. You'd be optimising against a broken instrument.
- Heavier editing helps, but not for the reason you think. Rewriting in your own voice makes the text better and less generic. That it also drifts away from whatever a detector keys on is a side effect, not a strategy.
- If you're on the receiving end of an accusation, the base rate is your argument. A detector score is not evidence. Ask what the tool's false positive rate is on writing like yours.
- If you run a policy that uses these tools, this is the study to read first. The cost of a false positive lands hardest on international students, and it lands through no fault of their writing.
None of which means disclosure doesn't matter. It means detection and disclosure are separate things: one is a broken measurement, the other is a choice you make. Our sister site covers where disclosure is actually required, including the EU AI Act transparency rules.
Does It Matter Which Language You Write In?
It matters, but not in the direction most people assume. The usual guess is that ChatGPT falls apart outside English and Claude holds steady. The evidence is messier than that, and worth walking through because the popular version is wrong in a specific way.
Start with the study that tested it head to head. Wójcik and colleagues published in Scientific Reports in 2025, running 198 questions from Polish dental licensing exams, the LDEK and LDEW, through each model three times in English and three times in Polish, 1,188 prompts per chatbot in all.
Claude was the most accurate overall, and its gap between English and Polish wasn't statistically significant (p = 0.278). ChatGPT-4 scored lower but was just as language-stable (p = 0.746), and it actually did slightly better in Polish. The model that cracked was Gemini, whose accuracy dropped substantially in Polish, a significant gap at p < 0.001.
So on this test the accuracy ranking and the language-stability ranking are two different things. Claude won the first. Claude and ChatGPT tied on the second, and the model people rarely worry about is the one that fell over.
But don't over-read it, for two reasons. It's 198 questions, which is small. And Polish is a well-resourced language with a large digital footprint, so it's close to the easy case.
Push further down the resource curve and the picture changes. Viet Dac Lai and colleagues evaluated ChatGPT across 37 languages and seven tasks in 2023, deliberately grouping languages by how much training data exists for them, from high-resource down to extremely low-resource. Performance generally fell as the resource level dropped, with the biggest gaps in lower-resource languages, though the authors note data size wasn't the only factor.
Put the two together and you get a rule of thumb that is more useful than "Claude is better at languages":
- Writing in a major European or East Asian language? Both handle it. Pick on prose quality, the same way you would in English.
- Writing in a smaller or less-digitised language? Expect degradation from either one, and check the output rather than trusting it. The gap here is between languages, not between the two products.
- Writing in English as a second language? This is the case worth flagging, and it connects to the detection section above. The Stanford finding that detectors misfire on non-native English writers means you can be penalised for writing that is entirely your own. Oddly, the same study found that running those essays through ChatGPT to enrich the vocabulary cut the false positive rate from 61.3 to 11.6 percent. That says the detectors measure vocabulary rather than authorship. It isn't a reason to launder your own writing through a model.
One honest caveat about all of this. Both studies mostly measure whether the model got the answer right, not whether it wrote well. Accuracy across languages and prose quality across languages are different questions, and the second one has far less research behind it. If your Spanish or Japanese copy has to actually sound good rather than just be correct, the only real test is the one in the next section, run in your own language.
Does Either One Work Better If Writing Is Hard for You?
This question gets skipped in almost every comparison, which is odd, because it's one of the few where the answer might actually change your life rather than your afternoon. If you're dyslexic, or writing in your second language, or you have ADHD and the blank page is the whole problem, "which model has nicer prose" isn't really what you're asking.
The honest answer is that neither has been properly tested for this, and what testing exists is more sobering than the marketing suggests.
The closest thing to real evidence is LaMPost, an AI email-writing prototype built specifically for adults with dyslexia and presented at the ACM SIGACCESS conference on accessibility in October 2022. The researchers tested it with 19 dyslexic adults using features like outlining, subject line generation and rewriting a selection. Two features landed well, rewriting and subject lines. But the study's conclusion was blunt: the current generation of language models may not clear the accuracy and quality thresholds that writers with dyslexia actually need.
That was an earlier generation of models and things have moved since. Still, the finding points at something that hasn't changed. A tool that produces confident, fluent text you can't easily verify is a different proposition when checking the output is the exact part you find hard. Fluency isn't the accessibility feature. Verifiability is, and neither Claude nor ChatGPT is built around that.
One more finding from that study worth holding onto: whether participants knew the AI was involved made no measurable difference to their sense of autonomy or self-efficacy. The worry that using assistance would make people feel less like the author didn't show up in the data.
There's a separate trap if English isn't your first language. As covered further up, the Stanford HAI work on detectors found they misclassify non-native English writing as AI-generated at high rates. So you can write something yourself and have it flagged. The perverse part is that polishing it with a model made the detectors less likely to flag it, which tells you what they're actually measuring. Neither tool solves that, and picking between them won't either.
If writing is genuinely difficult for you, judge these tools on a different axis than the reviews use. Does it let you talk instead of type. Does it break a task into steps rather than handing back a finished block. Does it explain a change so you learn the pattern, or just silently make it. Those questions matter more than which one has better rhythm, and the answer varies more by interface than by model.
Which Should You Choose for Your Writing?
Pick Claude if the words are the product. Essays, long articles, thought-leadership pieces, anything with a distinct voice, or work where a client will read every sentence. It gets you closest to a finished draft with the least editing. Pick ChatGPT if you need speed, volume, and range, like social content, quick drafts, outlines, and multi-format campaigns where done-fast beats perfectly-phrased.
But the smartest move for most writers is to use both, because both have genuine free tiers. Draft and brainstorm quickly in ChatGPT, then move the piece to Claude for a polish pass on tone and flow. Or start the long-form draft in Claude and use ChatGPT for the ten caption variations after. They cost nothing to try, so set both up and let each do what it does best. Still weighing it against Gemini too? Our Gemini vs Claude comparison and our guide to the best free AI chatbot cover the full field.
How Can You Test Which One Writes Better for You?
Because rankings move and taste is personal, the only comparison that stays true is the one you run yourself. It takes about twenty minutes and both free tiers are enough.
- Pick a real piece of your own work. Not a generic prompt. Something you've actually written and know how you'd want it to sound, so you can judge the output against a standard you hold rather than a vague sense of good.
- Give both the identical brief. Same prompt, same context, same constraints, pasted the same way. Any difference in what you feed them makes the test worthless.
- Make it long enough to matter. Ask for at least 600 words, and go longer if you can. Short outputs hide the thing that actually separates these models, which is whether the voice holds up across a whole piece or drifts halfway through. The LongGenBench finding above puts a number on this: instruction adherence starts slipping past about 4,000 tokens, which is roughly 3,000 words. So 600 words is a usable floor for a quick test, but if you routinely write long pieces, test at the length you actually write at. That's where the gap shows up.
- Include a constraint you care about. A banned word, a required structure, a specific reading level, a tone. Then check whether each one honored it all the way to the end or quietly dropped it after the first few paragraphs.
- Measure edit time, not first impressions. Take both drafts to publishable and time it. The one you finish faster is your answer, and it's often not the one that read better in the first paragraph.
Run that on three different kinds of work you actually do. You'll usually find the split isn't one winner but a division of labour, and knowing your own split is worth more than any leaderboard position.
What Else Do People Ask?
Is Claude better than ChatGPT for writing in 2026?
For most writing where voice and nuance matter, yes. Claude tends to produce more natural prose, keeps tone consistent across long pieces, and leans on fewer generic filler phrases than ChatGPT. But ChatGPT is faster for high-volume variations and handles a wider range of formats. Neither wins every task, so the better writer depends on whether you value polish or speed and range.
Does ChatGPT or Claude sound more human?
Claude usually sounds more human out of the box. Writers describe its prose as having better rhythm, smoother transitions, and a wider vocabulary, with less of the hedging and list-heavy structure ChatGPT defaults to. ChatGPT can match that quality, but it often takes a more detailed prompt to get there. For a first draft that reads like a person wrote it, Claude needs less coaxing.
Which is better for essays and long documents?
Claude, for anything over about 1,000 words. It holds a consistent argument and tone across a long piece without drifting or repeating itself, which is where ChatGPT is more likely to lose the thread. The difference shows up in structure: whether the thesis stated up front still governs the piece by the final section, or has quietly been replaced by a different one. For long-form coherence, Claude is the stronger pick.
Is ChatGPT better for anything in writing?
Yes, plenty. ChatGPT is faster when you need ten quick variations, a tight product description, a short email, or a structured outline. It also handles a wider range of tasks, generates images natively where Claude does not, and is often more willing to brainstorm loosely. For rapid, high-volume, or multi-format work, ChatGPT is frequently the more practical writing partner.
Can you use Claude and ChatGPT for free?
Yes, both have real free tiers with no credit card needed. ChatGPT even offers a logged-out mode with no account. ChatGPT's free plan now has unlimited everyday text chats, with separate limits on files, images and voice. Claude's free plan has usage limits that reset on a rolling five-hour window. Many people run both free tiers side by side and route each writing task to whichever one does it best.
Sources: Stanford HAI, 2025 AI Index Report, on top-model score convergence during 2024 (hai.stanford.edu/ai-index/2025-ai-index-report). Chiang et al., Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, UC Berkeley, 2024 (arxiv.org/abs/2403.04132). Singh et al., The Leaderboard Illusion, 2025, on private variant testing, uneven data allocation, and arena overfitting (arxiv.org/abs/2504.20879). Zheng, Chiang, Sheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023 (arxiv.org/abs/2306.05685), 3,000 expert votes and 30,000 human-preference conversations, for the over-80-percent judge-human agreement figure and the position, verbosity and self-enhancement bias numbers on 2023-era models. Wu, Hee, Hu and Lee, LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs, Singapore University of Technology and Design, 2024 (arxiv.org/abs/2409.02076), on instruction-adherence decay past roughly 4,000 tokens and the gap between long-input retrieval and long-form generation. Agarwal, Naaman and Vashistha, AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances, Cornell, ACM CHI 2025 (arxiv.org/abs/2409.11360), N = 118, on uneven efficiency gains and stylistic homogenization; suggestion-retention figures of 25 percent versus 19 percent via the Cornell Chronicle announcement. Noy and Zhang, Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence, MIT, Science 2023, 381(6654):187 to 192 (pubmed.ncbi.nlm.nih.gov/37440646), N = 453, on the 40 percent time reduction, 18 percent quality gain, compressed productivity distribution, and the shift from drafting toward idea generation and editing. OpenAI, How your data is used to improve model performance, help centre (help.openai.com/en/articles/5722486), and Anthropic, Is my data used for model training, privacy centre, consumer plans (privacy.claude.com/en/articles/10023580), on the training toggles, retention, and the safety-review exception. All linked above. Rankings on live leaderboards and data-training defaults both change frequently, so this guide describes how the comparison behaves and where to check the settings yourself rather than quoting a dated snapshot.
Find more free AI tools at SpotFreeAI.com and ready-made prompts at PromptCraftAsia.com.