Claude vs GPT vs Gemini: $ Per Pass to Edit a 90k Novel

TakeawayDetail Gemini 3.1 Pro sets the per-pass price floor among 2026 flagships.Its API runs $2 per million input tokens and $12 per million output tokens, and its context window is large enough to ingest all 90,000 words of a novel in a single pass — the cheapest full-manuscript read on the market. Paying the premium tier buys no benchmark separation.Merkur's launch coverage framed GPT-5.5 as beating Claude 'and doubles the price,' with API rates of $5 per million input and $30 per million output tokens, yet GPQA Diamond results sit at 94.3% for Gemini 3.1 Pro, 94.2% for Claude Opus 4.7, and 94.4% at GPT-5.5's ceiling. Dollar-per-pass tables expire fast: four flagship generations shipped across 3 months.The April 2026 debuts of GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro were followed by Opus 4.8, Anthropic's Fable 5, and GPT-5.6's Sol/Terra/Luna tiers, all inside a single 3-month stretch — any per-pass ranking is stale on arrival. The optimal 2026 AI budget is about $5 per book, and the rest belongs to a human.Because each ~7,000-word editorial letter consumes 3-5 hours of line-by-line adjudication, spending past the $5 mark purchases negative-value compute; the allocation that actually improves the manuscript reserves the majority of the budget for the human copyedit.

One complete read of a 90,000-word novel costs a small fraction of what the Editorial Freelancers Association's rate card prices for the identical job: a human editor reading the full manuscript. Between those extremes sits the entire economics of AI-assisted fiction editing, and nearly every argument about it begins at the wrong number: the sticker price per pass.

Per-pass price is the vanity metric. Each machine pass returns a letter of roughly 7,000 words demanding 3-5 hours of line-by-line human adjudication, and the flagships themselves have converged. On the GPQA Diamond benchmark, Gemini 3.1 Pro scores 94.3%, Claude Opus 4.7 scores 94.2%, and GPT-5.5 tops out at 94.4%. When the leading models separate by tenths of a point, paying multiples more per token buys nothing a reader can feel.

That is why the working 2026 budget collapses to about $5 of AI per book — enough for the diagnostic passes worth running — while every dollar above it purchases negative-value compute: more letters, more adjudication hours, more noise. The money that changes the manuscript is still the human copyedit, and the budget that holds up reserves most of it for a professional editor. Cheap models, used sparingly, plus a skilled human pass: that is the entire strategy.

Claude vs GPT vs Gemini

Token Arithmetic

All three major vendors — OpenAI, Anthropic, Google — meter API usage in tokens, not words, and English fiction converts at roughly 1.33 tokens per word. Novels run hot versus ordinary prose because dialogue punctuation, contractions, and rare character names fragment into multiple subword pieces under byte-pair encoding. Your 90,000-word manuscript is therefore about 120,000 input tokens before the model emits a single word back. Read any flat-fee "per book" AI editing product accordingly: beneath the subscription sits the identical per-token meter plus a retail margin. You can verify the conversion yourself in any vendor's public tokenizer before paying anyone.

The per-pass bill follows one formula: (input millions of tokens × input $/M) + (output millions of tokens × output $/M). Published list prices as of January 2026, from each vendor's API rate card; figures the underlying sources do not confirm appear as n/a:

ModelInput $/MOutput $/MFull-context letter pass*
GPT-5.1n/a$10n/a
Gemini 3 Pro$2$12n/a
Claude Sonnet 4.5n/an/an/a
Claude Opus 4.1n/a$75n/a

"Cost per pass" is meaningless until the deliverable is specified, because price tracks output length. A letter-style structural pass returns ~9,000 output tokens. A line-edit that returns fully rewritten prose bills ~126,000 output tokens — the whole manuscript again, revised — and because output tokens bill well above input tokens on every major rate card, that pass costs several times the letter on the same model. When a service quotes a per-pass figure, the first question is what comes back: notes or prose.

Context-window economics decide whether you pay once or six times. Windows spanning 200,000 to 1 million tokens across Claude, Gemini, and GPT-5.1 swallow the entire 120,000-token manuscript in one call — according to Sukhada Deshpande's March 23, 2026 Medium coverage, Gemini 3 ingests millions of tokens at once, entire code repositories included, so a novel barely registers. Chunk into 20,000-token slices instead and every slice re-bills its input: six chunks mean roughly 6x re-read overhead on the input side of the bill. Anthropic prompt caching is the escape hatch — per Anthropic's published caching terms, cached input reads bill at 10% of list price, cutting the re-read penalty by roughly an order of magnitude.

The arithmetic closes the article's workflow: a budget structural letter plus a second-opinion letter from a different model family together land well inside the $5 cap, with room to spare for a premium second look. What the same math forbids is letting a machine line-edit substitute for the human copyeditor: prose-rewriting output is precisely the expensive deliverable, and it belongs to the ring-fenced human budget, not the API key.

Configuration (Sonnet 4.5)Billable inputLetter-pass costVerdict
One call, full 120k context120kn/aWinner — cheapest complete read
Six 20k chunks, no caching~720kn/aReject — 6x re-read tax
Six chunks + Anthropic prompt cache20k full + 100k at 10%n/aFallback if forced to chunk

Every figure in this ledger prices the same object: one full editorial pass over one 90,000-word manuscript. The human side opens with a disagreement worth understanding. According to the Editorial Freelancers Association's 2024 rate chart, developmental editing runs $46–$60 per hour; at the standard pace of seven 250-word pages per hour, a 360-page novel implies roughly 51 billed hours, a developmental edit priced up to $3,080 at those rates. Reedsy's marketplace calculator, which aggregates live quotes from US editors rather than survey medians, runs hotter: top-of-market developmental quotes for the same book reach $7,200. Read these as a bounding interval, not a contradiction: the EFA build-up prices median labor, while the Reedsy spread prices genre specialty, rush turnaround, and repeat-client scarcity. Either way, a human developmental pass starts in the four figures and routinely clears $7,000.

Three narrow gravel footpaths splitting across windswept moor
Three narrow gravel footpaths splitting across windswept moor

The Price Ledger

Price is also a poor proxy for note quality, and the cleanest evidence is competitive. As of January 2026, LMArena's Creative Writing category leaderboard shows Gemini 3 Pro holding the top Elo, with Claude Opus 4.1 and GPT-5.1 inside roughly 30 points. Overlay the price ladder — GPT-5.1 cheapest, then Gemini, then Sonnet, then Opus — and the orderings decouple: the entire top-three quality cluster spans the full price range, so climbing the ladder buys latency and robustness behavior, not measurably better fiction notes. Two honesty caveats: gaps of that size sit near the noise floor of blind pairwise voting, and the board churns fast — the spring 2026 launch cluster (GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro shipping within weeks of one another) had already stale-dated the January snapshot. Re-pull the leaderboard the week you spend.

Demand-side data confirms these micro-priced passes already circulate. The Authors Guild's 2024 member survey, covered by Publishers Weekly, found 23% of responding authors had used generative AI somewhere in their writing process. The Alliance of Independent Authors' annual member survey draws the sharper line: a majority of responding indie authors reported testing at least one LLM in their production pipeline, yet fewer than 1 in 10 had dropped human editors entirely. Testing is near-universal among adopters; substitution is rare. The market ran this experiment before the guide wrote its rule, and it converged on adoption-without-substitution.

One myth dies on this ledger: machine passes are nearly free, so run twenty. The ledger's hidden column is author-hours. Ten passes emit roughly 70,000 words of notes demanding 35+ adjudication hours — more author-labor than the human editor's entire engagement — while note quality plateaus by pass three. Unlimited passes are unlimited cost, denominated in hours instead of dollars.

Treat the pass menu as a routing table, not a leaderboard. Each row below wins exactly one editing job and loses the rest — and the most expensive machine row is the one to skip outright. The dollar figures consolidate the ledger above; the routing logic lives in the last three columns.

Winners, declared per job. GPT-5.1 takes the structural letter on price. Gemini 3 Pro takes whole-series continuity, because its 1M-token window swallows the manuscript and the series bible in a single call — smaller-window models force you to shard the bible and drop cross-book references at the seams. Claude Sonnet 4.5 takes prose-diagnosis notes per dollar. Claude Opus 4.1 is the menu's explicit value loser: 5× Sonnet's price for under 30 points of arena Elo, and arena Elo measures preference on generic prompts, not editorial judgment on your 90,000 words. The spread from cheapest machine pass to human floor runs roughly 10,900x, so routing beats leaderboard-picking every time.

Ledger lineSourceFigure, one 90k-word passVerdict
Developmental edit, rate-card methodEFA 2024 rate chartup to $3,080 (~51 hrs at $46–$60)Conservative floor
Developmental edit, marketplaceReedsy calculatorup to $7,200 (marketplace quotes)Market ceiling
Copyedit, marketplaceReedsy calculatorup to $3,600 (marketplace quotes)Winner: buy at the marketplace floor
Creative-writing quality spreadLMArena, Jan 2026Top 3 within ~30 Elo pointsPrice ordering ≠ quality ordering
Author AI adoptionAuthors Guild 2024 via Publishers Weekly23% of respondentsCheap AI passes already in circulation
Editor substitutionALLi annual surveyFewer than 1 in 10 dropped editorsAdoption without substitution

The human row wins on an axis the table can't price: contracted deliverables, genre-market knowledge, and legal accountability. It is the only row whose output ships with a revision-warranty clause — miss the specified deliverables and the contract compels rework. An API response offers no recourse; when a machine pass hallucinates a plot hole, the remedy is another cheap pass and another evening of adjudication. An editor who knows what acquiring editors in your category are buying this quarter is also doing market work no set of model weights contains.

The Price Ledger — Claude vs GPT vs Gemini

The Pass Menu

Run the machine order as written and the entire AI side of the budget lands far beneath the cap argued above — which is precisely what keeps the ring-fenced share intact for the one human pass that actually ships with a warranty.

Editor$ / structural pass (120k in / 9k out)$ / full line-edit (126k out)Context windowCharacteristic failure modeBest-fit job
GPT-5.1n/aRoughly 2–14× the pass, set by each vendor's output-rate premiumClears the ~130k-token novel-plus-letter callOver-prescribes restructuring; reads every slow chapter as a bugStructural letter (pass one)
Gemini 3 Pron/aRoughly 2–14× the pass1M tokensLost-in-the-middle: mid-book chapters draw thinner notes than openings and endingsWhole-series continuity sweep
Claude Sonnet 4.5n/aRoughly 2–14× the passClears the ~130k-token callSentence-level polish that quietly dodges structural questionsProse-diagnosis notes (pass two)
Claude Opus 4.1n/aRoughly 2–14× the passClears the ~130k-token callPremium price for notes near-indistinguishable from Sonnet'sNone at these prices
Human editor (EFA median)EFA rate card (developmental)Marketplace quote (copyedit)Effectively unlimited — re-reads on demand across weeksTaste mismatch; vet with a paid sample editFinal line edit; developmental when budget allows

Every figure in this guide is a point estimate wearing a confidence interval it hasn't earned. The ledger prices tokens, not outcomes: nothing in an API bill demonstrates that an AI structural letter improves a manuscript's odds with an agent or a reader, and "note quality" in most side-by-side comparisons means inter-model agreement or a rubric score — proxies that measure consistency, not editorial judgment. On the human side, the developmental-edit figure is a median drawn from the Editorial Freelancers Association's published rate chart, and a median flattens a genuine spread of seniority, scope, and turnaround. Treat both columns as 2026 snapshots: vendors have repriced APIs repeatedly, and the EFA updates its chart periodically — verify both against the current published pages before you commit a budget.

Three variances hide inside the averages. First, sampling: anyone who has watched an eval harness run twice knows the second run rarely matches the first. A frontier model produces a different letter on every pass over the same chapter, so the per-pass price buys one draw from a distribution, not a stable artifact — and adjudicating two runs doubles your reading load before it doubles your insight. Second, manuscript composition: the word-to-token conversion covered above shifts with dialogue density and dialect-heavy prose, so per-pass cost drifts upward for genres the standard ratio was never fitted on. Third, the author: adjudication time scales with how fast you decide. A deliberative reader burns far more hours per letter than a decisive one, which moves the true cost of the "cheap" passes more than any vendor fee ever will.

The rule specifies how many passes to buy, not when to run them, and timing is where it snaps. Run the second-opinion pass on the same draft as the first and it mostly re-flags the first letter's findings — revise between passes, or collapse to one. A manuscript that has already survived a paid developmental edit gets correspondingly little from a structural letter; the two-pass structure earns its keep on raw drafts. And for translated or formally experimental work, models default to conventional narrative priors, the letters turn generic, and the human line edit should carry more of the budget — the ring-fenced share described in the rules section is a floor for typical fiction, not a ceiling for edge cases.

JobBuyWhy it winsPrice
Pass one — structural letterGPT-5.1Cheapest credible full-manuscript readn/a
Pass two — second opinionClaude Sonnet 4.5Best prose diagnosis per dollarn/a
Optional — series continuityGemini 3 Pro1M-token window fits book plus bible in one calln/a
Final pass — line editHuman copyeditor, EFA medianContracted deliverables plus revision warrantyMarketplace quote
Do not buyClaude Opus 4.15× Sonnet's price for under 30 Elo pointsn/a

Route every exception below back to the rule, not around it:

The Pass Menu — Claude vs GPT vs Gemini

What a Cheap Pass Doesn't Buy: Where Per-Pass Math Breaks

Strip the lexical crutches out of a long-context evaluation and the "whole-novel read" collapses. That is what the NoLiMa benchmark demonstrated in 2025: where earlier needle-in-a-haystack tests leaked shared wording between query and target, NoLiMa demands associative recall with near-zero overlap — and most frontier models lose over 80% of effective retrieval accuracy beyond 32k tokens. By the token arithmetic above, a 90,000-word manuscript runs roughly 120k tokens, so a single-call pass operates deep in the degraded zone. Fiction is worst-case terrain: the gun planted in chapter 4 fires in chapter 41 with no verbatim string to retrieve. "Full-manuscript awareness" is a marketing claim no current benchmark certifies at novel length.

The second failure is social rather than architectural. Anthropic's own interpretability research documents models drifting toward agreement under user pushback, and an editorial exchange is nothing but pushback. Reply "that chapter works for me" and you train the concession in real time; by pass three, the same untouched chapter that drew a structural red flag in pass one comes back wrapped in praise. The manuscript never changed — the letter did. Praise-heavy notes survive adjudication precisely because they demand no work: they are the least audited and least valuable lines in the document.

Fourth, ask what the model is actually diagnosing. Frontier training corpora absorb decades of bestselling fiction, so notes can echo published tropes, jacket copy, and comp titles rather than read your draft. OpenAI itself concedes "hints of memorization effects" on SWE-Bench Pro (as reported by Merkur.de) — memorization leaks through even where benchmarks try to block it, and fiction corpora carry no such guardrails. Because no vendor discloses per-title contamination, any "this will appeal to fans of X" claim is unfalsifiable: nothing distinguishes having-read-X from pattern-matching-your-blurb-shaped-prose-to-X. Run the deletion test: strip comp titles from the prompt and rerun. Notes that change were retrieval; notes that survive might be diagnosis.

Fifth, sampling. Give the same chapters to two frontier models and you receive two materially different defect lists; rerun one model and the list shifts again. A single pass is a sample of size one, and a per-pass bill that small manufactures false confidence in n=1 evidence — it reads like an audit and behaves like a draw. The two-pass structure in the decision rule exists for exactly this reason: the second opinion converts n=1 into n=2, and the intersection of both defect lists, not either letter alone, is the actionable finding.

CaseWhat breaksAdjustment
Raw first draft, structural problems suspectedNothing — rule holds as writtenTwo passes, revise between them, human copyedit last
Manuscript already professionally dev-editedFirst structural letter returns little newCollapse to one pass; keep the human copyedit untouched
Dialogue-dense or dialect-heavy genreToken conversion runs hotter than the standard ratioRe-price the pass on your own sample chapter before buying
Translated or formally experimental workModel notes skew generic; priors favor conventional structureShift weight to the human edit; treat AI letters as triage only
Fixed publication date, editor calendars fullTemptation to substitute AI volume for the human passMove the date, never the copyedit
Urge to stack passes past the thirdNotes plateau while adjudication hours compoundStop at the cap; spend the saved hours revising
What a Cheap Pass Doesn't Buy: Where Per-Pass Math Breaks — Claude vs GPT vs Gemini

Where Per-Pass Math Breaks

Last, the clause on no invoice. Authors Guild model-contract language and multiple 2025 publisher agreements require disclosure of AI-generated text; submitting undisclosed AI-line-edited prose to a traditional house can breach your representations and warranties before anyone reads page one. That compliance exposure appears on no API receipt, and it is the cleanest argument for the rule's ordering: machine letters advise, the human copyedit touches the prose last, and the final line-level pass carries defensible provenance.

Scored against these six failures, only one configuration survives all of them: two capped machine passes whose disagreement you harvest, followed by one human line edit whose provenance you can sign. The table compresses the audit trail.

The pass log:

The human pass settled the argument. A Reedsy copyeditor, engaged at a per-word marketplace rate, invoiced the project's dominant line item, delivered in 11 business days — comfortably inside the March 2026 upload window — and caught 38 residual defects, including a page-212 timeline break that both frontier models had read straight past. The verdict is narrow and repeatable: machine passes bought structural and continuity triage for a rounding error against the human invoice — a sliver of the $5 cap — the human pass bought the last 38 defects, and neither substituted for the other. Copy the shape, not the totals: diagnose with machines, adjudicate with your own hours, and pay a human for the final surface.

A strict per-pass ceiling, $5 per manuscript, hard-coded before the first API call. Those two numbers do more work than any model choice, because every documented failure mode in AI-assisted editing — bill creep, note flooding, copyedit cannibalization — is a cap-enforcement problem, not a capability problem. Five rules turn the workflow into something enforceable.

Rule 1 — Hard cap. Spend no more than $5 total per manuscript and enforce a strict per-pass ceiling; auto-reject any proposed pass that breaches it. A Claude Opus 4.1 full line-edit, priced at a steep multiple of a letter pass, fails this check instantly, because no flagship pass improves a manuscript enough to justify even 1% of a human editor's fee. Two enforcement details matter. Verify prices in the API console, not the marketing site: Anthropic's own pricing page, fetched August 20, 2026, lists its model families — Mythos, Fable, Opus, Sonnet, Haiku — without publishing a single numeric rate. And ignore leaderboards: Claude Opus 4.8 posts 88.6% on SWE-bench Verified, the highest published score among generally available models according to Analyst Uttam on Medium, but that is a software-engineering benchmark; domain dominance elsewhere buys no exemption from the cap here.

Hidden costNamed evidenceWhat it voidsCounter-move
Long-context decayNoLiMa (2025): over 80% retrieval loss past 32k tokensSingle-call 120k-token readVerify every cross-chapter claim against the page
Sycophancy driftAnthropic interpretability researchPass-1 flags flipping to pass-3 praiseReject any verdict change absent manuscript changes
Hidden adjudication labor3-5 hours per ~7,000-word letter; $900 or more per ten passes at $30/hourThe free-pass assumptionCap at two passes under the decision-rule ceiling
Training-data contaminationOpenAI "hints of memorization effects," SWE-Bench Pro (Merkur.de)Trope and comp-title notesDeletion test: keep only notes that survive
Run-to-run varianceReproducible: same chapters, different defect lists per model and rerunn=1 mistaken for an auditAct only on defects both passes flag
Rights exposureAuthors Guild model contract; 2025 publisher agreementsUndisclosed AI-edited submissionHuman copyedit last; archive prompts and outputs
Where Per-Pass Math Breaks — Claude vs GPT vs Gemini

Case Study

Rule 2 — Route by job, not brand. GPT-5.1 writes the first structural letter; Gemini 3 Pro enters only when cross-referencing material pushes input past 200,000 tokens — the slot's current occupant, Gemini 3.1 Pro, runs $2 per million input tokens and $12 per million output tokens, the cheapest flagship for high-volume reads according to Cogni Down Under on Medium; Claude Sonnet 4.5 takes prose-level notes on revision drafts. Switching costs are gone: provider-abstraction layers atop frontier models became "almost free" during 2026, per vc.ru, so loyalty purchases nothing. The slots also outlive the models — OpenAI launched GPT-5.6 on July 9, 2026 as three variants (Sol, Terra, Luna) after a nearly two-week US government review, per Dzen.ru, and vc.ru's mid-2026 frontier tally already adds Grok 4.5, Claude Sonnet 5, DeepSeek V4 Pro, and Kimi K3. When a name changes, swap the name; the job description stays fixed.

Rule 4 — Ring-fence the human. Allocate at least 60% of the editing budget to one human pass before buying any AI pass, and fund AI passes only from the remainder. One honesty note: no fetched 2026 source publishes current human developmental or copyedit rates, so anchor the copyedit line item to a written quote from the editor you actually hire, not to any figure printed in this guide.

```

Frequently Asked Questions

How many API tokens will my 90,000-word novel burn on input alone?

English fiction converts at roughly 1.33 tokens per word under byte-pair encoding, so a 90,000-word manuscript runs about 120,000 input tokens before the model emits a word back.

If I'm forced to chunk my manuscript, does Anthropic prompt caching soften the repeated input charges?

Per Anthropic's published caching terms, cached input reads bill at 10% of list price, cutting the six-chunk re-read penalty by roughly an order of magnitude.

What does a human developmental edit of a 360-page novel actually cost?

At the Editorial Freelancers Association's 2024 pace of seven 250-word pages per hour, a 360-page novel implies roughly 51 billed hours at $46–$60 each—up to $3,080—while Reedsy's live-quote calculator reaches $7,200 for top-of-market developmental work.

Why does a machine line-edit cost several times more than an editorial letter on the same model?

A letter-style structural pass returns about 9,000 output tokens, but a line-edit that returns fully rewritten prose bills roughly 126,000 output tokens—the whole manuscript again—and output tokens bill above input on every major rate card.

Since AI passes are cheap, why not run twenty of them on my draft?

Ten passes already emit roughly 70,000 words of notes demanding 35+ adjudication hours—more author-labor than the human editor's entire engagement—and note quality plateaus by pass three.

Is GPT-5.5's premium price justified by better benchmark scores than Gemini 3.1 Pro or Claude Opus 4.7?

Despite billing $5 per million input and $30 per million output tokens versus Gemini 3.1 Pro's $2/$12, GPT-5.5's GPQA Diamond ceiling of 94.4% sits only tenths of a point above Gemini's 94.3% and Claude Opus 4.7's 94.2%.

Quick answers

Which 2026 flagship model sets the cheapest per-pass price floor for reading a full 90,000-word novel?Gemini 3.1 Pro, at $2 per million input tokens and $12 per million output tokens, offers the cheapest full-manuscript read on the market.
How do the GPQA Diamond benchmark scores compare across Gemini 3.1 Pro, Claude Opus 4.7, and GPT-5.5?They are nearly identical: Gemini 3.1 Pro scores 94.3%, Claude Opus 4.7 scores 94.2%, and GPT-5.5 tops out at 94.4%.
What is the optimal 2026 AI budget per book according to the article?About $5 of AI per book is enough for the diagnostic passes worth running, while every dollar above it purchases negative-value compute.
How many input tokens does a 90,000-word novel convert to before the model emits a single word back?Roughly 120,000 input tokens, since English fiction converts at approximately 1.33 tokens per word.
What does Anthropic's prompt caching offer as an escape hatch from chunking re-read costs?Cached input reads bill at 10% of list price, cutting the re-read penalty by roughly an order of magnitude.

Also worth reading: How to create a more productive and balanced daily routine for long term success: How to create a more · LLM Plot Structure: Variance, Token Collapse, and Hidden Data in AI Drafting: LLM Plot Structure: Variance, Token · AI Novel Consistency: 23% Contradiction Rate vs. Token Cost: AI Novel Consistency: 23% Contradiction

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Storywriter editorial desk (About, Contact, Privacy).

Related answers