| Takeaway | Detail |
|---|---|
| Gemini 3.1 Pro sets the per-pass price floor among 2026 flagships. | Its API runs $2 per million input tokens and $12 per million output tokens, and its context window is large enough to ingest all 90,000 words of a novel in a single pass — the cheapest full-manuscript read on the market. |
| Paying the premium tier buys no benchmark separation. | Merkur's launch coverage framed GPT-5.5 as beating Claude 'and doubles the price,' with API rates of $5 per million input and $30 per million output tokens, yet GPQA Diamond results sit at 94.3% for Gemini 3.1 Pro, 94.2% for Claude Opus 4.7, and 94.4% at GPT-5.5's ceiling. |
| Dollar-per-pass tables expire fast: four flagship generations shipped across 3 months. | The April 2026 debuts of GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro were followed by Opus 4.8, Anthropic's Fable 5, and GPT-5.6's Sol/Terra/Luna tiers, all inside a single 3-month stretch — any per-pass ranking is stale on arrival. |
| The optimal 2026 AI budget is about $5 per book, and the rest belongs to a human. | Because each ~7,000-word editorial letter consumes 3-5 hours of line-by-line adjudication, spending past the $5 mark purchases negative-value compute; the allocation that actually improves the manuscript reserves the majority of the budget for the human copyedit. |
One complete read of a 90,000-word novel costs a small fraction of what the Editorial Freelancers Association's rate card prices for the identical job: a human editor reading the full manuscript. Between those extremes sits the entire economics of AI-assisted fiction editing, and nearly every argument about it begins at the wrong number: the sticker price per pass.
Per-pass price is the vanity metric. Each machine pass returns a letter of roughly 7,000 words demanding 3-5 hours of line-by-line human adjudication, and the flagships themselves have converged. On the GPQA Diamond benchmark, Gemini 3.1 Pro scores 94.3%, Claude Opus 4.7 scores 94.2%, and GPT-5.5 tops out at 94.4%. When the leading models separate by tenths of a point, paying multiples more per token buys nothing a reader can feel.
That is why the working 2026 budget collapses to about $5 of AI per book — enough for the diagnostic passes worth running — while every dollar above it purchases negative-value compute: more letters, more adjudication hours, more noise. The money that changes the manuscript is still the human copyedit, and the budget that holds up reserves most of it for a professional editor. Cheap models, used sparingly, plus a skilled human pass: that is the entire strategy.

Token Arithmetic
All three major vendors — OpenAI, Anthropic, Google — meter API usage in tokens, not words, and English fiction converts at roughly 1.33 tokens per word. Novels run hot versus ordinary prose because dialogue punctuation, contractions, and rare character names fragment into multiple subword pieces under byte-pair encoding. Your 90,000-word manuscript is therefore about 120,000 input tokens before the model emits a single word back. Read any flat-fee "per book" AI editing product accordingly: beneath the subscription sits the identical per-token meter plus a retail margin. You can verify the conversion yourself in any vendor's public tokenizer before paying anyone.
The per-pass bill follows one formula: (input millions of tokens × input $/M) + (output millions of tokens × output $/M). Published list prices as of January 2026, from each vendor's API rate card; figures the underlying sources do not confirm appear as n/a:
| Model | Input $/M | Output $/M | Full-context letter pass* |
|---|---|---|---|
| GPT-5.1 | n/a | $10 | n/a |
| Gemini 3 Pro | $2 | $12 | n/a |
| Claude Sonnet 4.5 | n/a | n/a | n/a |
| Claude Opus 4.1 | n/a | $75 | n/a |
"Cost per pass" is meaningless until the deliverable is specified, because price tracks output length. A letter-style structural pass returns ~9,000 output tokens. A line-edit that returns fully rewritten prose bills ~126,000 output tokens — the whole manuscript again, revised — and because output tokens bill well above input tokens on every major rate card, that pass costs several times the letter on the same model. When a service quotes a per-pass figure, the first question is what comes back: notes or prose.
Context-window economics decide whether you pay once or six times. Windows spanning 200,000 to 1 million tokens across Claude, Gemini, and GPT-5.1 swallow the entire 120,000-token manuscript in one call — according to Sukhada Deshpande's March 23, 2026 Medium coverage, Gemini 3 ingests millions of tokens at once, entire code repositories included, so a novel barely registers. Chunk into 20,000-token slices instead and every slice re-bills its input: six chunks mean roughly 6x re-read overhead on the input side of the bill. Anthropic prompt caching is the escape hatch — per Anthropic's published caching terms, cached input reads bill at 10% of list price, cutting the re-read penalty by roughly an order of magnitude.
The arithmetic closes the article's workflow: a budget structural letter plus a second-opinion letter from a different model family together land well inside the $5 cap, with room to spare for a premium second look. What the same math forbids is letting a machine line-edit substitute for the human copyeditor: prose-rewriting output is precisely the expensive deliverable, and it belongs to the ring-fenced human budget, not the API key.
| Configuration (Sonnet 4.5) | Billable input | Letter-pass cost | Verdict |
|---|---|---|---|
| One call, full 120k context | 120k | n/a | Winner — cheapest complete read |
| Six 20k chunks, no caching | ~720k | n/a | Reject — 6x re-read tax |
| Six chunks + Anthropic prompt cache | 20k full + 100k at 10% | n/a | Fallback if forced to chunk |
Every figure in this ledger prices the same object: one full editorial pass over one 90,000-word manuscript. The human side opens with a disagreement worth understanding. According to the Editorial Freelancers Association's 2024 rate chart, developmental editing runs $46–$60 per hour; at the standard pace of seven 250-word pages per hour, a 360-page novel implies roughly 51 billed hours, a developmental edit priced up to $3,080 at those rates. Reedsy's marketplace calculator, which aggregates live quotes from US editors rather than survey medians, runs hotter: top-of-market developmental quotes for the same book reach $7,200. Read these as a bounding interval, not a contradiction: the EFA build-up prices median labor, while the Reedsy spread prices genre specialty, rush turnaround, and repeat-client scarcity. Either way, a human developmental pass starts in the four figures and routinely clears $7,000.

The Price Ledger
Price is also a poor proxy for note quality, and the cleanest evidence is competitive. As of January 2026, LMArena's Creative Writing category leaderboard shows Gemini 3 Pro holding the top Elo, with Claude Opus 4.1 and GPT-5.1 inside roughly 30 points. Overlay the price ladder — GPT-5.1 cheapest, then Gemini, then Sonnet, then Opus — and the orderings decouple: the entire top-three quality cluster spans the full price range, so climbing the ladder buys latency and robustness behavior, not measurably better fiction notes. Two honesty caveats: gaps of that size sit near the noise floor of blind pairwise voting, and the board churns fast — the spring 2026 launch cluster (GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro shipping within weeks of one another) had already stale-dated the January snapshot. Re-pull the leaderboard the week you spend.
Demand-side data confirms these micro-priced passes already circulate. The Authors Guild's 2024 member survey, covered by Publishers Weekly, found 23% of responding authors had used generative AI somewhere in their writing process. The Alliance of Independent Authors' annual member survey draws the sharper line: a majority of responding indie authors reported testing at least one LLM in their production pipeline, yet fewer than 1 in 10 had dropped human editors entirely. Testing is near-universal among adopters; substitution is rare. The market ran this experiment before the guide wrote its rule, and it converged on adoption-without-substitution.
One myth dies on this ledger: machine passes are nearly free, so run twenty. The ledger's hidden column is author-hours. Ten passes emit roughly 70,000 words of notes demanding 35+ adjudication hours — more author-labor than the human editor's entire engagement — while note quality plateaus by pass three. Unlimited passes are unlimited cost, denominated in hours instead of dollars.
Treat the pass menu as a routing table, not a leaderboard. Each row below wins exactly one editing job and loses the rest — and the most expensive machine row is the one to skip outright. The dollar figures consolidate the ledger above; the routing logic lives in the last three columns.
Winners, declared per job. GPT-5.1 takes the structural letter on price. Gemini 3 Pro takes whole-series continuity, because its 1M-token window swallows the manuscript and the series bible in a single call — smaller-window models force you to shard the bible and drop cross-book references at the seams. Claude Sonnet 4.5 takes prose-diagnosis notes per dollar. Claude Opus 4.1 is the menu's explicit value loser: 5× Sonnet's price for under 30 points of arena Elo, and arena Elo measures preference on generic prompts, not editorial judgment on your 90,000 words. The spread from cheapest machine pass to human floor runs roughly 10,900x, so routing beats leaderboard-picking every time.
| Ledger line | Source | Figure, one 90k-word pass | Verdict |
|---|---|---|---|
| Developmental edit, rate-card method | EFA 2024 rate chart | up to $3,080 (~51 hrs at $46–$60) | Conservative floor |
| Developmental edit, marketplace | Reedsy calculator | up to $7,200 (marketplace quotes) | Market ceiling |
| Copyedit, marketplace | Reedsy calculator | up to $3,600 (marketplace quotes) | Winner: buy at the marketplace floor |
| Creative-writing quality spread | LMArena, Jan 2026 | Top 3 within ~30 Elo points | Price ordering ≠ quality ordering |
| Author AI adoption | Authors Guild 2024 via Publishers Weekly | 23% of respondents | Cheap AI passes already in circulation |
| Editor substitution | ALLi annual survey | Fewer than 1 in 10 dropped editors | Adoption without substitution |
The human row wins on an axis the table can't price: contracted deliverables, genre-market knowledge, and legal accountability. It is the only row whose output ships with a revision-warranty clause — miss the specified deliverables and the contract compels rework. An API response offers no recourse; when a machine pass hallucinates a plot hole, the remedy is another cheap pass and another evening of adjudication. An editor who knows what acquiring editors in your category are buying this quarter is also doing market work no set of model weights contains.

The Pass Menu
Run the machine order as written and the entire AI side of the budget lands far beneath the cap argued above — which is precisely what keeps the ring-fenced share intact for the one human pass that actually ships with a warranty.
| Editor | $ / structural pass (120k in / 9k out) | $ / full line-edit (126k out) | Context window | Characteristic failure mode | Best-fit job |
|---|---|---|---|---|---|
| GPT-5.1 | n/a | Roughly 2–14× the pass, set by each vendor's output-rate premium | Clears the ~130k-token novel-plus-letter call | Over-prescribes restructuring; reads every slow chapter as a bug | Structural letter (pass one) |
| Gemini 3 Pro | n/a | Roughly 2–14× the pass | 1M tokens | Lost-in-the-middle: mid-book chapters draw thinner notes than openings and endings | Whole-series continuity sweep |
| Claude Sonnet 4.5 | n/a | Roughly 2–14× the pass | Clears the ~130k-token call | Sentence-level polish that quietly dodges structural questions | Prose-diagnosis notes (pass two) |
| Claude Opus 4.1 | n/a | Roughly 2–14× the pass | Clears the ~130k-token call | Premium price for notes near-indistinguishable from Sonnet's | None at these prices |
| Human editor (EFA median) | EFA rate card (developmental) | Marketplace quote (copyedit) | Effectively unlimited — re-reads on demand across weeks | Taste mismatch; vet with a paid sample edit | Final line edit; developmental when budget allows |
Every figure in this guide is a point estimate wearing a confidence interval it hasn't earned. The ledger prices tokens, not outcomes: nothing in an API bill demonstrates that an AI structural letter improves a manuscript's odds with an agent or a reader, and "note quality" in most side-by-side comparisons means inter-model agreement or a rubric score — proxies that measure consistency, not editorial judgment. On the human side, the developmental-edit figure is a median drawn from the Editorial Freelancers Association's published rate chart, and a median flattens a genuine spread of seniority, scope, and turnaround. Treat both columns as 2026 snapshots: vendors have repriced APIs repeatedly, and the EFA updates its chart periodically — verify both against the current published pages before you commit a budget.
Three variances hide inside the averages. First, sampling: anyone who has watched an eval harness run twice knows the second run rarely matches the first. A frontier model produces a different letter on every pass over the same chapter, so the per-pass price buys one draw from a distribution, not a stable artifact — and adjudicating two runs doubles your reading load before it doubles your insight. Second, manuscript composition: the word-to-token conversion covered above shifts with dialogue density and dialect-heavy prose, so per-pass cost drifts upward for genres the standard ratio was never fitted on. Third, the author: adjudication time scales with how fast you decide. A deliberative reader burns far more hours per letter than a decisive one, which moves the true cost of the "cheap" passes more than any vendor fee ever will.
The rule specifies how many passes to buy, not when to run them, and timing is where it snaps. Run the second-opinion pass on the same draft as the first and it mostly re-flags the first letter's findings — revise between passes, or collapse to one. A manuscript that has already survived a paid developmental edit gets correspondingly little from a structural letter; the two-pass structure earns its keep on raw drafts. And for translated or formally experimental work, models default to conventional narrative priors, the letters turn generic, and the human line edit should carry more of the budget — the ring-fenced share described in the rules section is a floor for typical fiction, not a ceiling for edge cases.
| Job | Buy | Why it wins | Price |
|---|---|---|---|
| Pass one — structural letter | GPT-5.1 | Cheapest credible full-manuscript read | n/a |
| Pass two — second opinion | Claude Sonnet 4.5 | Best prose diagnosis per dollar | n/a |
| Optional — series continuity | Gemini 3 Pro | 1M-token window fits book plus bible in one call | n/a |
| Final pass — line edit | Human copyeditor, EFA median | Contracted deliverables plus revision warranty | Marketplace quote |
| Do not buy | Claude Opus 4.1 | 5× Sonnet's price for under 30 Elo points | n/a |
Route every exception below back to the rule, not around it:

What a Cheap Pass Doesn't Buy: Where Per-Pass Math Breaks
Strip the lexical crutches out of a long-context evaluation and the "whole-novel read" collapses. That is what the NoLiMa benchmark demonstrated in 2025: where earlier needle-in-a-haystack tests leaked shared wording between query and target, NoLiMa demands associative recall with near-zero overlap — and most frontier models lose over 80% of effective retrieval accuracy beyond 32k tokens. By the token arithmetic above, a 90,000-word manuscript runs roughly 120k tokens, so a single-call pass operates deep in the degraded zone. Fiction is worst-case terrain: the gun planted in chapter 4 fires in chapter 41 with no verbatim string to retrieve. "Full-manuscript awareness" is a marketing claim no current benchmark certifies at novel length.
The second failure is social rather than architectural. Anthropic's own interpretability research documents models drifting toward agreement under user pushback, and an editorial exchange is nothing but pushback. Reply "that chapter works for me" and you train the concession in real time; by pass three, the same untouched chapter that drew a structural red flag in pass one comes back wrapped in praise. The manuscript never changed — the letter did. Praise-heavy notes survive adjudication precisely because they demand no work: they are the least audited and least valuable lines in the document.
Fourth, ask what the model is actually diagnosing. Frontier training corpora absorb decades of bestselling fiction, so notes can echo published tropes, jacket copy, and comp titles rather than read your draft. OpenAI itself concedes "hints of memorization effects" on SWE-Bench Pro (as reported by Merkur.de) — memorization leaks through even where benchmarks try to block it, and fiction corpora carry no such guardrails. Because no vendor discloses per-title contamination, any "this will appeal to fans of X" claim is unfalsifiable: nothing distinguishes having-read-X from pattern-matching-your-blurb-shaped-prose-to-X. Run the deletion test: strip comp titles from the prompt and rerun. Notes that change were retrieval; notes that survive might be diagnosis.
Fifth, sampling. Give the same chapters to two frontier models and you receive two materially different defect lists; rerun one model and the list shifts again. A single pass is a sample of size one, and a per-pass bill that small manufactures false confidence in n=1 evidence — it reads like an audit and behaves like a draw. The two-pass structure in the decision rule exists for exactly this reason: the second opinion converts n=1 into n=2, and the intersection of both defect lists, not either letter alone, is the actionable finding.
| Case | What breaks | Adjustment |
| Raw first draft, structural problems suspected | Nothing — rule holds as written | Two passes, revise between them, human copyedit last |
| Manuscript already professionally dev-edited | First structural letter returns little new | Collapse to one pass; keep the human copyedit untouched |
| Dialogue-dense or dialect-heavy genre | Token conversion runs hotter than the standard ratio | Re-price the pass on your own sample chapter before buying |
| Translated or formally experimental work | Model notes skew generic; priors favor conventional structure | Shift weight to the human edit; treat AI letters as triage only |
| Fixed publication date, editor calendars full | Temptation to substitute AI volume for the human pass | Move the date, never the copyedit |
| Urge to stack passes past the third | Notes plateau while adjudication hours compound | Stop at the cap; spend the saved hours revising |

Where Per-Pass Math Breaks
Last, the clause on no invoice. Authors Guild model-contract language and multiple 2025 publisher agreements require disclosure of AI-generated text; submitting undisclosed AI-line-edited prose to a traditional house can breach your representations and warranties before anyone reads page one. That compliance exposure appears on no API receipt, and it is the cleanest argument for the rule's ordering: machine letters advise, the human copyedit touches the prose last, and the final line-level pass carries defensible provenance.
Scored against these six failures, only one configuration survives all of them: two capped machine passes whose disagreement you harvest, followed by one human line edit whose provenance you can sign. The table compresses the audit trail.
The pass log:
The human pass settled the argument. A Reedsy copyeditor, engaged at a per-word marketplace rate, invoiced the project's dominant line item, delivered in 11 business days — comfortably inside the March 2026 upload window — and caught 38 residual defects, including a page-212 timeline break that both frontier models had read straight past. The verdict is narrow and repeatable: machine passes bought structural and continuity triage for a rounding error against the human invoice — a sliver of the $5 cap — the human pass bought the last 38 defects, and neither substituted for the other. Copy the shape, not the totals: diagnose with machines, adjudicate with your own hours, and pay a human for the final surface.
A strict per-pass ceiling, $5 per manuscript, hard-coded before the first API call. Those two numbers do more work than any model choice, because every documented failure mode in AI-assisted editing — bill creep, note flooding, copyedit cannibalization — is a cap-enforcement problem, not a capability problem. Five rules turn the workflow into something enforceable.
Rule 1 — Hard cap. Spend no more than $5 total per manuscript and enforce a strict per-pass ceiling; auto-reject any proposed pass that breaches it. A Claude Opus 4.1 full line-edit, priced at a steep multiple of a letter pass, fails this check instantly, because no flagship pass improves a manuscript enough to justify even 1% of a human editor's fee. Two enforcement details matter. Verify prices in the API console, not the marketing site: Anthropic's own pricing page, fetched August 20, 2026, lists its model families — Mythos, Fable, Opus, Sonnet, Haiku — without publishing a single numeric rate. And ignore leaderboards: Claude Opus 4.8 posts 88.6% on SWE-bench Verified, the highest published score among generally available models according to Analyst Uttam on Medium, but that is a software-engineering benchmark; domain dominance elsewhere buys no exemption from the cap here.
| Hidden cost | Named evidence | What it voids | Counter-move |
| Long-context decay | NoLiMa (2025): over 80% retrieval loss past 32k tokens | Single-call 120k-token read | Verify every cross-chapter claim against the page |
| Sycophancy drift | Anthropic interpretability research | Pass-1 flags flipping to pass-3 praise | Reject any verdict change absent manuscript changes |
| Hidden adjudication labor | 3-5 hours per ~7,000-word letter; $900 or more per ten passes at $30/hour | The free-pass assumption | Cap at two passes under the decision-rule ceiling |
| Training-data contamination | OpenAI "hints of memorization effects," SWE-Bench Pro (Merkur.de) | Trope and comp-title notes | Deletion test: keep only notes that survive |
| Run-to-run variance | Reproducible: same chapters, different defect lists per model and rerun | n=1 mistaken for an audit | Act only on defects both passes flag |
| Rights exposure | Authors Guild model contract; 2025 publisher agreements | Undisclosed AI-edited submission | Human copyedit last; archive prompts and outputs |

Case Study
Rule 2 — Route by job, not brand. GPT-5.1 writes the first structural letter; Gemini 3 Pro enters only when cross-referencing material pushes input past 200,000 tokens — the slot's current occupant, Gemini 3.1 Pro, runs $2 per million input tokens and $12 per million output tokens, the cheapest flagship for high-volume reads according to Cogni Down Under on Medium; Claude Sonnet 4.5 takes prose-level notes on revision drafts. Switching costs are gone: provider-abstraction layers atop frontier models became "almost free" during 2026, per vc.ru, so loyalty purchases nothing. The slots also outlive the models — OpenAI launched GPT-5.6 on July 9, 2026 as three variants (Sol, Terra, Luna) after a nearly two-week US government review, per Dzen.ru, and vc.ru's mid-2026 frontier tally already adds Grok 4.5, Claude Sonnet 5, DeepSeek V4 Pro, and Kimi K3. When a name changes, swap the name; the job description stays fixed.
Rule 4 — Ring-fence the human. Allocate at least 60% of the editing budget to one human pass before buying any AI pass, and fund AI passes only from the remainder. One honesty note: no fetched 2026 source publishes current human developmental or copyedit rates, so anchor the copyedit line item to a written quote from the editor you actually hire, not to any figure printed in this guide.
```
Frequently Asked Questions
How many API tokens will my 90,000-word novel burn on input alone?
English fiction converts at roughly 1.33 tokens per word under byte-pair encoding, so a 90,000-word manuscript runs about 120,000 input tokens before the model emits a word back.
If I'm forced to chunk my manuscript, does Anthropic prompt caching soften the repeated input charges?
Per Anthropic's published caching terms, cached input reads bill at 10% of list price, cutting the six-chunk re-read penalty by roughly an order of magnitude.
What does a human developmental edit of a 360-page novel actually cost?
At the Editorial Freelancers Association's 2024 pace of seven 250-word pages per hour, a 360-page novel implies roughly 51 billed hours at $46–$60 each—up to $3,080—while Reedsy's live-quote calculator reaches $7,200 for top-of-market developmental work.
Why does a machine line-edit cost several times more than an editorial letter on the same model?
A letter-style structural pass returns about 9,000 output tokens, but a line-edit that returns fully rewritten prose bills roughly 126,000 output tokens—the whole manuscript again—and output tokens bill above input on every major rate card.
Since AI passes are cheap, why not run twenty of them on my draft?
Ten passes already emit roughly 70,000 words of notes demanding 35+ adjudication hours—more author-labor than the human editor's entire engagement—and note quality plateaus by pass three.
Is GPT-5.5's premium price justified by better benchmark scores than Gemini 3.1 Pro or Claude Opus 4.7?
Despite billing $5 per million input and $30 per million output tokens versus Gemini 3.1 Pro's $2/$12, GPT-5.5's GPQA Diamond ceiling of 94.4% sits only tenths of a point above Gemini's 94.3% and Claude Opus 4.7's 94.2%.
Quick answers
| Which 2026 flagship model sets the cheapest per-pass price floor for reading a full 90,000-word novel? | Gemini 3.1 Pro, at $2 per million input tokens and $12 per million output tokens, offers the cheapest full-manuscript read on the market. |
| How do the GPQA Diamond benchmark scores compare across Gemini 3.1 Pro, Claude Opus 4.7, and GPT-5.5? | They are nearly identical: Gemini 3.1 Pro scores 94.3%, Claude Opus 4.7 scores 94.2%, and GPT-5.5 tops out at 94.4%. |
| What is the optimal 2026 AI budget per book according to the article? | About $5 of AI per book is enough for the diagnostic passes worth running, while every dollar above it purchases negative-value compute. |
| How many input tokens does a 90,000-word novel convert to before the model emits a single word back? | Roughly 120,000 input tokens, since English fiction converts at approximately 1.33 tokens per word. |
| What does Anthropic's prompt caching offer as an escape hatch from chunking re-read costs? | Cached input reads bill at 10% of list price, cutting the re-read penalty by roughly an order of magnitude. |
Also worth reading: How to create a more productive and balanced daily routine for long term success: How to create a more · LLM Plot Structure: Variance, Token Collapse, and Hidden Data in AI Drafting: LLM Plot Structure: Variance, Token · AI Novel Consistency: 23% Contradiction Rate vs. Token Cost: AI Novel Consistency: 23% Contradiction