Voice Research · 12 September 2026
AI writing worth reading.
Research on AI writing, editing rules and review quality.
What the evidence changes
Review the reader's problem
Relevance, information density and tone helped explain expert judgments in Measuring AI ‘Slop’. Agreement was stronger on problematic passages than on a binary label. Shaib et al., preprint, revised 2026.
For Voice, the review should identify passages that prevent the reader from understanding, deciding or acting. Word matches can locate candidates; the page's purpose determines whether they need changing.
Better individual texts can still become more alike
In a randomized short-story experiment, AI ideas improved ratings while assisted stories became more similar. This is evidence for assessing usefulness and distinctiveness separately. Doshi and Hauser, 2024.
In a contextual review, compare repeated framing across a page and across proposed alternatives. Three rewrites that merely swap synonyms offer little choice. A deliberate rhythm or an accepted headline can remain effective.
A familiar word is a clue with a narrow scope
Corpus research detects abrupt vocabulary shifts; it cannot identify the origin of an individual passage. Kobak et al., 2025. An exploratory study did not find an overall human preference for the conspicuous words it tested. Juzek and Ward, 2025.
A rule therefore needs a condition, a reader cost and an exception. Banning a word everywhere would erase useful terminology along with empty phrasing.
Revision needs an explicit success criterion
Self-Refine improved its tested tasks, but a reasoning study found that unsupported self-correction could introduce errors. The tasks and feedback conditions differ. Madaan et al., 2023; Huang et al., ICLR 2024.
Our implication is to require a quoted problem and compare the revision with the source. Preserve evidence, qualifications and accepted decisions. Shorter output, a model's approval and a high acceptance rate each leave important questions unanswered.
Quality and authorship need different evidence
Historical detector studies found sensitivity to language background, generator changes and mixed editing. Their reported rates apply to the tested conditions. Liang et al., 2023; RAID, 2024; MixSet, 2024.
Voice returns local writing signals and review rules without an AI-authorship probability. A recorded generation run can document a known contribution; a Git commit by itself does not establish who composed its sentences.
Practice articles and the “vibe” debate
These firsthand essays and directories put names to frustrating writing habits. Read them as examples and editorial arguments. Their model labels and proposed remedies rarely come with controlled comparisons.
Rules from these observations
Unrequested interpretation informs meta-framing, reader-goal. Formulaic contrasts and repeated rhetorical structures inform unsupported-contrast, forced-template, repetition. Review the wording against the reader’s task before proposing a change.
Implementation: The catalogue now includes conditional heading guidance derived from Furze’s examples and the editing research: identify the section’s subject, preserve useful questions and imagery, and avoid implying results the section does not contain. Full-page rhetorical repetition still requires a host-run contextual review.
See the Library’s exact analysis coverageProblem Patterns in AI: Beyond HallucinationsAsk what each heading or explanation contributes to the reader's task. Compare emphasis with the supplied source and remove unsupported interpretation.
- Observation
- Furze describes Claude adding interpretive subtitles to requested slide headings. His proposed categories cover faithfulness, judgment, communication and aesthetics alongside truth: an accurate addition can still distort emphasis or waste attention.
- Limits
- A practitioner's taxonomy and examples, with an explicitly informal poll. It does not measure model failure rates or establish the proposed technical causes.
Use with Voice
A host can supply the brief alongside generated headings for an agent to review. The library prepares instructions; it does not compare their emphasis.
composeWritingReviewPrompt— Frame heading review around the supplied audience and goal.selectWritingRules— Select reader-goal, task-relevance and meta-framing instructions.
What is missing: A context builder connecting page purpose, source statements and heading roles, followed by a review result with verifiable references.
Related patterns: Unnecessary interpretive framing · Taxonomy before the task · Optional-detail detours · Invented cause and benefit · Forced triads and canned Q&A
Why ChatGPT writes like that: A rhetorical analysis of AI ‘slop’Review recurrence and meaning across the page. Preserve a deliberate, accepted headline; suggest alternatives when a rhetorical structure invents an opposition or substitutes rhythm for substance.
- Observation
- Gorrie analyzes parallelism, antithesis and three-part constructions, including a disclosed Claude Sonnet 4 letter prompt. These devices have a human rhetorical history; concentration and predictable repetition can make them feel synthetic.
- Limits
- Illustrative rhetorical analysis, not a sampled corpus or controlled comparison between providers. An effective contrast or three-part headline is not automatically a defect.
Use with Voice
Current matches can start a rhetorical review, but they do not recognize every antithesis, parallel structure or deliberate three-part headline.
findWritingSignals— Locate supported unsupported-contrast phrase patterns.getWritingRule— Retrieve contrast exceptions and conditional examples.
What is missing: Compare rhetorical choices against the page goal and recurring structures across Git files, preserving effective contrasts while offering substantively different alternatives.
Related patterns: Not X, but Y · Forced triads and canned Q&A · Summary loops · Formatting on every sentence
On AI-assisted writing, AI slop, and the line between themEvaluate whether assistance preserves the intended claim. Keep factual qualifications and author choices visible when comparing revisions, including decisions to retain the original.
- Observation
- Amatriain describes AI-assisted editing as useful for language and structure while warning that revisions can shift emphasis, caveats and claims. He discloses assistance in the essay itself and places responsibility for its ideas with the author.
- Limits
- A first-person position, including anecdotal detector experience. It does not establish a general productivity gain, an authorship threshold or detector accuracy.
Use with Voice
The prompt expresses preservation requirements; the overlay lets a reviewer inspect quoted passages. Neither API proves that a suggested rewrite preserves emphasis.
composeWritingReviewPrompt— Request alternatives that preserve facts, qualifications and accepted decisions.mountVoiceReview— Display host-supplied findings beside the original passage.
What is missing: Compare claim changes between alternatives and protect recorded author decisions against the exact source revision.
Related patterns: Taxonomy before the task · Claim without evidence · Hedge stacking · Ornamental wording
The ‘Negative Contrast Trap’: Why AI Writing Overuses ‘Not X, But Y’Flag the quoted contrast, establish whether its rejected alternative is relevant, then offer a direct statement. Retain contrasts that resolve an actual misconception.
- Observation
- Hughes demonstrates requesting an affirmative rewrite of a negative contrast and suggests checking repeated sentence openings after generation. The article provides an accessible example of converting a stylistic complaint into an editing instruction.
- Limits
- The examples and illustrative heuristics are not an evaluated detector. Claims about frequency and training mechanisms are hypotheses here, rather than measured results.
Use with Voice
The matcher covers selected prefixes, not every “not X, but Y” construction. A reviewer must establish whether the rejected alternative matters.
findWritingSignals— Find the catalogue's supported English and German contrast openings.composeWritingReviewPrompt— Ask for grounded alternatives under unsupported-contrast.
What is missing: Construct the relevant misconception context and compare an affirmative rewrite with the original without discarding a useful distinction.
Related patterns: Not X, but Y · Summary loops · Arrow-chain shorthand
Tropes: AI Writing Pattern DirectoryUse such observations to generate conditional review hypotheses. Supply counterexamples and reader costs instead of importing a universal banned-word or punctuation list.
- Observation
- This practitioner directory names recurring word, sentence, tone and composition patterns and distributes prompt instructions. Its examples include negative parallelism, announcement-like conclusions and diluting one point through repeated restatement.
- Limits
- A strongly opinionated directory, not a validated authorship benchmark. Its prompt includes categorical bans that can conflict with useful explanation or enumeration. The site's previewed tools do not establish efficacy.
Use with Voice
Voice exposes its own curated rules; it neither imports this directory nor executes its prompt files. Composition-level judgments still require a reviewer.
getWritingRule— Inspect a conditional rule with legitimate keep examples.selectWritingRules— Choose relevant catalogue rules instead of a universal ban list.
What is missing: Benchmark selected instructions against both unwanted repetition and legitimate enumeration, recording false positives and preservation decisions.
Related patterns: Not X, but Y · Unnecessary interpretive framing · Summary loops · Unprioritized laundry list · Forced triads and canned Q&A
Wikipedia: Signs of AI writingReview claims and cited evidence alongside wording. Keep genre-specific rules separate, and distinguish malformed citation artifacts from ordinary punctuation or vocabulary.
- Observation
- Editors collect examples of inflated significance, superficial analysis, vague attribution, promotional language and leaked assistant markup. The advice page explicitly treats its list as observations and warns that visible signs can conceal deeper sourcing problems.
- Limits
- Community advice, not Wikipedia policy or a detector benchmark; some examples depend on encyclopedic conventions. Human texts can share these features, and the page flags outdated coverage of recent models.
Use with Voice
An agent can receive cited material separately. Current signals neither resolve references nor identify all malformed citation artifacts; encyclopedic conventions need host context.
composeWritingReviewPrompt— Select claims-evidence and reported-evidence for a sourcing review.findWritingSignals— Locate the catalogue's limited claim and verification phrases.
What is missing: Connect each claim to retrievable evidence and validate source locators and citation artifacts within the selected document genre.
Related patterns: Claim without evidence · Invented cause and benefit · Ornamental wording · Unnecessary interpretive framing · Unsupported completion claim
slopt: Stop sounding like ChatGPTEvaluate counter-prompts on cases where the original already serves the reader task. A larger instruction catalogue alone does not demonstrate better writing, and English claims do not qualify German behavior.
- Observation
- A commercial catalogue packages anti-pattern instructions as prompt or skill files. Its public examples use punctuation and rhetorical prohibitions; the FAQ acknowledges that prompts cannot guarantee compliance.
- Limits
- Marketing claims about pattern counts, crawls and model coverage were not independently verified. No auditable evaluation supporting the advertised effect was established from the page; no paid material or product was tested.
Use with Voice
Consumers can inspect the generated prompt before using their own agent. Voice does not include or execute this vendor's paid material or promise prompt compliance.
composeWritingReviewPrompt— Build local, selected-rule instructions with conditional examples.selectWritingRules— Keep the instruction set specific to the writing task.
What is missing: A reproducible English/German prompt comparison using preserved originals, warranted revisions and legitimate exceptions to stylistic prohibitions.
Related patterns: Forced triads and canned Q&A · Formatting on every sentence · Unsupported completion claim · Ornamental wording
Studies: methods, findings and limits
The collection covers text quality, lexical drift, homogenization, agreeable feedback, online accusations and authorship detection. Each record separates the measured setting from its implication for Voice and names the API functions available today. Preprints are identified in their limits.
Analyses supported by the studies
Editing studies inform word-choice, format-fit, task-relevance: inspect awkward phrasing, decorative wording and missing useful detail. Studies of claims and agreement inform claims-evidence, calibrated-uncertainty: compare the original and proposed claims against the supplied evidence.
Implementation: These are contextual prompt instructions. The package does not implement the papers’ statistical models, semantic fidelity checks or calibrated judges. Lexical drift and authorship studies constrain interpretation; they do not turn a word match into an authorship finding.
See the Library’s exact analysis coverageCan AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsWhich edits address ornamental, imprecise or redundant AI-generated writing?
- Method
- Eight writers informed seven edit categories; eighteen writers edited 1,057 creative-writing paragraphs from GPT-4o, Claude 3.5 Sonnet and Llama 3.1-70B.
- Findings
The taxonomy separates distracting ornament, insufficient specificity, awkward wording and redundant exposition.
Adding relevant detail can improve writing even when it increases length.
- Limits
Creative genres and subjective expert edits do not establish technical-heading failure rates. Paragraph-level evaluation can miss whole-document problems; factual hallucinations were not studied.
- Implication for Voice
- Review what wording contributes to the reader's task; retain useful metaphor, rhythm and specificity.
Use with Voice
The composer prepares instructions. It does not detect these semantic categories or implement the study's editing model.
selectWritingRules— Select reader-goal, word-choice and repetition guidance.composeWritingReviewPrompt— Supply audience and purpose for contextual review.
What is missing: Compare quoted problems and alternatives against page context, including deliberate keep cases and expert disagreement.
Related patterns: Taxonomy before the task · Ornamental wording · Optional-detail detours · Summary loops
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference ModelsDo model judges reward surface features beyond human preferences?
- Method
- Controlled English counterfactual pairs vary five features; four reward models and three LLM judges are compared with three human judgments per case.
- Findings
Reward models overvalued jargon and vagueness relative to human judgments; LLM judges also showed preference distortions.
Counterfactual training reduced measured miscalibration.
- Limits
Synthetic single-turn pairs, selected historical models and noisy human labels limit transfer. The study cannot explain the cause of a particular Voice heading.
- Implication for Voice
- Benchmark usefulness and specificity independently of a judge's preference for polished presentation.
Use with Voice
These instructions support review; they neither calibrate a judge nor establish human acceptance.
getWritingRule— Retrieve contextual wording and format guidance.composeWritingReviewPrompt— Request concrete reader impact and quoted evidence.
What is missing: Compare controlled alternatives with blinded human adjudication; test whether decorative phrasing changes model judgments despite equivalent substance.
Related patterns: Taxonomy before the task · Ornamental wording · Formatting on every sentence · Unsupported completion claim
Measuring AI ‘Slop’ in TextWhich observable text problems contribute to judgments of AI slop?
- Method
- Nineteen expert responses informed a taxonomy. Three final copy-editors annotated English news and MS MARCO answers; regressions related span labels to overall judgments.
- Findings
Relevance, information density and tone were prominent predictors; factuality and structure also mattered.
Annotators agreed more on problematic spans than on the binary slop label; automatic judges reproduced human judgments poorly.
- Limits
Preprint with a small, calibrated expert panel and two domains. Subjective labels and limited English coverage constrain generalization.
- Implication for Voice
- Assess separate, contextual problems with quoted evidence; preserve reviewer disagreement instead of assigning an overall slop score.
Use with Voice
The API supports a review organized around selected problems. It does not implement the paper's annotation scheme or an overall slop classifier.
composeWritingReviewPrompt— Request separate goal-linked findings with exact quotations.mountVoiceReview— Show host-supplied span findings for inspection.
What is missing: Build page context, validate quoted evidence and coverage, and retain differing reviewer decisions without collapsing them into a quality score.
Related patterns: Taxonomy before the task · Optional-detail detours · Claim without evidence · Summary loops · Formatting on every sentence
‘That's AI Slop, You Bot!’ Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated CommentsDo online AI accusations track writing features or social boundaries?
- Method
- Approximately 25 million Hacker News/Reddit comments, January 2023–May 2026; regex screening, Claude Opus 4.7 coding, and 421 accused Reddit comments compared with 2,048 matched controls.
- Findings
Pejorative accusations grew. Table 4 associates longer mean tokens with lower accusation odds (OR 0.78, p<0.001); the other five tested prose markers were not statistically significant.
- Limits
Preprint using English platforms, model-coded categories and a limited lexicon; authorship was not independently established.
The prose claims none of four distinguishing markers predicted accusations, but Table 4 reports a significant mean-token-length coefficient. This inconsistency limits the stated null conclusion.
- Implication for Voice
- Separate a reader's negative reaction from claims about authorship; request the concrete wording, relevance or evidence problem.
Use with Voice
A host can turn a complaint into a review request. The library assigns no authorship label and does not assess a commenter's credibility.
getWritingRule— Retrieve evidence and uncertainty guidance for a disputed accusation.composeWritingReviewPrompt— Request a concrete quoted problem and its reader impact.
What is missing: Separate observed wording, reviewer interpretation and supporting evidence in validated findings, including conflicting evidence and unresolved claims.
Related patterns: Claim without evidence · Hedge stacking · Ornamental wording
Why Does ChatGPT ‘Delve’ So Much? Exploring the Sources of Lexical Overrepresentation in Large Language ModelsWhy are some words disproportionately frequent in generated scientific abstracts?
- Method
- PubMed frequency analysis; 9,953 GPT-3.5 reconstructions of 10,000 abstracts; Llama-2 Base/Chat entropy comparison; an exploratory preference study recruiting 201 participants in India.
- Findings
Twenty-one focal words combined rising corpus frequency with GPT-3.5 overuse.
Model differences were compatible with post-training effects, but participants showed no overall preference for focal-word abstracts; delve-initial items were disfavored.
- Limits
The online study was underpowered after exclusions and splitting conditions. Forced vocabulary could alter meaning; unavailable training data prevents causal attribution to RLHF.
- Implication for Voice
- Treat lexical cues as contextual review candidates, without attributing them to a single training mechanism or banning words.
Use with Voice
The current matcher does not implement the study's focal-word list or detect “delve.” Its candidates require a separate contextual judgment.
getWritingRule— Retrieve word-choice guidance and legitimate vocabulary exceptions.findWritingSignals— Locate only Voice's published stock-word patterns.
What is missing: Evaluate vocabulary alternatives against meaning, domain and audience, then benchmark warranted changes and legitimate original wording by language.
Related patterns: Ornamental wording · Invented cause and benefit · Hedge stacking
Delving into LLM-assisted writing in biomedical publications through excess vocabularyCan aggregate word-frequency changes estimate LLM assistance in biomedical abstracts?
- Method
- Analyze 15.1 million English PubMed abstracts from 2010–2024; compare 2024 word occurrence rates with conservative extrapolations from 2021–2022, after removing contaminating metadata.
- Findings
Style words increased abruptly. Under the paper's assumptions, excess vocabulary implied a 13.5% lower-bound estimate of LLM-assisted 2024 abstracts.
- Limits
Corpus-level estimation cannot label individual abstracts or distinguish direct assistance from humans adopting fashionable words. Editing practices and publication delays complicate subgroup comparisons.
- Implication for Voice
- Use vocabulary shifts to motivate review hypotheses, never as evidence that a particular sentence is wrong or machine-authored.
Use with Voice
Hosts can collect these local results manually. The API performs no corpus-frequency analysis, historical comparison or estimate of AI assistance.
findWritingSignals— Return candidate wording and coverage for one supplied text.getWritingRule— Explain why a lexical cue needs contextual review.
What is missing: Review vocabulary recurrence across versioned Git documents with explicit corpus scope, distinguishing inconsistent terminology from deliberate repetition.
Related patterns: Ornamental wording · Claim without evidence · Hedge stacking
Generative AI enhances individual creativity but reduces the collective diversity of novel contentCan assistance improve individual writing while making a collection less diverse?
- Method
- Preregistered randomized experiment: 293 UK participants wrote eight-sentence stories with no GPT-4 idea, one available idea, or up to five. Six hundred evaluators supplied ratings; embeddings measured similarity.
- Findings
Access to ideas improved novelty and usefulness ratings, especially for lower baseline-creativity writers, while assisted stories became more similar to one another.
- Limits
Short fiction, nonprofessional participants and fixed prompts without iterative conversation; neither professional documentation nor long-term creativity was tested.
- Implication for Voice
- Evaluate usefulness and distinctiveness separately; keep effective assistance while checking repeated framing across pages or alternatives.
Use with Voice
Prompt instructions ask for distinctness; no API generates alternatives, measures their similarity or ranks usefulness automatically.
composeWritingReviewPrompt— Request one to three distinct alternatives when change is warranted.selectWritingRules— Include reader-goal and repetition in the review.
What is missing: Compare alternatives against the original goal and scan repeated framing across project pages, keeping usefulness and distinctiveness as separate review criteria.
Related patterns: Summary loops · Taxonomy before the task · Ornamental wording
Towards Understanding Sycophancy in Language ModelsCan agreeable feedback outrank evidence or truthful correction?
- Method
- Prompt perturbations tested Claude 1.3/2.0, GPT-3.5/GPT-4 and Llama-2-70B-Chat. Analysis included 15,000 human-preference pairs and a 266-misconception experiment.
- Findings
User preferences shifted critiques and answers; some correct answers were abandoned after challenge.
Humans generally preferred helpful truthful replies, but persuasive agreement with misconceptions sometimes won, particularly on harder questions.
- Limits
Historical models, model-assisted judgments and a proof-of-concept misconception set. Human raters could not fact-check externally; these are not current-model failure rates.
- Implication for Voice
- Require reasons and evidence for both keeping and changing text; agreement with the requester is insufficient validation.
Use with Voice
A host can use these instructions with its chosen agent. Prompt wording alone neither verifies truth nor prevents agreement-driven revisions.
composeWritingReviewPrompt— Require evidence-linked reasons for keeping or changing a passage.getWritingRule— Expose claims-evidence and calibrated-uncertainty guidance.
What is missing: Check revised claims against source evidence and benchmark whether changing the requester's expressed preference changes an otherwise identical review.
Related patterns: Claim without evidence · Invented cause and benefit · Hedge stacking · Unsupported completion claim
GPT detectors are biased against non-native English writersCan detector scores confuse language background with machine authorship?
- Method
- Seven off-the-shelf detectors, accessed March 2023, evaluated 91 human TOEFL essays and 88 US eighth-grade essays; additional experiments changed vocabulary using ChatGPT.
- Findings
Average false positives were approximately 61% for TOEFL essays, versus roughly 5% for US essays.
Vocabulary enrichment reduced TOEFL false positives; self-editing also reduced detection of generated essays.
- Limits
Small, unmatched educational corpora and historical detector versions; findings do not establish identical bias in every language, genre or current detector.
- Implication for Voice
- Preserve non-native expression and judge reader impact directly; fluency or predictability cannot establish authorship.
Use with Voice
These APIs support contextual review without an authorship score. Localized instructions do not establish equivalent review quality for different language backgrounds.
getWritingRule— Show word-choice exceptions before treating wording as a defect.composeWritingReviewPrompt— Supply audience, goal and English or German instructions.
What is missing: Test accepted originals and warranted edits across language backgrounds and domains, reporting false positives without using fluency as an origin label.
Related patterns: Ornamental wording · Claim without evidence · Hedge stacking
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsDo detector results survive changes in generator, domain and sampling?
- Method
- Benchmark 12 detectors on over six million generated examples spanning 11 generators, eight domains, four decoding settings and 11 attacks; compare detection at a fixed 5% false-positive rate.
- Findings
Unseen generators, sampling, repetition penalties and text modifications reduced detection performance. Detector rankings depended on the allowed false-positive rate.
- Limits
Core coverage is English; multilingual extensions cover news. Models age, and optimizing against a public benchmark can undermine apparent out-of-domain generalization.
- Implication for Voice
- Report locale, domain and tested source conditions; never translate a detector benchmark score into a writing-quality score.
Use with Voice
Returned coverage describes lexical scanning, not measured detection accuracy. The library does not import RAID or run detector evaluations.
findWritingSignals— Expose scanned rules, unscanned rules, language and truncation.composeWritingReviewPrompt— Request explicit review coverage limits.
What is missing: Evaluate review prompts across declared domains, locales and edited inputs; validate result provenance and distinguish fixture coverage from measured review outcomes.
Related patterns: Unsupported completion claim · Claim without evidence · Hedge stacking
LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected?What happens when humans and models both edit a document?
- Method
- MixSet combines polishing, completion, rewriting and adaptation across six text types, using GPT-4, Llama-2-70B and eight human editors. Experiments compare binary/three-class detection and transfer between operations.
- Findings
Subtle editing challenged detectors; training on mixed examples improved some results, but transfer depended on editing operation and generator.
- Limits
Constructed English scenarios use few generators; some humanization is model-simulated. Binary experiments label all mixed text machine-generated, a study convention rather than an authorship fact.
- Implication for Voice
- Retain source revisions and accepted edits; review each passage's purpose without forcing collaborative writing into a binary origin label.
Use with Voice
A host can review a collaboratively edited passage without assigning an origin label. The overlay checks text mapping but does not maintain revision history.
mountVoiceReview— Anchor host-supplied findings to matching rendered text units.composeWritingReviewPrompt— Request source references while preserving accepted decisions.
What is missing: Bind decisions to Git revisions and validated source spans, detecting when subsequent edits invalidate the passage a reviewer accepted.
Related patterns: Claim without evidence · Hedge stacking · Unsupported completion claim
Counter-strategies and their evidence
A counter-prompt is an intervention to test. These studies examine training incentives, iterative feedback and human control; none qualifies the current Voice prompts across Claude, Codex and Gemini.
How the prompts use this evidence
Each selected rule contributes its symptom, reader impact, editing instruction, conditional before/after example and an acceptable keep example. The composer adds audience and goal, source boundaries, justified alternatives and coverage requirements.
Implementation: The host runs the prompt. The quality benchmark evaluates supplied results against frozen inputs and separate reviewer judgments. The package does not fine-tune models, run a self-revision loop or treat a model’s approval as independent verification.
See the Library’s exact analysis coverageDisentangling Length from Quality in Direct Preference Optimizationcompare alternatives at similar lengths and assess task completion separately. Treat length as a diagnostic, preserving explanations and qualifications the reader needs.
- Method
- Pythia 2.8B was trained with standard and length-regularized DPO on dialogue and summarization preferences. Evaluation compared 256 generated answers per setting using GPT-4 judgments.
- Finding
- Regularization reduced answer length and improved preference results against settings producing similarly long answers. Uncontrolled preference optimization amplified verbosity.
- Limits
- One model size, two datasets and a model judge. This training intervention does not establish that a shortness prompt improves documentation.
Use with Voice
A host can request alternatives without imposing a blanket shortness rule. The API does not optimize a model, judge completion or compare answer lengths.
composeWritingReviewPrompt— State the reader task and select relevance or repetition guidance.
What is missing: Compare alternatives using explicit task completion and retained qualifications alongside length, with fixtures that distinguish useful explanation from unnecessary expansion.
Related patterns: Summary loops · Optional-detail detours · Unprioritized laundry list · Formatting on every sentence
Self-Refine: Iterative Refinement with Self-Feedbackrequest a specific, evidence-linked critique before alternatives. Keep revision bounded, compare against the original and allow the reviewer to retain it.
- Method
- The same model generated, critiqued and revised outputs using task-specific prompts across seven tasks. Evaluation combined task metrics, model judgments and blinded author judgments.
- Finding
- Refinement improved the tested task results; human judges preferred refined outputs in dialogue, sentiment reversal, acronym generation and code readability.
- Limits
- English tasks and mainly proprietary 2023 models. Human evaluation usually used one author judgment per example. Better dialogue can also mean more elaborate text.
Use with Voice
Hosts must submit prompts and manage any revision loop themselves. maxFindings requests an output limit; it neither enforces compliance nor bounds inference spending.
composeWritingReviewPrompt— Prepare a bounded critique request with grounded alternatives.selectWritingRules— Narrow a subsequent review to the relevant rule IDs.
What is missing: Validate each critique and compose scoped review batches with explicit usage budgets before a host runs another round.
Related patterns: Taxonomy before the task · Summary loops · Claim without evidence · Optional-detail detours
Large Language Models Cannot Self-Correct Reasoning Yeta model's critique is a proposal, not verification. Check revised claims against source evidence and retain the original when the critique lacks support.
- Method
- Experiments compared initial answers and up to two self-correction rounds on reasoning benchmarks, separating feedback with known correct answers from model-only feedback.
- Finding
- Without correctness labels, the tested models often changed correct answers into errors. Stronger initial prompts also reduced apparent benefits attributed to refinement.
- Limits
- The ICLR 2024 study tests older models and reasoning tasks. Its authors explicitly distinguish style preferences, where self-correction may help.
Use with Voice
These instructions can guide a second opinion, but the library has no fact checker and cannot establish that a revised answer is better.
composeWritingReviewPrompt— Ask for missing evidence and justified keep decisions.getWritingRule— Retrieve claims-evidence and reported-evidence requirements.
What is missing: Verify each proposed claim change against supplied evidence and compare the revision with the original before accepting any correction.
Related patterns: Claim without evidence · Unsupported completion claim · Invented cause and benefit · Hedge stacking
CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilitiesmake alternatives optional and record keep, accept and edit decisions. Evaluate author intent directly instead of treating acceptance rate as success.
- Method
- Sixty-three writers completed 1,445 English writing sessions with GPT-3 suggestions. The interface logged requests, selections, dismissals and edits; surveys measured ownership and satisfaction.
- Finding
- Perceived ownership correlated with the fraction written by the writer. Editing activity alone showed little relationship with ownership; collaboration varied considerably between writers.
- Limits
- This CHI 2022 dataset offers correlations, not a randomized test of intent preservation. Participants were qualified crowd workers writing short assigned stories and essays.
Use with Voice
Hosts provide the choice panel and persist its actions; mounting the library alone does not record acceptance, dismissal or editing history.
mountVoiceReview— Expose selection and finding callbacks to the host's review interface.composeWritingReviewPrompt— Request optional alternatives or a justified decision to keep.
What is missing: Preserve author decisions with source versions and compare alternatives against stated intent, without treating acceptance counts as a quality measure.
Related patterns: Taxonomy before the task · Claim without evidence
Compose a small review brief
The current Library combines a selected rule profile with the audience and goal. A host can add supported facts, source text and accepted decisions as separately identified context. The prompt asks for specific findings and up to three alternatives tied to an identified problem; keeping the text remains an available decision.
Integrate writing tools in your application · Use the typed API
Libraries and engines
Capabilities, licenses and maintenance were checked against project-owned sources on 12 September 2026. The integration assessments below are our technical judgment. This comparison examines documented capabilities. It does not report installed integrations or comparative performance results.
Adapters for additional static analysis
Vale can supply Git-owned terminology and style rules; textlint can supply markup-aware rule execution. LanguageTool is a candidate for English/German grammar. Harper supplies English browser-local proofreading. The records below describe other options and their exact limits.
Implementation: No listed engine is bundled or connected yet. An adapter must retain engine/version, source text and verified spans. Engine findings can be displayed with @voice/review after host validation; contextual heading quality and factual support still require separate review.
See the Library’s exact analysis coverage| Project | Execution and API | Language coverage |
|---|---|---|
| Vale | Local Go CLI; files or stdin, JSON diagnostics; no model required. | English defaults; German requires appropriate rules and a Hunspell dictionary. |
| textlint | Local JavaScript CLI, createLinter/lintText API, or universal @textlint/kernel; plugins determine external calls. | Language coverage comes from plugins; English rules are listed, but no bundled German grammar pack was verified. |
| retext / retext-simplify | Local ESM plugin for unified/retext in Node, Deno or browsers; no model required. | English simplification examples; Latin-script parsing also accepts German, but no German simplification pack was verified. |
| write-good | Local Node function writeGood(text, options) and CLI; no model required. | English defaults; README links a separate basic German extension, schreib-gut, not evaluated here. |
| LanguageTool | Local Java engine or self-hosted HTTP JSON API; cloud service is a separate option. | English and German language modules, plus additional languages and regional variants. |
| Harper | Local Rust engine and harper.js WebAssembly package; WorkerLinter is recommended for interactive browsers. | English only; German is not currently supported. |
| proselint | Local Python LintFile API and CLI; JSON diagnostics and pre-commit integration; no model required. | English prose; German coverage is not documented in the reviewed README. |
| promptfoo | Local Node CLI/evaluate API; deterministic assertions or model-assisted grading. Configured remote providers receive evaluation inputs. | Language-neutral test fixtures; English/German evaluation quality depends on the chosen assertions and models. |
Valesuitable for Git-owned style packs and CI; adapt its diagnostics to Voice findings and preserve human decisions.
Official repository · MIT license
- Execution
- Local Go CLI; files or stdin, JSON diagnostics; no model required.
- Languages
- English defaults; German requires appropriate rules and a Hunspell dictionary.
- Capabilities
Versioned YAML rules, terminology vocabularies and configurable severity.
Markup scopes distinguish prose, headings and code; substitution rules supply suggestions.
- Limits
Style consistency is its stated focus, rather than comprehensive grammar checking.
Assessment: pattern matches cannot establish factual truth, page strategy or AI authorship.
- Maintenance snapshot · 2026-09-12
- Latest GitHub release: v3.21.0, published 2026-09-09; release API checked.
Use with Voice
There is no bundled Vale adapter. A host can already display manually converted results, provided the unit text matches the rendered document exactly.
mountVoiceReview— Render findings after a host maps external diagnostics to Voice units.
What is missing: Adapt Vale diagnostics with rule provenance, severity and markup-aware source positions; validate UTF-16 spans and attach suggestions to the reviewed source version.
Related patterns: Ornamental wording · Summary loops · Invented compound labels · Formatting on every sentence
textlinta close JavaScript integration fit for reusable rule packs, source diagnostics and reviewable fixes.
Official repository · MIT license
- Execution
- Local JavaScript CLI, createLinter/lintText API, or universal @textlint/kernel; plugins determine external calls.
- Languages
- Language coverage comes from plugins; English rules are listed, but no bundled German grammar pack was verified.
- Capabilities
Markdown/plain-text parsing, custom rules, filters and formatters.
Fix APIs return proposed results without writing files.
- Limits
No rules ship in the core; each plugin needs separate quality and license review.
Assessment: AST rules alone do not judge positioning or authorship.
- Maintenance snapshot · 2026-09-12
- Latest release v15.8.0: 2026-08-01; API documentation updated 2026-09-09.
Use with Voice
Voice does not currently load textlint rules or call its fix API. A consumer owns parsing, plugin selection and conversion before supplying findings.
mountVoiceReview— Display host-converted lint findings without modifying source text.
What is missing: Convert textlint messages and proposed fixes into validated Voice findings, retaining plugin identity, exact ranges and source-version preconditions for later review.
Related patterns: Ornamental wording · Unnecessary interpretive framing · Summary loops · Formatting on every sentence
retext / retext-simplifyreuse the suggestion/message structure for optional word-choice hints; retain source spans and a keep-as-written choice.
Official repository · MIT license
- Execution
- Local ESM plugin for unified/retext in Node, Deno or browsers; no model required.
- Languages
- English simplification examples; Latin-script parsing also accepts German, but no German simplification pack was verified.
- Capabilities
Emits positioned VFile messages with ruleId, actual wording and expected alternatives.
Supports ignored phrases and composition with other prose plugins.
- Limits
Assessment: phrase substitutions cannot determine whether a technical term is necessary or a claim is supported.
- Maintenance snapshot · 2026-09-12
- Latest simplify release: 8.0.0, published 2023-09-10; no newer GitHub release observed.
Use with Voice
No retext plugin runs inside Voice. Consumers currently need their own VFile-message conversion and decision interface.
getWritingRule— Retrieve word-choice guidance for technical terms and acceptable originals.mountVoiceReview— Display a suggestion after host conversion to a finding.
What is missing: Adapt positioned suggestions into validated UTF-16 findings and compare replacements with the original meaning, preserving an explicit decision to keep the phrase.
Related patterns: Ornamental wording · Mechanism without an operation · Invented compound labels
write-gooduseful as a small comparison baseline for local signals, with context-sensitive review outside the engine.
Official repository · MIT license
- Execution
- Local Node function writeGood(text, options) and CLI; no model required.
- Languages
- English defaults; README links a separate basic German extension, schreib-gut, not evaluated here.
- Capabilities
Flags passive voice, repeated words, clichés, adverbs and wordy phrases.
Returns reason, index and length; supports disabled checks, custom checks and whitelists.
- Limits
The project describes its method as naive.
Assessment: blanket passive/adverb warnings can reject useful prose and cannot establish AI authorship.
- Maintenance snapshot · 2026-09-12
- GitHub releases page contains no releases; latest package publication date remains unverified.
Use with Voice
Voice neither calls write-good nor inherits its passive-voice or adverb checks. Its published patterns form a separate, limited rule set.
findWritingSignals— Provide Voice's own local cue results as a comparison baseline.mountVoiceReview— Render externally supplied findings after host mapping.
What is missing: Adapt enabled checks with engine attribution and evaluate their false positives on useful passive constructions, necessary qualifications and accepted wording.
Related patterns: Ornamental wording · Summary loops · Hedge stacking · Forced triads and canned Q&A
LanguageToola bilingual proofreading adapter candidate alongside Voice's contextual review; declare the chosen local/cloud mode.
Official repository · LGPL-2.1-or-later license
- Execution
- Local Java engine or self-hosted HTTP JSON API; cloud service is a separate option.
- Languages
- English and German language modules, plus additional languages and regional variants.
- Capabilities
Grammar, spelling and style diagnostics with offsets and replacement candidates.
Annotated text preserves position relationships around markup.
- Limits
Self-hosted server excludes cloud AI rules and synonyms.
Assessment: proofreading does not establish project strategy or AI authorship.
- Maintenance snapshot · 2026-09-12
- Changelog dates 6.8 to 2026-05-05 and marks 6.9 unreleased; ZIP distribution switched to snapshots after 6.6.
Use with Voice
The review library starts no Java process or LanguageTool request. A consumer must currently choose the engine mode and supply mapped findings itself.
mountVoiceReview— Display proofreading findings converted by the host.
What is missing: Add an explicit local or cloud adapter with locale selection, validated positions across annotated markup and reviewable replacements bound to the source version.
Related patterns: Arrow-chain shorthand · Ornamental wording · Summary loops · Formatting on every sentence
Harperpromising for English browser-local proofreading and explicit dismissals; German needs another engine.
Official repository · Apache-2.0 license
- Execution
- Local Rust engine and harper.js WebAssembly package; WorkerLinter is recommended for interactive browsers.
- Languages
- English only; German is not currently supported.
- Capabilities
Async lint API, replacement suggestions, configurable rules and custom dictionaries.
Context-sensitive ignored-lint hashes can be exported and imported.
- Limits
Documentation still labels harper.js early access with an unstable API.
Assessment: verify Unicode span conversion against Voice's UTF-16 contract; grammar findings are not authorship evidence.
- Maintenance snapshot · 2026-09-12
- Latest GitHub release: v2.10.0, published 2026-09-09; release API checked.
Use with Voice
Voice currently starts no Harper worker and imports no ignored-lint state. English proofreading requires a consumer-owned integration.
mountVoiceReview— Render host-supplied proofreading findings on matching text units.
What is missing: Adapt worker results with tested Unicode-to-UTF-16 conversion and preserve ignore decisions alongside engine and source versions, so changed passages can be reviewed again.
Related patterns: Arrow-chain shorthand · Ornamental wording · Summary loops
proselintuseful reference and CI baseline for explicit negative writing patterns, especially jargon and empty narration.
Official repository · BSD-3-Clause license
- Execution
- Local Python LintFile API and CLI; JSON diagnostics and pre-commit integration; no model required.
- Languages
- English prose; German coverage is not documented in the reviewed README.
- Capabilities
Checks clichés, corporate jargon, metadiscourse, repetition and typography.
Granular per-file configuration; diagnostic spans and optional replacements.
- Limits
Assessment: advice against hedging can conflict with necessary uncertainty; evaluate selected checks on Voice examples.
Heuristic matches do not establish authorship or factual correctness.
- Maintenance snapshot · 2026-09-12
- Latest GitHub release: v0.16.0, published 2025-11-14; repository currently describes itself as a mirror.
Use with Voice
No Python engine or proselint configuration is included in the review library. Consumers must currently convert and attribute its results.
getWritingRule— Retrieve calibrated-uncertainty exceptions when assessing a hedging warning.mountVoiceReview— Display host-converted diagnostics for contextual review.
What is missing: Adapt enabled checks and benchmark jargon, metadiscourse and hedging warnings against examples where technical terminology or uncertainty is necessary.
Related patterns: Ornamental wording · Unnecessary interpretive framing · Summary loops · Hedge stacking
promptfooregression-test Voice counter-prompts against Git-owned examples and accepted decisions; separate deterministic checks from model judgments.
Official repository · MIT license
- Execution
- Local Node CLI/evaluate API; deterministic assertions or model-assisted grading. Configured remote providers receive evaluation inputs.
- Languages
- Language-neutral test fixtures; English/German evaluation quality depends on the chosen assertions and models.
- Capabilities
Compares prompt/model outputs using regex, JSON schemas, custom functions and rubrics.
Supports saved-output evaluation and CI integration.
- Limits
An evaluation framework, not a prose linter or validated authorship detector.
Assessment: scores inherit fixture coverage and judge limitations.
- Maintenance snapshot · 2026-09-12
- Latest GitHub release: 0.123.0, published 2026-09-10; release API checked.
Use with Voice
Voice prepares inputs but has no promptfoo adapter or evaluation runner. Provider calls, assertions and saved-output comparisons remain host responsibilities.
composeWritingReviewPrompt— Generate versioned prompt text for a consumer-owned evaluation fixture.selectWritingRules— Choose a stable rule subset for comparisons.
What is missing: Run Git-owned fixtures that record catalogue and provider settings, preservation decisions, separate deterministic and model judgments, and measured usage.
Related patterns: Claim without evidence · Unsupported completion claim · Optional-detail detours · Forced neutrality and false balance
How the research becomes a review
Voice's rule catalogue stores 20 negative patterns with reader costs, examples, exceptions and counter-prompts. Four profiles select six rules each. The browser's local scanner covers eight rules; it reports exact UTF-16 spans and explicitly identifies rules requiring semantic review.
| Question | Evidence to supply | Current Voice boundary |
|---|---|---|
| Does this wording obscure the point? | Exact quote, surrounding text, reader task and the relevant rule. | Local candidate signals and composed review prompts are available in the writing-rules package. |
| Does the whole page serve its goal? | Page purpose, audience, complete content, accepted positioning and supporting sources. | A host must run and validate a contextual analysis. The local scanner does not assess page strategy. |
| Is the project consistent? | Selected files at a known revision, terminology, cross-page claims and review decisions. | Git can retain reviewed records. Automated project-wide SEO and visual review are not implemented. |
| Did AI compose this passage? | Known generation provenance or a separately qualified attribution method. | The package supplies no authorship verdict or probability. |
The rule package and the review UI have separate responsibilities. @voice/writing-rules selects rules, composes prompts and identifies surface cues. @voice/review maps supplied findings onto text. Studio currently runs its existing four local checks; the expanded catalogue is not automatically connected to its backend.
Measure a remedy against preserved intent
Our proposed evaluation unit is a source passage plus its purpose, supporting facts and accepted decisions. Compare the unchanged source with revisions from a versioned prompt. Reviewers should assess task completion, factual preservation, unnecessary text and unwanted changes separately. Include legitimate examples of each pattern, German and English material, and human-edited model drafts.
These are qualification criteria for future model comparisons. The current tests verify the catalogue, prompt composition, local matching and browser behavior; they do not measure improved writing quality from model use.
Explore the pattern catalogue · See the review UI · Read the detection boundaries
Writing tools and HTML visualization
| Package | Responsibility | Result |
|---|---|---|
@voice/writing-rules | Rule knowledge, prompt composition, static cues and a provider-neutral LLM tool interface. | Inspectable instructions, exact candidate spans and host-supplied model evidence. |
@voice/review | Render host-supplied findings and feedback in HTML; handle text selection and source mapping. | Annotations and callbacks for your review interface. |
The host connects these packages: it supplies source context and available models, executes an admitted model call, validates its findings and persists decisions. Studio provides a local application around the review workflow.
What the static scanner does
findWritingSignals runs 16 fixed English/German regular expressions across raw text, covering 8 of 20 rules. It returns the matched quote, UTF-16 start/end positions, rule ID, covered and unscanned rules, and truncation status. For example, it can locate it is worth noting or könnte möglicherweise.
The scanner does not parse sentences or markup, exclude code/quotations, interpret negation, check facts or assess the purpose of a heading. not guaranteed can still produce a candidate. The rejected heading Which model and prompt earn their cost? produced zero local signals: its problem requires context.
All 20 rules and their analysis method
| Rule | Review concern | Implemented method |
|---|---|---|
reader-goal | Taxonomy before the task | Contextual prompt |
current-state | Roadmap inside the usage guide | Contextual prompt |
process-history | Development diary | Contextual prompt |
meta-framing | Unnecessary interpretive framing | Fixed phrase candidates + contextual prompt |
concrete-subject | Mechanism without an operation | Fixed phrase candidates + contextual prompt |
invented-labels | Invented compound labels | Contextual prompt |
unsupported-contrast | Not X, but Y | Fixed phrase candidates + contextual prompt |
calibrated-uncertainty | Hedge stacking | Fixed phrase candidates + contextual prompt |
unsupported-causality | Invented cause and benefit | Fixed phrase candidates + contextual prompt |
claims-evidence | Claim without evidence | Fixed phrase candidates + contextual prompt |
repetition | Summary loops | Contextual prompt |
format-fit | Formatting on every sentence | Contextual prompt |
complete-sentences | Arrow-chain shorthand | Contextual prompt |
word-choice | Ornamental wording | Fixed phrase candidates + contextual prompt |
task-relevance | Optional-detail detours | Contextual prompt |
reported-evidence | Unsupported completion claim | Fixed phrase candidates + contextual prompt |
unearned-praise | Unearned praise and automatic agreement | Contextual prompt |
false-balance | Forced neutrality and false balance | Contextual prompt |
exhaustive-checklist | Unprioritized laundry list | Contextual prompt |
forced-template | Forced triads and canned Q&A | Contextual prompt |
What enters the review prompt
composeWritingReviewPrompt includes the selected rule instructions, examples and keep cases. Rule evidence links and the complete exception list remain available through the catalogue and the rule tool. Research records become package behavior only through a reviewed catalogue or analyzer change and a new build.
Catalogue voice-writing-rules/0.2.2 incorporates the heading-feedback lesson into wording, format and framing guidance. Broader analyses—cross-page repetition, claim preservation, alternative diversity and heading-to-content alignment—remain contextual review tasks, with results to test against independently reviewed examples.
Use Voice through LLM tool calls · Reproduce a quality evaluation · Use the two packages
Model selection and prompt structure
Compare review quality across models and prompt structures, then compare cost. Use the same source cases and selected rules. Record missed issues, unnecessary changes, justified keep decisions, preserved facts and intent.
Evaluate the result
The offline quality evaluator validates recorded responses against the requested rules, source version and exact quoted spans. Independent human reviewers assess the decisions and every proposed alternative. Missing judgments remain unassessed; disagreement remains visible. A justified request for missing evidence can pass editorial review while leaving the task incomplete.
Quality is evaluated even when cost is unknown. Sparse authored labels supply separate rule-and-quote diagnostics; agreement with those labels cannot replace human judgment or establish a model ranking.
Compare prompt organizations
| Organization | Question to test | Fair comparison |
|---|---|---|
| One rule per request | Does narrow focus recover problems a bundle misses? | Count every request and deduplicate findings for the complete review. |
| Two related rule groups | Does grouping retain focus while sharing context? | Use the same six rules, split into two fixed groups. |
| One six-rule profile | Can one compact review preserve the same useful findings? | Use the same cases, rules and total finding allowance. |
| All twenty rules | Does broader review justify extra instructions and output? | Compare twenty rules separately and together in their own cohort. |
| Cheaper review, selective escalation | When does a second model improve the result enough to pay for both? | Define escalation triggers and audit cases that were not escalated. |
The first four organizations have a local experiment preparer. Selective escalation is a later experiment. None has model-quality results yet.
How to assemble a review prompt
Keep the rule set, response contract and reviewed keep examples stable within an experiment. Supply the reader goal, relevant facts, accepted decisions and versioned source separately. Ask for exact quotes, rule IDs and concise reasons; generate alternatives only for justified changes.
Review instructions + selected rules + keep examples
+ audience and reader goal
+ versioned source, relevant facts and retained decisions
→ findings, coverage and justified keep/change choicescomposeWritingReviewPrompt currently puts audience and goal before the rule text; the host supplies the document separately. A reusable cached rubric prefix requires a different serialization and its own comparison. Long rule sets, compressed examples and reordered context can change what a model notices. Measure those changes rather than assuming that fewer tokens preserve quality.
Local prompt measurements
Requests constructed with the Library cover 30 authored English/German cases. Each comparison uses the same rules, source, audience and goal. The requests were measured locally; none was sent to a model.
| Same scope | Separate / bundled requests | Separate UTF-8 bytes | Bundled UTF-8 bytes | Input reduction |
|---|---|---|---|---|
| 6 rules | 180 / 30 | 386,724 | 191,479 | 50.5% |
| 20 rules | 600 / 30 | 1,281,035 | 539,104 | 57.9% |
Bytes are not tokens. No tokenizer was available for this run. The measurement excludes provider message framing, cache behavior, output, reasoning, retries and latency. Separate requests also allow more findings in total under the same per-request limit. These results establish reduced repeated input, not equivalent review quality or a percentage cost saving.
The same fixtures produced 24 local scanner candidates: 16 agreed with the authored review labels and 8 occurred on deliberate keep cases. All 4 semantic problem cases lacked a candidate. These examples demonstrate why local signals need context; they do not estimate performance on a representative corpus.
Open measurements, source hashes and pricing JSON
Model candidates and token prices
These candidates were selected from official API documentation. Their ability to recognize Voice's patterns has not been measured. Rates below are USD per million ordinary input or billed output tokens, checked 2026-09-12, for standard short-context text requests.
| Model | Input / 1M tokens | Output / 1M tokens | Source |
|---|---|---|---|
gpt-5.6-luna | $0.20 | $1.20 | Official pricing |
claude-haiku-4-5-20251001 | $1.00 | $5.00 | Official pricing |
gemini-3.5-flash-lite | $0.30 | $2.50 | Official pricing |
gpt-6-astra | $10.00 | $50.00 | Official pricing |
Test Luna, Haiku and Flash-Lite as efficiency candidates on the same cases and rule scope. Include Astra in the same comparison; human review resolves disagreements. Record exact model versions and settings, including reasoning effort. English results do not establish German performance.
Select models from quality evidence
Compare available models within the same reviewed source set, language, rule scope and declared quality requirements. Two results can support different choices:
| Selection | Required evidence |
|---|---|
| Lowest cost meeting the quality requirements | Compare the cost of all attempts per accepted complete review. Include failed outputs and retries. Unknown cost or zero accepted complete reviews produces no cost ratio. |
| Highest observed issue detection | Compare independently confirmed issues found against the same reviewed issue set. Report missed issues, false changes and correct keep decisions alongside detection. A larger finding count alone does not establish better detection. |
If those selections differ, report both with their measured scope. Neither establishes that a model catches every issue. The current pilot has no measured model-quality results or qualified model recommendation.
Reproduce the evaluation
The evaluator reads a frozen experiment and run record: exact model identities, parameters, prompts, sources, every attempt, selected final responses and human judgments. Hash checks prevent silently evaluating an old run against changed inputs. Planned cases remain in the denominator when answers fail or never arrive.
The same recorded inputs and evaluator version produce the same offline report. New model responses can vary and need separate repetitions. Reports keep English and German, rule coverage and prompt conditions separate. Follow the benchmark method, file layout and JSON contracts.
Cache, batching and output costsWhat a token price table leaves out
Cache reuse needs a matching eligible prefix. Write charges, read charges and retention differ between providers; explicit Gemini caches also incur storage cost. The pricing snapshot records these conditions and an unresolved Google Batch cache-rate discrepancy. OpenAI Caching; Claude Caching; Gemini Caching.
Combining rules in one prompt differs from a provider's asynchronous Batch API. Batch discounts can suit offline benchmarks; completion windows make them a separate choice from an interactive review. OpenAI Batch; Claude Batch; Gemini Batch.
The host should record actual ordinary input, cache writes, cache reads and billed output or reasoning separately. Report cost per accepted complete review, with useful findings, latency and errors as separate measures. A smaller visible answer does not by itself prove a lower bill.
What research supports
Rule bundling shares one passage across checks. Sample packing shares one rubric across passages. A provider’s asynchronous Batch API schedules separate requests. The studies below test sample packing or context use; their results motivate our experiments without establishing the best Voice configuration.
Batch Prompting: Efficient Inference with Large Language Model APIs · peer-reviewed industry paper
Groups independent questions under shared demonstrations and returns indexed answers. Main experiments use code-davinci-002 and ten QA, arithmetic and NLI/NLU datasets. Reusing demonstrations reduces repeated input. Larger batches often reduce accuracy; the best size varies by task.
Limits: No EN/DE editorial review, current-model comparison or rewrite-quality measurement. Output-heavy work and long inputs can weaken the benefit. This tests multiple samples, not multiple rules on one sample.
Use with Voice — our assessment: selectWritingRules and composeWritingReviewPrompt: Freeze a consistent rubric and reproduce the prompt for each condition.
Proposed API work: composeWritingReviewBatch needs passage IDs and per-passage results; evaluateReviewPrompts should measure omissions, wrong assignments, quality and total cost across batch sizes.
Read the primary source · Sections 2–4, section 6 and Appendix B/Table 6; checked 2026-09-12.
Lost in the Middle: How Language Models Use Long Contexts · peer-reviewed TACL paper
Moves relevant information within controlled multi-document QA and key-value retrieval inputs. Models include MPT, LongChat, GPT-3.5 and Claude-1.3. QA performance commonly declines for central evidence. Claude performs almost perfectly on synthetic retrieval; repeating the query helps retrieval more than QA.
Limits: Older models and retrieval tasks do not establish failure of a twenty-rule editorial prompt. Context-window capacity alone does not establish that every relevant constraint is used.
Use with Voice — our assessment: selectWritingRules and composeWritingReviewPrompt: Make intended coverage explicit and reproduce the selected instructions.
Proposed API work: evaluateReviewPrompts should move important facts, rules and keep cases between beginning, middle and end while holding content constant. composeWritingReviewBatch should retain page context and measure position sensitivity.
Read the primary source · Sections 2–4 and section 5 retrieval case study; checked 2026-09-12.
Towards Cost-effective LLMs Routing with Batch Prompting · May 2026 preprint
Jointly selects the model and number of independent queries per prompt. Tests Qwen3 4B/14B/32B and Gemma3 4B/12B/27B on six benchmarks with separate training, validation and test partitions. Reports improved cost–accuracy trade-offs in many settings. Batching tolerance varies by model and task; larger batches can produce malformed outputs.
Limits: Not validation of Voice reviews or current proprietary models. Calibration requires evidence. Scheduling time excludes model API latency, and incomplete batches complicate shared-input cost accounting.
Use with Voice — our assessment: selectWritingRules and composeWritingReviewPrompt: Freeze rules and prompt versions across paired model comparisons.
Proposed API work: evaluateReviewPrompts should compare model and prompt structure jointly. composeWritingReviewBatch can produce execution candidates; routing remains in Token Economy and the host.
Read the primary source · Sections 2–6, including implementation and ablations; checked 2026-09-12.
Use this with the current API
Use selectWritingRules to fix the review scope and composeWritingReviewPrompt to build comparable requests. findWritingSignals supplies local candidate spans; an empty result must not skip requested semantic review. The host executes models and records their results. The offline evaluator validates supplied responses and evaluates the recorded human judgments.
Use the current Library functions · Read the LLM tool integration example
Sources and reusable research data
This is a curated collection of 7 practice sources, 11 studies, 4 strategy studies and 8 tools, selected for writing quality, review decisions and implementation fit. It is not a systematic review or a provider ranking. Source records include the sections read, limitations and related rule IDs; tool records link official documentation, licenses and release evidence.
The JSON records are maintained in Git under website/research/. They contain original summaries and links; third-party prompts, papers and library code are not bundled. Open a record set to inspect or reuse the structured data.
Source indexOriginal articles, papers and repositories
- Problem Patterns in AI: Beyond Hallucinations — Leon Furze (2026-07-26) · Research note
- Why ChatGPT writes like that: A rhetorical analysis of AI ‘slop’ — Colin Gorrie (2025-07-09) · Research note
- On AI-assisted writing, AI slop, and the line between them — Xavier Amatriain (2026-08-30) · Research note
- The ‘Negative Contrast Trap’: Why AI Writing Overuses ‘Not X, But Y’ — Ernan Hughes · Research note
- Tropes: AI Writing Pattern Directory — ossama.is · Research note
- Wikipedia: Signs of AI writing — WikiProject AI Cleanup contributors · Research note
- slopt: Stop sounding like ChatGPT — slopt · Research note
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits — Tuhin Chakrabarty, Philippe Laban et al. (2024-09-22) · Research note
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models — Anirudh Bharadwaj, Chaitanya Malaviya et al. (2025-06-05) · Research note
- Measuring AI ‘Slop’ in Text — Chantal Shaib, Tuhin Chakrabarty et al. (2025-09-23) · Research note
- ‘That's AI Slop, You Bot!’ Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments — Jason Miklian, John E. Katsos (2026-06-10) · Research note
- Why Does ChatGPT ‘Delve’ So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models — Tom S. Juzek, Zina B. Ward (2025-01) · Research note
- Delving into LLM-assisted writing in biomedical publications through excess vocabulary — Dmitry Kobak, Rita González-Márquez et al. (2025-07-02) · Research note
- Generative AI enhances individual creativity but reduces the collective diversity of novel content — Anil R. Doshi, Oliver P. Hauser (2024-07-12) · Research note
- Towards Understanding Sycophancy in Language Models — Mrinank Sharma, Meg Tong et al. (2023-10-20) · Research note
- GPT detectors are biased against non-native English writers — Weixin Liang, Mert Yuksekgonul et al. (2023-07-10) · Research note
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors — Liam Dugan, Alyssa Hwang et al. (2024-08) · Research note
- LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected? — Qihui Zhang, Chujie Gao et al. (2024-06) · Research note
- Disentangling Length from Quality in Direct Preference Optimization — Ryan Park, Rafael Rafailov et al. (2024-08) · Research note
- Self-Refine: Iterative Refinement with Self-Feedback — Aman Madaan, Niket Tandon et al. (2023) · Research note
- Large Language Models Cannot Self-Correct Reasoning Yet — Jie Huang, Xinyun Chen et al. (2024-03-14) · Research note
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities — Mina Lee, Percy Liang et al. (2022-01-25) · Research note
- Vale · Research note
- textlint · Research note
- retext / retext-simplify · Research note
- write-good · Research note
- LanguageTool · Research note
- Harper · Research note
- proselint · Research note
- promptfoo · Research note