Vahtian

A mirror, not a verdict

See the patterns in your own writing.

Paste a section. You get a list of sentences worth looking at again, in the order you wrote them, each with the reason and the case for leaving it alone.

There is no score. This is a writing tool, not an AI detector: it cannot tell whether text was written by a person or a machine, and nothing it produces may be used to accuse anyone of anything. Runs in your browser, nothing is uploaded, nothing is stored.

Telling the tool which section this is turns off the rules that do not apply there. Hedging belongs in a Discussion; a nominal, procedural sentence belongs in Methods. Nothing leaves the browser either way: no request, no account, no telemetry, no storage.

What this looks for

  • Structural: the antithesis family, both inside one sentence ("not just X, but Y") and split across two, where one sentence denies and the next supplies the replacement; restatement in place of development, parallelism overload, artificial comprehensiveness, signposting, formatting in place of argument, and a repeated paragraph template.
  • Language: vocabulary that could sell any product, empty intensifiers, generic abstraction, and stacked hedges.
  • Evidence: a claim handed to unnamed studies, a universal generalisation with nothing behind it, and a paragraph carrying no numeral, proper noun, or citation.
  • Expect findings on good writing. The rules were calibrated against a doctoral thesis written and defended before generative models existed, by a multilingual clinician whose prose from those years is scored around 60% "AI" by detectors. About 2500 words of it came back with four findings: one worth acting on, and three carrying their own case for leaving the sentence alone. Four more turned out to be faults in the rules, and the rules changed. That is one run on one thesis, not a measured error rate, and there is no benchmark here to quote. The story is in Writing discipline when detectors judge your prose.
  • Built for manuscript prose. The Evidence group assumes the citation conventions of a paper. Run an essay, a blog post, or a grant narrative through it and those rules will fire far more often, because ordinary prose carries claims without a citation in every sentence. That is the tool being out of its range, not your writing being weak.
  • Not style, and not voice. Sentence rhythm, sentence length, repeated openings, and nominalisation density are left alone on purpose. Penalising them makes everyone's prose converge, and they track writing in a second or third language more than they track quality. The tools that score those are the ones that flag a careful multilingual writer at 61%.

Each finding gives you the anchor, the sentence, why the pattern weakens the writing, what to do instead, and the conditions under which it is right as written. Counts appear per group, with a rate per 100 words so a long Discussion is comparable with a short Abstract. The groups are never added together, because a single number about writing is a grade, and a grade about writing gets read as a grade about the writer.

What it can't do. It cannot tell whether a person or a machine wrote the text, and no output may be used to accuse anyone. It reads shapes in sentences, not meaning: a flagged sentence may be exactly right, which is why every finding carries the case for keeping it. Good writing produces findings too, and no error rate has been measured for this. Clearing every item is not the goal. A draft with no findings is not a good paper, and a draft with twenty is not a bad one.

Checking whether the sources support the sentences instead? That is CiteVahti. Checking that the reference list resolves? That is the reference check.