Skip to main content

AI-Resistant Assessment

ChatGPT can pass most editing tests. Not ours.

Every generic grammar and editing test you can buy online can be completed by a language model in seconds. That's a problem when you're trying to hire a real editor. Here's how our assessments handle it.

1. Question design that rewards human judgement

Multiple-choice grammar questions are easy for an LLM to handle. So our editing and proofreading tests are built around ambiguous and judgement-dependent tasks: which of these three technically-correct sentences is the right register for a scientific journal? Which of these edits preserves the author's voice? Which of these fact-checking moves is worth the editor's time?

Current models answer these noticeably worse than competent human editors because the right answer depends on context, audience, and editorial taste. We measure the gap continuously and retire items where it closes.

2. Tab-switch detection

During the test, we log every time the candidate leaves our browser tab. Excessive tab switching (especially around paste events) is a strong signal of external-tool consultation. The score report surfaces these events to hiring managers with timestamps.

3. Paste-event logging

We log paste events on every text-entry field. Copy-pasting an entire answer — or several large chunks in succession — is flagged. Candidates are told upfront that paste events are recorded; the honest ones don't paste.

4. Typing-cadence analysis

Humans type in bursts, with pauses for thought, backspaces, and re-edits. Text pasted from an LLM lands in a single event with no typing rhythm. Our stylometric module flags responses whose keystroke cadence doesn't match natural typing.

5. Human review on writing tasks

Every writing, editing, and proofreading submission is reviewed by human eyes. Our reviewers are editors themselves — they can spot the stylistic fingerprint of a language model in a few sentences, and they do.

6. Time-per-question anomalies

We measure per-question time. A candidate who sits on a hard question for 4 minutes and then produces a pristine answer in 8 seconds is flagged. So is the opposite: someone who speed-runs questions an experienced editor would normally pause on.

We publish our detection accuracy

Our most recent internal audit measured how reliably our combined signals flag a known-LLM-assisted submission:

  • True-positive rate (LLM use correctly flagged): 89%
  • False-positive rate (honest candidates wrongly flagged): 2.1%
  • Sample: 400 assessments, 200 with scripted LLM use

No detector is perfect, and we're transparent about that. Flags are decision support — final hiring decisions stay with the hiring manager.

Want to see how it works?

Book a 15-minute walkthrough — we'll show you flagged and unflagged sample submissions side by side.

Book a demo →

Editors from these organisations have used our services since 1998

Reuters BBC Oxford University Press Penguin Random House Springer Microsoft Suncor Energy United Nations Fisher Investments IBM The Home Depot KODAK CHEVRON