AI-Resistant Assessment
ChatGPT can pass most editing tests. Not ours.
Every generic grammar and editing test you can buy online can be completed by a language model in seconds. That's a problem when you're trying to hire a real editor. Here's how our assessments handle it.
1. Question design that rewards human judgement
Multiple-choice grammar questions are easy for an LLM to handle. So our editing and proofreading tests are built around ambiguous and judgement-dependent tasks: which of these three technically-correct sentences is the right register for a scientific journal? Which of these edits preserves the author's voice? Which of these fact-checking moves is worth the editor's time?
Current models answer these noticeably worse than competent human editors because the right answer depends on context, audience, and editorial taste. We measure the gap continuously and retire items where it closes.
2. Tab-switch detection
During the test, we log every time the candidate leaves our browser tab. Excessive tab switching (especially around paste events) is a strong signal of external-tool consultation. The score report surfaces these events to hiring managers with timestamps.
3. Paste-event logging
We log paste events on every text-entry field. Copy-pasting an entire answer — or several large chunks in succession — is flagged. Candidates are told upfront that paste events are recorded; the honest ones don't paste.
4. Typing-cadence analysis
Humans type in bursts, with pauses for thought, backspaces, and re-edits. Text pasted from an LLM lands in a single event with no typing rhythm. Our stylometric module flags responses whose keystroke cadence doesn't match natural typing.
5. Human review on writing tasks
Every writing, editing, and proofreading submission is reviewed by human eyes. Our reviewers are editors themselves — they can spot the stylistic fingerprint of a language model in a few sentences, and they do.
6. Time-per-question anomalies
We measure per-question time. A candidate who sits on a hard question for 4 minutes and then produces a pristine answer in 8 seconds is flagged. So is the opposite: someone who speed-runs questions an experienced editor would normally pause on.
We publish our detection accuracy
Our most recent internal audit measured how reliably our combined signals flag a known-LLM-assisted submission:
- True-positive rate (LLM use correctly flagged): 89%
- False-positive rate (honest candidates wrongly flagged): 2.1%
- Sample: 400 assessments, 200 with scripted LLM use
No detector is perfect, and we're transparent about that. Flags are decision support — final hiring decisions stay with the hiring manager.
Want to see how it works?
Book a 15-minute walkthrough — we'll show you flagged and unflagged sample submissions side by side.
Book a demo →