Skip to main content

AI-Resistant Assessment

ChatGPT can pass most editing tests. Not ours.

Every generic grammar and editing test you can buy online can be completed by a language model in seconds. That's a problem when you're trying to hire a real editor. Here's how our assessments handle it.

1. Question design that rewards human judgment

Multiple-choice grammar questions are easy for an LLM to handle. So our editing and proofreading tests are built around ambiguous and judgment-dependent tasks: which of these three technically-correct sentences is the right register for a scientific journal? Which of these edits preserves the author's voice? Which of these fact-checking moves is worth the editor's time?

Current models answer these noticeably worse than competent human editors because the right answer depends on context, audience, and editorial taste. We also track how every multiple-choice question performs across real sittings: how often it is answered correctly, and whether the people who get it right are the people who do well on the rest of the test. A question that nearly everyone gets right, or that no longer separates strong candidates from weak ones, is retired from the bank.

2. Tab-switch detection

During the test, we log every time the candidate leaves our browser tab or exits fullscreen, with the time. Repeated switching, especially around paste events, is a signal of external-tool consultation. The score report and the candidate review page show each event with its timestamp, and the sitting is flagged at three or more.

3. Paste-event logging

We record every paste into an answer field: when it happened and how much text it carried (the text itself is not needed; the answer is saved anyway). A single paste of 200 characters or more, or three or more pastes of 50 characters or more, flags the sitting. Candidates are told on the start page that pastes are recorded.

4. Typing-cadence analysis

Humans type in bursts, with pauses for thought, backspaces, and re-edits. Text pasted from an LLM lands in a single event with no typing rhythm. We record the timing of keystrokes in answer fields (counts and intervals, never the keys themselves) and flag an answer when most of its text did not arrive by typing, when the rhythm is too regular to be a person over a long run, or when the sustained speed is faster than people type.

5. Human review on writing tasks

Every writing, editing, and proofreading submission is reviewed by human eyes. Our reviewers are editors themselves — they can spot the stylistic fingerprint of a language model in a few sentences, and they do.

6. Time-per-question anomalies

On the multiple-choice tests we record how long each question was on screen. Three or more correct answers given in under five seconds, faster than the question can be read, flags the sitting. The report shows the median time per question alongside.

What a flag means

We have not published detection-rate figures, because we have not run a study that would support them. What we can say is how the flags are used. Each of the four checks above appears on the report as flagged, not flagged, or not recorded (a sitting taken before the check existed), with a one-line explanation of what was measured.

No detector is perfect, and we're transparent about that. Flags are decision support — final hiring decisions stay with the hiring manager.

Want to see how it works?

Book a 15-minute walkthrough — we'll show you flagged and unflagged sample submissions side by side.

Book a demo →

Editors from these organizations have used our services since 1998

Reuters BBC Oxford University Press Penguin Random House Springer Microsoft Suncor Energy United Nations Fisher Investments IBM The Home Depot KODAK CHEVRON