Research
Validity & research
How our assessments are built and scored, and — just as important for anyone evaluating them — which validation studies we have carried out and which we have not.
What we test
Seven assessments cover the editorial skill domains our clients hire for:
- Grammar Test — mechanical accuracy, multiple choice against an answer key.
- Editing Test — structural and stylistic editing judgement, marked as tracked changes.
- Proofreading Test — detection and correction of planted errors in a passage.
- Writing Test — an original written response, scored by a human against a published rubric.
- Industry Vocabulary Test — sector terminology, multiple choice against an answer key.
- MS Word Test — application knowledge plus a document-formatting task.
- Full Editing Assessment — the editing, proofreading and grammar tests in one timed session.
Items are written against a competency model drawn from published professional standards for editing, and each item is tagged with the construct it is meant to measure so that coverage can be checked.
How scoring works
- Multiple-choice answers are marked against a stored answer key. The score is the number correct out of the number of questions served, not out of the number the candidate chose to answer.
- Editing and proofreading submissions are reconstructed from the candidate's tracked changes and compared against a master-corrected version of the same passage. The score reflects which corrections were actually made, not how many edits the candidate happened to type.
- The Writing Test is scored by a person against a five-dimension, 100-point rubric — Content 15, Structure 20, Style 15, Grammar 25, Mechanics 25 — which is published in full in our AI notice.
- Open-response work — writing, editing, proofreading and the MS Word document task — is reviewed by a person before the result is released to the employer.
- The pass mark is a setting you choose. Nothing about a candidate other than the work they submitted enters a score.
Benchmarking
Scores are reported as a percentile against other candidates who took the same test at the same difficulty, where we hold enough completed results for that comparison to mean anything. A cohort below our minimum size returns no percentile rather than an unreliable one. The pool grows as assessments are completed; we do not publish a headline figure for its size, because any number we printed here would be out of date the day after.
What we have not done
We would rather you heard this from us than found it out in a procurement review:
- No criterion-validity study. We have not measured the correlation between test score and any on-the-job outcome. Clients can record hiring and retention outcomes against a candidate, and some do, but we have not analysed that data and we publish no correlation coefficient.
- No published reliability statistics. We do not currently compute internal consistency (Cronbach's alpha) or test-retest reliability for any of the seven assessments.
- No inter-rater agreement programme. Human-scored work is reviewed by one reviewer. We do not double-score a sample or compute agreement between reviewers.
- No independent bias audit. We hold no demographic data about candidates, so we cannot compute selection rates by protected class. Our full position, and what it means if you hire in New York City, is on our bias auditing page.
None of this stops the assessments being useful evidence about a candidate's editing skill. It does mean that if you need documented validation evidence for a regulated selection process, you do not yet have it from us, and you should plan accordingly.
Questions about methodology
If you are evaluating this platform and need detail we have not published, ask — [email protected]. We would rather answer a hard question than have you assume an answer.