Skip to main content

Whitepaper

Building an Editorial Benchmarking Standard

Published 22 September 2026

Every Editors Canada certification exam is marked twice. Two markers score each paper independently in a double-blind process, and a third marker assesses the work when the first two disagree about a pass.1 The markers work from "a very detailed answer key that contains a range of correct responses for each question."1 A marking analyst then reviews all of the marked exams "to ensure that the marking is consistent and reliable."2 Those three arrangements describe most of what an editorial benchmark needs, and an employer can build a smaller version of each one.

A benchmark starts as a written level of work

Editors Canada describes its Professional Editorial Standards as "statements about levels of performance that editors are expected to demonstrate."3 The document names employers among its users, who rely on it to know what to expect from the editors they hire, to develop job descriptions, and to create performance evaluation tools.3 Its definition of copy editing, for example, covers work that "corrects spelling, usage, grammar, and punctuation, and maintains consistency within the text."3

An in-house benchmark can take the same form. It is a short written statement of the work an editor at a given level should produce, tied to the documents that level handles. The Standards also observe that "sometimes the level of edit requested is not the level of edit required."3 A benchmark that names the level in writing removes that uncertainty for the candidate and for the people judging the work.

Reference answers turn the level into a key

A written level becomes usable once it is attached to real passages and model answers. The Editors Canada key lists a range of correct responses rather than a single one.1 Many editing problems have more than one acceptable solution, and a key that allows only one will mark a competent alternative as an error.

The testing profession's Standards for Educational and Psychological Testing address the same risk. When judges are expected to apply particular criteria in scoring, "it is important to ascertain whether they are, in fact, applying the appropriate criteria."4 A key with its accepted alternatives written down in advance is the record that makes that check possible. It also lets a later reviewer see why a candidate lost or kept a mark.

Agreement between markers is measured

Two editors reading the same test will sometimes disagree, and a benchmark is only as consistent as the people applying it. The statistical literature calls the extent to which raters "assign the same score to the same variable" interrater reliability.5 Percent agreement is the simplest figure to compute. Jacob Cohen criticized it in 1960 because it does not account for agreement that would occur by chance, and McHugh's review recommends reporting both percent agreement and the kappa statistic.5 An employer with two or more editors can run this check on a set of past tests before relying on the benchmark for hiring. The passages the markers scored differently show which parts of the written level, or which entries in the key, need more detail.

A cutoff and a percentile are different comparisons

Standard setting has two broad approaches. An absolute standard judges each candidate against a fixed external standard. A relative standard compares candidates with the others who took the same test, so the outcome "is dependent on the performance of the group."6 Norm-referenced methods carry a known drawback: "some examinees will always fail irrespective of their performance."6

The federal rules on employment testing point toward the absolute approach for a pass mark. The Uniform Guidelines on Employee Selection Procedures say that cutoff scores "should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force."7 Editors Canada sets its certification pass at approximately 80 percent and describes a pass as "an indication of excellence."1 That figure marks a professional certification level, which is a different threshold from the minimum a job requires on the first day.

Percentiles have a place when the comparison group suits the candidate. The International Test Commission's guidelines warn against conclusions drawn from norms "that are not relevant to the people being tested or are outdated."8 EditingTests.com reports a percentile only where it holds enough completed results for the same test at the same difficulty, and it returns no percentile for a smaller group. A candidate's answers and results are released to the employer who sent the invitation, and EditingTests.com does not sell personal data.

Where the benchmark stops

A benchmark makes judgments consistent. It does not by itself show that the test predicts performance in the job, which is a separate question of validity. The Guidelines add that evidence sufficient for a pass or fail screen "may be insufficient to support the use of the same procedure on a ranking basis."7 They also treat a selection rate for any race, sex, or ethnic group that falls below four-fifths of the highest group's rate as general evidence of adverse impact.7 EditingTests.com holds no demographic data about candidates, and its bias audit page explains what that means for an employer.

How a benchmark fits an employer's legal obligations is a matter for that employer's counsel. Human-scored work on EditingTests.com is currently reviewed by one reviewer, as the validity and research page states. The agreement check described above is therefore one an employer would run on its own markers.

Notes

  1. Editors Canada, "Frequently asked questions about Editors Canada Professional Certification," accessed September 22, 2026. https://editors.ca/professional-development/certification/certification-faq/ ↩
  2. Editors Canada, "Professional Certification exams," last modified September 2, 2025. https://editors.ca/professional-development/certification/exams/ ↩
  3. Editors Canada, Professional Editorial Standards 2024, introduction and standards A2 and A3. https://editors.ca/wp-content/uploads/2024/05/EditorsCanada_ProfessionalEditorialStandards_2024.pdf ↩
  4. American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, Standards for Educational and Psychological Testing, 2014 edition, chapter 1. https://www.testingstandards.net/open-access-files.html ↩
  5. Mary L. McHugh, "Interrater Reliability: The Kappa Statistic," Biochemia Medica 22, no. 3 (2012): 276-282. https://www.biochemia-medica.com/en/journal/22/3/10.11613/BM.2012.031 ↩
  6. Sanju George, M. Sayeed Haque, and Femi Oyebode, "Standard Setting: Comparison of Two Methods," BMC Medical Education 6 (2006): 46. https://pmc.ncbi.nlm.nih.gov/articles/PMC1578558/ ↩
  7. U.S. Equal Employment Opportunity Commission and others, Uniform Guidelines on Employee Selection Procedures, 29 CFR Part 1607, sections 1607.4 and 1607.5, CFR 2024 edition. https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XIV/part-1607 ↩
  8. International Test Commission, ITC Guidelines on Test Use, version 1.2, October 8, 2013, guideline 2.6.5. https://www.intestcom.org/files/guideline_test_use.pdf ↩

Editors from these organizations have used our services since 1998

Reuters BBC Oxford University Press Penguin Random House Springer Microsoft Suncor Energy United Nations Fisher Investments IBM The Home Depot KODAK CHEVRON