Work Sample Test: A Recruiter Design and Review Guide
August 12, 2026 · 8 min read
A work sample test asks a candidate to do a bounded piece of work that resembles an important part of the role. The result can give a recruiter and hiring manager a new evidence source: not another claim about capability, but an observable response to a shared task. That value disappears when the prompt is vague, the exercise measures work the team plans to teach later, or reviewers invent their standards after seeing the submissions.
The U.S. Office of Personnel Management guidance on work samples says these tests require candidates to perform tasks or work activities that mirror the tasks employees perform on the job. It also limits the method to competencies applicants are expected to possess when they enter the role. Those two ideas give recruiters a practical design boundary: test real work the candidate must be ready to perform upon entry, not a generic puzzle and not material the new hire would normally learn after joining.
This guide covers the operating workflow: choosing the evidence question, designing a realistic task, preparing the review guide, running a pilot, collecting independent ratings, and carrying the result into a human-owned decision.
Start with one evidence question
Do not begin by asking the hiring manager what assignment they want candidates to complete. Begin with the unresolved decision: which important competency is difficult to examine through a resume or interview alone, and what observable work would make that competency easier to discuss?
Use the role's current task record as the source. The job analysis guide for recruiters shows how to connect important tasks to entry competencies and acceptable evidence. A work sample should inherit that connection rather than create a new requirement halfway through the process.
Write a one-sentence evidence question. For example: “Can the candidate turn a set of support cases into a prioritized response plan and explain the tradeoffs?” This is narrower than “strategic thinking” and more useful than “complete a customer-success exercise.” It names the work product, the judgment inside it, and the explanation reviewers need to inspect.
Choose a task that resembles the role without copying it
A realistic task does not need every detail of the live job. It needs the essential action, constraints, information pattern, and output. Build a small scenario that preserves those features while removing organization-specific material that is unnecessary for evaluation.
For each element, ask whether it changes the evidence:
- Input: What brief, dataset, request, or problem would someone in the role receive?
- Action: What must the candidate prioritize, create, diagnose, or explain?
- Output: What artifact or live response will reviewers examine?
- Constraint: Which time, information, tool, or stakeholder limit is genuinely part of the work?
- Follow-up: What question lets the candidate explain a choice without changing the original task?
Remove decorative complexity. A large packet, obscure terminology, or several unrelated deliverables can make the exercise harder without making it more representative. If one task cannot examine every criterion, keep it focused and use the rest of the interview plan for the remaining evidence.
Decide whether the competency is required at entry
OPM distinguishes work samples from exercises that measure broader capabilities and advises using work samples when the tested competency is needed when the person enters the position. Apply that check before finalizing the prompt.
Ask three questions:
- Will the person perform this kind of task early in the role?
- Does the team expect the person to arrive with this competency?
- Would a weaker result change the next human review, or merely identify a normal development area?
If the answer to the second or third question is no, move the item out of the test. A work sample should not quietly turn a learnable preference into an entry gate. Record that boundary in the design note so later reviewers know what the exercise does and does not establish.
Write instructions another recruiter can administer
The candidate prompt should state the scenario, expected output, available materials, allowed tools, completion conditions, submission method, and what will happen next. Use the same core prompt and materials for everyone assessed for the same role version.
Separate candidate instructions from reviewer notes. Reviewers may need the competency map, common interpretations, rating anchors, and follow-up guidance. Candidates need a clear task and enough context to perform it. Mixing the two often produces a prompt full of evaluation language that encourages people to mimic the rubric instead of doing the work.
If candidates may use AI or another assistance tool, state the operating condition clearly and apply it consistently. Decide what reviewers are examining: the final artifact, the reasoning walkthrough, the use of supplied information, or a combination. Do not infer a tool-use rule after the submission arrives.
Build the review guide before the first submission
The OPM guidance on writing assessments connects work samples to job analysis, critical competencies, and standardized reviewing and scoring procedures. It also recommends a scoring rubric and reviewers with relevant expertise for writing samples. The operational lesson applies across work products: decide how evidence will be examined before anyone sees a candidate result.
Use a small set of dimensions tied directly to the evidence question. For the support-case example, the guide might include prioritization logic, use of supplied evidence, action clarity, and explanation of tradeoffs. Give each dimension observable anchors rather than adjectives:
- Direct evidence: the response uses relevant case details, states an order, and explains why.
- Partial evidence: the order is understandable but one material tradeoff or source detail is missing.
- Not established: the response gives a conclusion without enough visible reasoning to evaluate the criterion.
“Not established” is not a claim that the candidate lacks the competency. It says this artifact did not provide the required evidence. Keep that distinction in the review record and route any material uncertainty to the designated person.
Pilot the task with people who know the work
Run the exercise internally before using it with candidates. Ask at least two people familiar with the role to complete or walk through the task using only the supplied materials. Then have intended reviewers apply the guide independently.
The pilot should answer practical questions: Are the inputs sufficient? Does the task produce the intended evidence? Do reviewers interpret the anchors similarly? Does an irrelevant detail dominate the result? Can someone complete the exercise in the stated conditions? Revise the prompt or guide when the pilot exposes a design problem; do not solve ambiguity by telling live candidates different things one at a time.
Preserve the role version, prompt version, review-guide version, and pilot notes. If the job or exercise changes, name the new version so candidate records remain understandable.
Review independently, then discuss differences
Give trained reviewers the same artifact, rubric, and evidence-note fields. Each reviewer should record a rating, the part of the submission that supports it, and any unresolved question before seeing another reviewer's conclusion. OPM notes that work-sample scores can come from trained assessors observing behavior or from measured task outcomes; your review record should make clear which evidence source produced each rating.
After independent review, compare meaningful differences. Ask whether reviewers noticed different evidence, interpreted an anchor differently, or applied an unstated preference. The evidence-first interview debrief offers the same useful sequence: preserve individual observations, examine disagreement against the criterion, and let the designated human owner choose the next action.
Do not average away a design problem. Repeated disagreement may mean the anchor is unclear or the task produces several reasonable outputs. Fix the guide for the next consistent group and record how the current submissions will be handled.
Connect the result to the existing candidate record
A work sample is one stage in a broader evidence chain. Keep the original resume screen, interview notes, work-sample artifact, criterion ratings, and reviewer explanations distinguishable. The candidate scorecard guide shows how shared criteria can organize those sources without pretending they are interchangeable.
The handoff should state what the exercise examined, what the submission demonstrated, what remains unknown, and who owns the next decision. Do not let a single score replace the evidence notes. A hiring manager should be able to open the record and see why the result matters to the role.
Use AI as a checked preparation aid
A team may test AI for narrow preparation tasks such as organizing verified role tasks, checking for duplicate rubric language, formatting a prompt from approved inputs, or grouping submission excerpts under criterion IDs. Validate the proposed workflow on representative material and keep it only when reviewers can trace every output back to the approved role record or original submission.
Every output remains provisional. A recruiter or subject-matter reviewer should compare it with the approved role record and original submission, restore missing context, reject unsupported inferences, and apply the final rating guide. AI should not invent the tested competency, silently change an anchor, or decide whether a candidate advances.
Work sample test checklist
- Does the exercise answer one documented evidence question?
- Is the tested competency required when the person enters the role?
- Does the task mirror an important action, input, constraint, and output?
- Are unnecessary details removed?
- Will candidates receive comparable instructions and materials?
- Are tool-use conditions explicit?
- Was the review guide written before submissions arrived?
- Do rating anchors describe observable evidence?
- Has the task been piloted by people familiar with the work?
- Do reviewers record independent ratings and source notes?
- Are repeated differences used to improve the design?
- Does a named person own every next-step and hiring decision?
Make the exercise produce reviewable evidence
A useful work sample test is not the biggest assignment a hiring team can design. It is the smallest realistic task that reveals evidence about a critical entry competency under comparable conditions. Link it to the role, define the review guide first, pilot the workflow, preserve independent observations, and carry the result forward as one checked source. The test supports judgment; it does not replace the people responsible for the hiring decision.
When the resume screen is ready, see Resume Autopsy for a human-reviewed candidate comparison built around role criteria and supporting resume evidence.
Related reading
Free recruiting tools
Put the ideas to work — free, no signup
Check a job description for bias and clarity, build a sourcing search string, or size your screening cost — all in your browser.