Pymetrics and Harver: Game-Based Talent Assessment
On this page 11 sections

Pymetrics is now part of Harver. Its game-based assessments use short tasks to generate behavioral measures that may be compared with role profiles. The approach can add structured evidence to a hiring process, but it should not be described as neuroscience proving a candidate’s potential or as a bias-free replacement for human judgment.
The product and ownership record
Harver announced its acquisition of Pymetrics in August 2022. That ownership change matters because product support, contracting, and data processing now sit within a broader assessment platform.
Harver’s current gamified-assessments page says the assessment contains more than 12 games and takes about 25 minutes. It also reports a 98% completion rate and says its matching data draw on 100 million assessments. Those are vendor statements. Buyers should request the measurement definitions, population, study dates, and evidence for the job families they intend to assess.
What a game-based assessment can measure
A game records observable actions under defined rules. Those actions can be transformed into scores, and scores can be tested for relationships with job-relevant outcomes. The validity of the process depends on the chain between task, construct, score, job analysis, and decision threshold.
No interface removes the need to ask basic assessment questions:
- What construct is each game intended to measure?
- Is the score reliable for the relevant applicant population?
- Which job outcome was used for validation?
- Does the relationship hold outside the original employer or role?
- What happens when disability, device access, language, or unfamiliarity with games changes performance?
Calling a task engaging or neuroscience-based does not answer those questions.
Audit-AI’s contribution
Pymetrics released the Audit-AI repository as an open-source toolkit for measuring group differences in algorithmic outputs. Its implementation guidance discusses tests such as selection-rate comparisons.
This is useful for detecting a possible disparity. It does not establish that a model is fair, job-related, or legally compliant. Results change with group definitions, sample sizes, thresholds, missing data, and the outcome chosen for comparison. A clean aggregate result can also hide problems in a specific job, location, or stage.
The US Equal Employment Opportunity Commission’s AI and ADA resources make clear that employers need an accommodation process and can remain responsible for a vendor tool’s effects.
A defensible implementation
Start with a documented job analysis and a clear reason for adding the assessment. Run a pilot without allowing the score to make an automatic rejection. Compare completion, candidate feedback, subgroup outcomes, interview evidence, and later performance with the existing process. Review false negatives, not only average predictive results.
Candidates should receive plain-language notice about the assessment, the data collected, how it affects the decision, available accommodations, retention, and deletion. Recruiters need a review path for technical failures and unusual results.
Separate the interface from the measurement claim
Game-like presentation can make an assessment feel different from a conventional test, but the evidentiary questions are the same. A candidate action is captured, transformed into one or more features, combined into a score, and used in a decision. Every step can introduce error. The task may be unfamiliar, the feature may be unstable, the construct may be poorly defined, or the hiring threshold may discard useful information.
The buyer should request a measurement map for each deployed game:
| Link in the chain | Evidence to request | Failure to watch for |
|---|---|---|
| task to observed behavior | instructions, device requirements, test-retest data | the interface changes the behavior being measured |
| behavior to construct | construct definition and supporting studies | a broad trait label exceeds what the actions show |
| construct to job | job analysis and validity evidence | the trait is not important for this role |
| score to decision | threshold rationale and incremental value | a convenient cutoff creates avoidable false negatives |
This map prevents a common overclaim: that a behavioral task directly reveals a person’s durable potential. It produces a score under specific conditions. The employer must still show why that score is relevant to the work and how much weight it deserves.
Read vendor metrics correctly
Harver’s published completion and assessment-volume figures can indicate operational scale, but they do not establish validity. A 98% completion rate needs a denominator, population, period, and definition of completion. A large historical dataset can support research, yet a model trained on past outcomes may reproduce old selection choices or fail when the job changes.
Ask for evidence at the level used in production: job family, country, language, device, applicant group, model version, and pass rule. A correlation reported across multiple roles may not hold for a specific one. An average result can hide both a strong use case and a weak use case.
The federal Uniform Guidelines Q&A explains that written or oral assertions of validity are not a substitute for evidence when a procedure has adverse impact. It also notes that findings from one situation do not automatically transfer when jobs, work behaviors, criteria, or samples differ. That makes local monitoring part of responsible use, even when the vendor supplies a strong technical report.
Fairness review needs more than one ratio
Audit-AI can calculate useful group comparisons, but a production review should examine the full funnel. Compare completion, score distributions, selection rates, technical-failure rates, accommodation use, overrides, and later outcomes. Use confidence intervals or other uncertainty information where sample sizes are small. Do not conclude that a missing statistical signal proves equal experience.
The Society for Industrial and Organizational Psychology’s recommendations for AI-based employee selection are a useful independent reference for evaluating validation and use. They are professional guidance, not a legal safe-harbor or a certification of Pymetrics.
Investigate outliers rather than optimizing one aggregate number. A particular browser, translated instruction, or timed interaction may disproportionately affect a smaller group. Keep enough event detail to diagnose the problem while limiting collection and retention to what is necessary.
Accessibility and candidate recourse
A timed, visual, or motor interaction can measure disability-related interaction with the interface rather than a job-relevant capability. The accommodation path should therefore appear before the assessment, be easy to use, and not require a candidate to disclose more medical information than necessary. Alternative formats should measure the same job requirement as closely as possible.
When a session fails, recruiters need a documented reset and review process. A second attempt should not quietly change the norm group or decision rule. Candidates should be able to ask what stage the assessment affected, correct account or technical data, and reach a person when the result is contested.
Data governance after the acquisition
Because Pymetrics is part of Harver, buyers should verify the current contracting entity, subprocessors, data locations, support owner, and product-specific retention rules. Historical Pymetrics documentation may not describe the current service. Confirm whether raw game events, derived features, scores, recordings, and research datasets have different retention or deletion behavior.
Version control matters as much as contractual control. Record the assessment version, norm, role profile, threshold, and configuration used for every decision. If the vendor updates a model or content library, decide whether the change requires a new pilot, candidate notice, or validation review.
A pilot that can answer a real question
Start with a question such as: does the assessment add job-relevant evidence beyond the structured interview without creating unacceptable completion or subgroup differences? Predefine the sample, outcomes, review period, and escalation threshold. Run the score in shadow mode before using it to reject candidates.
At the end, report operational outcomes separately from predictive evidence. Candidate completion can improve while the score adds little information. Recruiter time can fall while false negatives rise. A responsible launch may approve some roles, revise others, and reject the product for jobs where the interface or evidence is a poor fit.
Bottom line
Pymetrics brought a distinctive game format and public bias-audit tooling into Harver’s assessment portfolio. That combination can support a more structured process. It does not remove the employer’s duties to validate the assessment, monitor outcomes, provide accommodations, and keep a person accountable for the decision.