Criteria sells pre-employment assessments and employee-development tools. Its portfolio includes cognitive aptitude, personality, emotional-intelligence, skills, and risk assessments, as well as the Coach Bo development product. Each instrument should be evaluated for its stated construct and intended use; they are not interchangeable measures of overall talent.

What the CCAT measures

Criteria’s CCAT product page describes a 50-item, 15-minute cognitive aptitude test. The vendor says the assessment has been administered more than 10 million times and cites a 985-person norm sample. These details explain format and vendor reach. They do not by themselves show that a cutoff is valid for a particular role.

Cognitive measures may correlate with training and job performance, but their use can also produce group differences. Employers should start with job analysis, use the least restrictive cutoff supported by evidence, provide accommodations, and assess whether the test adds value beyond structured work samples and interviews.

Emotional-intelligence and game-based products

Criteria’s emotional-intelligence overview describes Emotify as measuring the ability to perceive, understand, and manage emotions through interactive tasks. That is narrower than claiming the test measures a candidate’s character, empathy, or culture fit.

Interactive presentation may improve the testing experience for some candidates. It can also introduce device, language, familiarity, and accessibility effects. A pilot should examine completion and score patterns by relevant group and testing environment.

Development is a separate decision context

Criteria presents Coach Bo as an AI-supported coaching product and its Develop platform as a tool for employee growth. Development tools operate in a different context from selection tests. Employees need to know whether coaching conversations or derived insights can be seen by managers, used in performance decisions, or retained after they leave.

An employer should not quietly reuse development data for hiring, promotion, or discipline. Purpose, access, retention, and consent need to be explicit.

Evidence buyers should request

For every assessment, ask for the technical manual, reliability evidence, validation design, norm population, accessibility information, adverse-impact analyses, and recommended retest interval. Confirm which evidence applies to the current product version and the target job.

Track results after launch. Useful measures include completion, accommodation requests, subgroup selection rates, reviewer agreement, stage conversion, later performance, and candidate complaints. Investigate false negatives and unusual results instead of relying only on average correlations.

Treat each product as a separate measurement system

Criteria’s catalog spans constructs and decision contexts. A cognitive aptitude test, an emotional-intelligence task, a personality inventory, and an employee coaching tool do not become one coherent “talent score” simply because they share a platform. Combining them without a job-based rationale can make the result harder to explain and validate.

Product categoryNarrow question it may informEvidence the employer needsMisuse to avoid
Cognitive aptitudeperformance on reasoning problems under timed conditionsjob relevance, reliability, norm fit, cutoff rationaletreating speed as universal potential
Skills or work sampleability to perform defined job tasksrepresentative content and scoring consistencyusing generic questions for a specialized role
Emotional intelligenceperformance on the stated emotion-processing tasksconstruct evidence and relevance to critical workinferring empathy, integrity, or culture fit
Personalityself-reported or otherwise measured tendenciesscale reliability and role-linked criterion evidencediagnosing candidates or enforcing an ideal type
Development coachingprompts and feedback for employee growthprivacy, role separation, employee understandingsilently reusing coaching data for selection

The measurement plan should say why an instrument is needed and how it changes a decision. If the hiring team cannot describe that use without broad labels such as “quality” or “fit,” it is not ready to set a cutoff.

How to evaluate the CCAT

The CCAT’s published format makes speed a material part of the observed score. That may be defensible for some jobs, but the employer should not assume that a 15-minute reasoning test represents every form of learning or complex problem solving. Identify which critical tasks require similar reasoning under similar constraints, and which do not.

Norms answer how a score compares with a reference group. They do not determine the right threshold for a job. Ask when the norm sample was collected, how it was composed, whether the current form is comparable, and how score precision varies across the scale. Then examine the decision consequence: how many applicants near the cutoff would be rejected, and what other evidence could correct a false negative?

The federal agencies’ Uniform Guidelines Q&A defines validation as demonstrating job relatedness and notes that a procedure developed elsewhere still has to be justified for the employer’s particular use. It also distinguishes content, criterion-related, and construct evidence. This is an independent standard for the buyer’s review, not a certification of Criteria or any competing assessment.

Emotional intelligence requires careful language

Emotify’s interactive tasks may reduce some limitations of asking candidates to describe themselves, but they do not make interpretation unlimited. Buyers should inspect the precise construct definitions, the scoring model, reliability, norms, language adaptations, and the relationship between each scale and important job behavior.

Do not convert an assessment output into a clinical or moral judgment. A result should not be described as proving that a candidate lacks empathy or cannot lead. If the role requires recognizing emotion in a defined customer interaction, use that work context and compare the assessment with a structured simulation or interview. Divergent evidence should trigger review, not automatic exclusion.

Adverse impact and alternatives

Cognitive and other assessments can produce different selection rates across groups. Monitoring should occur at the configured job and stage, not only across the whole company. Record the applicant pool, completions, valid scores, selection rule, progression, and accommodations. Small samples need uncertainty and should not be declared safe merely because a threshold was not crossed.

The Society for Industrial and Organizational Psychology’s recommendations for AI-based assessments provide an independent professional reference for evidence, implementation, and monitoring. Employers should also examine less adverse alternatives that measure the same requirement, including different weights, thresholds, formats, or structured work samples. A high predictive relationship does not end that inquiry.

Accessibility, security, and candidate experience

Timed, visual, game-like, or language-heavy tests can create barriers unrelated to the job. Test the candidate journey with keyboards, screen readers, magnification, mobile devices, slower connections, and approved extra time. Put the accommodation route before the test and ensure requesting it does not silently change candidate status.

The EEOC’s AI and ADA resources explain why employer responsibility remains even when software comes from a vendor. Record technical failures and manual resets, and give candidates a person who can review a disputed result.

Security review should distinguish response data, derived scores, reports, development conversations, and model inputs. Ask about subprocessors, encryption, access logs, retention, backups, deletion, research reuse, incident response, and exports. The data needed to administer a test may not be needed indefinitely.

Keep development data out of hidden selection

Coach Bo and the Develop platform may be valuable when employees understand the purpose and can use feedback for growth. Trust breaks if private reflections, prompt history, or inferred weaknesses later appear in performance, promotion, or termination decisions without a clear policy.

Set access by role and purpose. Managers may need an agreed development goal, while HR administrators may need usage and safety information; neither necessarily needs a verbatim coaching conversation. Define whether the employer, Criteria, or an underlying model provider can use the content for product improvement. Provide a deletion and correction path and state what happens when employment ends.

A responsible pilot decision

Run the assessment in shadow mode for representative roles before allowing it to reject applicants. Predefine the job evidence, sample, outcome, subgroup checks, accessibility tests, and stopping rules. Compare the configured assessment with a baseline such as a structured interview or work sample, not with intuition alone.

At the decision meeting, separate four conclusions: the software operated reliably, candidates completed it, scores showed job-relevant evidence, and the hiring outcome improved. Passing one does not prove the others. Approve use at the level the evidence supports, which may mean a limited role, an advisory score, or no deployment at all.

Bottom line

Criteria offers several structured assessment formats and a distinct development product. The portfolio can help replace informal impressions with consistent evidence, but the label “scientific” is not a blanket guarantee. Defensible use requires job relevance, current validation, accessible alternatives, careful data governance, and a human decision owner.