# Meta Rewrote Performance Reviews After Tokenmaxxing

> Meta's September guidance shifted attention from AI usage to contribution as companies link AI skills to promotion without shared standards. A five-level rubric separates tool activity from judgment, verification, and team impact.

- Published: 2026-09-22
- Author: Gene Dai
- Canonical: [https://digidai.github.io/2026/09/22/meta-tokenmaxxing-performance-review-guidance/](https://digidai.github.io/2026/09/22/meta-tokenmaxxing-performance-review-guidance/)
- Topics: Artificial Intelligence, Performance Management, AI Skills, Meta, Career Development, Future of Work, Deep Investigation

---

![Two colleagues review an evidence board while a pile of blank AI tokens sits beside the headline AI REVIEWS BEYOND TOKEN COUNTS.](/images/articles/meta-tokenmaxxing-performance-review-guidance/cover-v1.jpg)

_AI-generated editorial illustration._

On September 2, Meta put two sharply different accounts of AI at work into public view.

Its engineering team described an internal agent that turns expert corrections into tested updates. The system pauses at
defined checkpoints, escalates ambiguous cases and leaves authority with a human specialist. Meta said the six-week
project reduced some assessments from days to minutes, though it did not publish the domain, sample size or error rate
needed to compare that claim outside the company.

That same day,
[WIRED reported](https://www.wired.com/story/meta-pushes-its-new-ai-agent-on-employees-but-eases-off-on-tokenmaxxing/)
that Meta had changed its performance-review guidance. References to employees' AI usage and an "AI Native" designation
gave way to language focused on contribution. Engineers were told that AI adoption dashboards and token counts would not
be used to evaluate impact.

Meta spokesperson Tracy Clayton told WIRED that contributions had always been the basis of evaluation and that the
labels were never used for performance ratings. Some employees described the effect differently: the revised language,
they said, reduced pressure to use AI where it did not fit.

Together, the two accounts expose a management problem. AI activity is extremely visible. Prompts, tokens and generated
files arrive as ready-made numbers. Evidence of contribution sits farther downstream. It requires a reviewer to
reconstruct task choice, errors, uncertainty, outcomes and the work shared with colleagues.

Companies are already pulling those behaviors into promotion and pay decisions. Their standards have not caught up. The
resulting gap affects more than HR. It changes which work employees perform, which mistakes they surface, how managers
allocate credit and whether an AI budget buys useful capacity or an internal contest.

This article uses public company material, reporting, surveys, research and current regulation as of September 22, 2026.
It does not establish how any individual Meta employee was rated. The proposed rubric is an editorial framework, not a
validated test or legal compliance standard.

## Meta changed the review language in September

Meta's revised wording arrived after months in which token usage became a visible symbol of AI adoption across the
technology sector. A token is a small unit processed by a language model. It belongs in a cost calculation. Inside a
performance conversation, however, it can easily become a proxy for effort, enthusiasm or modernity.

[Associated Press reporting in July](https://apnews.com/article/ai-token-openai-anthropic-corporate-31bb80ac1cd7862d05f6397177d826b1)
described "tokenmaxxing" as an office fad that began to lose force when AI bills rose without a matching increase in
useful output. Meta had an internal competition that rewarded usage. An employee-created leaderboard assigned labels
such as "Token Legend" before it was taken down, according to WIRED and reporting it cited.

A usage dashboard reaches leadership before the finance team can calculate the value of an improved process. Managers
can see whether employees opened a tool long before customer outcomes, error rates or delivery cycles become measurable.
During a rollout, activity answers an immediate board question: is anybody using what we bought?

Activity can be useful for adoption support. A team with zero usage may lack access, training or an applicable use case.
A sudden drop can reveal a broken integration. Cost and security teams need to know where calls originate. Those uses do
not turn a count into evidence that one employee deserves a stronger performance rating.

There is a stronger case for a temporary usage target than critics sometimes allow. A company can buy thousands of
licenses, announce voluntary training and discover months later that employees never found time to practice. Waiting for
a revenue result before diagnosing adoption would leave leaders blind. A completion target for a bounded exercise,
combined with role-level access data, can show where training or product support has failed.

The boundary should be explicit. A rollout diagnostic answers whether people had access and tried the tool. A career
criterion answers whether an employee performed work at the next level. The first may use aggregate sessions or course
completion during a defined period. The second needs work evidence. Keeping the two measures separate preserves the
useful signal without turning practice volume into permanent rank.

That separation also protects employees who discover that a tool is a poor fit. A credible adoption program needs room
for a documented refusal, followed by a safer alternative. Otherwise the only acceptable result of an experiment is
continued use, and the exercise stops being an experiment.

Meta's own September engineering article supplies a more demanding model. Its
[organizational second-brain system](https://engineering.fb.com/2026/09/02/ml-applications/organizational-second-brain-ai-learns-from-experts/)
records expert feedback, proposes minimal changes, reruns the scenario that exposed the problem, tests for regressions
and sends a reviewed change back to a human expert. The claimed benefit comes from a controlled loop, not from the
volume of model calls.

That loop creates several kinds of evidence. A reviewer can inspect the original question, the proposed correction, the
targeted replay, the regression result and the final approval. The process also recognizes an answer that should be
escalated. A person who improves this system could use those artifacts to explain the contribution without presenting a
single token count.

The public account stops short of a benchmark. Meta reports that specialists found the outputs useful almost all the
time and that individual assessment time fell from days to minutes. It does not disclose the underlying counts,
independent validation or the cost of building and maintaining the system. The case supports a design principle, not a
general productivity claim.

Meta is still asking employees to build and use advanced AI systems. The revised guidance removes a simple consumption
measure from the account of individual contribution. Employees and managers must now supply the missing evidence.

## Token counts made a bad career signal

Promotion metrics give employees a target. Token counts invite gaming because the counted unit sits far from the result
a company usually wants.

An employee can ask a model for ten versions of a memo instead of one. A developer can send a large context window when
a smaller request would do. A team can automate work that nobody needs. Each action raises visible consumption. None
proves that the final work became more accurate, faster, safer or more valuable.

Vincent Gusdorf of Moody's Ratings told AP that companies became more cautious as bills accumulated. That financial
pressure exposes one flaw in treating usage as a career signal. The employee who spends fewer tokens because the task is
well scoped can look less active than the employee who sends repeated, poorly bounded requests. A cost-saving behavior
may score worse on the adoption chart.

Access also varies. Engineers with generous model allowances can generate a large trace. A payroll specialist handling
sensitive data may face stricter controls. A salesperson traveling between customer meetings may have fewer suitable
tasks than a researcher working in documents all day. A blanket usage target converts job design and access policy into
an apparent difference in motivation.

Time away from work creates a sharper denominator problem. In July, 26 current and former Meta employees filed a
[federal complaint](https://storage.courtlistener.com/recap/gov.uscourts.cand.474171/gov.uscourts.cand.474171.1.0.pdf)
alleging that AI-token usage and other activity measures contributed to layoff selection and disadvantaged people who
had taken protected leave or received disability accommodations. The complaint describes allegations, not judicial
findings. Meta denied the allegations. A
[July 17 order](https://docs.justia.com/cases/federal/district-courts/california/candce/3%3A2026cv07122/474171/25)
denied temporary emergency relief and left the underlying claims unresolved.

The case shows why the denominator belongs in any review design. A raw count can reflect eligible working time, tool
access, job family and project assignment before it reflects skill. Normalizing by days worked would address only one
part of that problem. It would not tell a reviewer whether the work needed AI or whether the output survived scrutiny.

Duolingo encountered the same issue from another direction. The company tested making AI use part of performance
reviews, then backed away. CEO Luis von Ahn said employees should be judged on doing their jobs well and that the
company had pushed the tool into cases where it did not fit, according to
[an April account of his comments](https://www.techradar.com/ai-platforms-assistants/were-trying-to-push-something-that-in-some-cases-did-not-fit-duolingos-ceo-changes-course-on-ai-at-work).

AI can still contribute to performance by improving a process, finding a pattern or expanding service coverage. The
record has to follow the work past the prompt box. A review should distinguish appropriate abstention from resistance,
and disciplined use from expensive theater.

## Promotion rules are moving faster than standards

Two 2026 surveys capture employers at different points in this transition.

[HiBob surveyed 1,200 AI decision-makers](https://www.hibob.com/research/ai-skills-2026-report/) across six market
groups between February 3 and March 12. Sixty-seven percent said their organizations linked AI skills to promotion
criteria, and 50% tied those skills to performance ratings. Yet only 36% considered direct managers highly prepared to
build AI capability on their teams.

Respondents ranked proactive output review and documentation of workflow decisions among the most important everyday
behaviors, at 52% each. Those findings point away from raw usage. They also come from a vendor survey of managers and
leaders personally involved in AI decisions at organizations with 50 to 5,000 employees. The sample is useful for
understanding active buyers. It cannot establish that two-thirds of all employers have formal AI promotion rules.

[The Conference Board reported](https://www.conference-board.org/press/corporate-america-hasnt-moved-beyond-early-AI-adoption-yet)
a different result from more than 250 HR leaders. Fifty-six percent said AI fluency played little or no role in
advancement. Sixty percent placed their organizations in early experimentation, while 11% reported more advanced
integration. Among workers in the broader study, 52% thought stronger AI skills would affect their promotion prospects
to at least a moderate degree.

The surveys use different samples and questions, so their percentages should not be blended. Their tension still
matters. Employees have reason to believe AI skill will affect their careers. Many employers have started attaching it
to talent decisions. A large share of HR leaders say it has little practical weight, while managers remain uncertain
about how to assess it.

Pay data adds pressure without resolving the definition. PwC's
[2026 Global AI Jobs Barometer](https://www.pwc.com/gx/en/news-room/press-releases/2026/pwc-2026-ai-jobs-barometer.html)
analyzed more than one billion job advertisements in 27 countries and territories. It reported a 62% average wage
premium for postings requiring specific AI skills, up from 57% in the prior report. Jobs requiring those skills grew 69%
against 9% for the wider job market.

That is a hiring-market comparison, not a promise that an existing employee receives a 62% raise after learning an AI
tool. The category includes scarce technical capabilities and varies widely by industry. Its value here is directional:
companies are paying for some AI-linked skills while their internal ladders still struggle to describe the behavior that
merits advancement.

PwC also examined 2.4 million U.S. entry-level postings. Highly AI-exposed roles were seven times more likely to ask for
skills usually associated with senior work, including judgment and leadership. Openings for that group grew 35% from
2019, while other entry-level openings fell 10%. Those figures describe job advertisements, not the actual performance
or promotion rate of early-career employees. They still raise an urgent ladder question. If junior workers face senior
expectations sooner, managers need a way to observe how judgment develops rather than rewarding polished output alone.

Pete Brown, PwC's global workforce leader, framed the problem in developmental terms: routine work once acted as an
apprenticeship, while AI-exposed roles now demand judgment and leadership earlier. His interpretation comes from the
firm that produced the analysis. It is not a measured account of how any one employer trains junior staff. It does name
the missing input in a promotion rubric. Employees need repeated chances to make and explain decisions before the
organization can rate their judgment.

Another broad survey shows what is at stake in that transition. Mercer's
[Global Talent Trends 2026](https://www.mercer.com/about/newsroom/mercer-s-global-talent-trends-2026-report/), based on
nearly 12,000 executives, HR leaders, investors and employees surveyed in late 2025, found that 65% of executives
expected 11% to 30% of their workforce to be redeployed or reskilled because of AI over the next two years. Fifty-three
percent of employees worried that they lacked skills for future work. These are expectations and concerns, not observed
redeployment. They show how much career movement may depend on definitions that are still being written.

Robin Erickson, head of human capital research at The Conference Board, described large AI investment arriving before
measurable workplace impact. That lag explains why leaders reach for usage. It also makes promotion criteria risky. An
employee can be graded on an input while the company is still unable to show what the input changed.

## Judgment leaves a different evidence trail

AI-assisted work contains several decisions before a document, code change or customer response appears. The employee
decides whether the task belongs with a model, what material it may see, how the request is bounded and when a human
should take over. After the output arrives, someone decides what to check and whether the work is fit to ship.

Each decision can leave evidence that a usage meter misses. A task log can show why AI was selected. A source record can
show what supported the output. A correction can show that the reviewer detected an error. A rejected draft can show
that the employee stopped unsafe or low-quality work instead of maximizing throughput.

![A worker selects a task, an abstract AI workbench produces a draft, and a reviewer checks and corrects it before placing the result in a shared team tray.](/images/articles/meta-tokenmaxxing-performance-review-guidance/interior-v1.jpg)

_AI-generated editorial illustration. The sequence separates tool activity from task choice, verification and reusable
team work._

Research on human review makes the correction step important. In a 2026 experiment published by Harvard Data Science
Review, [2,784 participants checked AI-extracted values](https://hdsr.mitpress.mit.edu/pub/nrcn4h7d/release/2) from
corporate emissions reports. Participants corrected fewer errors when flagging a problem also required typing the right
value. People with more favorable attitudes toward automation accepted more incorrect suggestions. A bonus for accuracy
did not materially change performance in that task.

The experiment involved crowdworkers and data extraction, not corporate promotion reviews. It does not tell a manager
how to rate an engineer or marketer. It does show that review quality depends on workflow design and the cost of
correction. Telling an employee to "check the AI" is weak evidence if the interface makes rejection harder than
acceptance and the performance system rewards speed.

A February IZA discussion paper by Rainer Michael Rilke and Dirk Sliwka adds another boundary. Their
[experiments with language models as evaluators](https://www.iza.org/index.php/de/publications/dp/18371/when-algorithms-rate-performance-do-large-language-models-replicate-human-evaluation-biases)
found that models reproduced familiar rating patterns when standards were subjective and individuals were evaluated
separately. With noisy but objective performance signals, the models produced more dispersed and accurate assessments
than human raters in the experimental setting.

The experimental result supports stronger evidence. It does not support automatic grading. A model can organize
documented signals or test a rating against a rubric. It cannot repair a job whose outcome was never defined. When the
input rewards visible volume, the evaluator processes visible volume efficiently.

Useful evidence also includes effects on other people. A reusable workflow can reduce repeated setup for a team. A clear
failure note can prevent colleagues from trusting the same bad output. A training session can move capability beyond the
person who built the first prompt. These contributions often matter more at higher levels, where promotion depends on
multiplying the work of others.

The employee receiving AI-assisted work has a perspective that usage data cannot capture. A support specialist may
inherit a larger queue because an upstream tool generated more cases. A junior engineer may receive a patch without the
reasoning needed to learn from it. A subject expert may spend the afternoon correcting outputs while the workflow owner
claims the saved morning. Their records belong beside the creator's time estimate.

This is also where credit becomes a management decision. The person who built a workflow may deserve recognition for
task design. Reviewers deserve recognition for detecting errors and preserving quality. Colleagues who document a
failure contribute even when the experiment is stopped. A rubric that records only the person closest to the tool will
systematically hide the labor that made the result usable.

The record must keep costs and harms in view. Useful output can still expose confidential data, create unequal access or
send correction work to a less visible colleague. A manager needs to see who performed the review and who absorbed the
exceptions. Otherwise the employee who generates the most work receives credit while another person quietly makes it
safe.

## A five-level rubric for AI-assisted work

This five-level rubric is a proposed operating tool for review design, not a universal competency standard. A
customer-support role, a laboratory role and a software role require different evidence. Regulated or safety-sensitive
work may set a higher minimum before any output can be used.

The levels describe observable responsibility rather than tool fluency in the abstract. A reviewer should rate the
strongest level supported by repeated work, not by a self-description or a single demonstration.

| Level                     | Observable behavior                                                                                                    | Evidence for a review packet                                                                                       | Common false signal                                           |
| ------------------------- | ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------- |
| 0. Appropriate abstention | Recognizes when the task, data or policy makes AI use unsuitable and chooses a safer method                            | Recorded reason, policy reference, manual fallback and result                                                      | Low usage interpreted as low capability                       |
| 1. Assisted execution     | Uses an approved tool for a bounded task and checks basic correctness before delivery                                  | Before-and-after sample, disclosed tool use, source check and corrected output                                     | Prompt count or polished prose treated as skill               |
| 2. Reliable workflow      | Defines inputs, acceptance criteria, review steps, escalation and a fallback for repeated work                         | Task specification, error log, time and cost baseline, rejection examples and outcome measure                      | Gross time saved without correction or rework                 |
| 3. Workflow design        | Improves the division of work between people and AI, tests failure cases and changes the process when evidence is weak | Versioned procedure, evaluation set, subgroup or edge-case review, incident record and owner sign-off              | A reusable prompt presented without maintenance cost          |
| 4. Team leverage          | Helps other people perform the workflow, improves shared standards and shows durable benefit beyond personal output    | Adoption by named teams, training evidence, quality trend, access plan, saved capacity destination and review date | Personal visibility or tool evangelism counted as team impact |

Level 0 belongs in the framework because restraint can be skilled behavior. An employee handling protected data or an
ambiguous employment decision should not lose credit for refusing an unsuitable tool. The reviewer can ask whether the
person identified the constraint and found a workable path, rather than rewarding inactivity by default.

Level 1 captures competent personal use. A before-and-after example can show what the tool contributed and what the
employee corrected. This level should be common. It does not yet prove that the workflow works reliably or that other
people can use it.

Level 2 introduces an operating boundary. The employee defines acceptable output, measures the whole task and keeps a
fallback. Cost belongs here. So do errors and rejected attempts. A workflow that saves ten drafting minutes and creates
twenty minutes of review has not delivered positive capacity merely because generation was fast.

Level 3 fits roles responsible for systems and processes. Evaluation sets, edge cases and incident records make the work
inspectable. The employee earns credit for redesigning a weak workflow or stopping it when the evidence changes. This is
where judgment becomes more visible than usage.

Level 4 requires a result beyond personal productivity. The team can repeat the process, people receive access and
training, exceptions have owners and the released time has a destination. A manager should be able to name the work that
changed. If the benefit disappears when one enthusiastic employee leaves, the contribution may still be valuable, but it
is not yet durable team leverage.

The same employee can sit at different levels for different tasks. A senior designer may build a Level 3 research
workflow and remain at Level 1 for image generation. A finance analyst may show Level 4 impact in monthly close work
while appropriately avoiding AI for a restricted transaction. Averaging those examples into one generic "AI native"
score would throw away useful information.

Promotion decisions should also retain the ordinary job standard. AI-assisted work has to meet the same or a higher bar
for customer value, accuracy, reliability and professional responsibility. Tool skill cannot compensate for weak role
performance. It can change how that performance is achieved and how much responsibility the employee can carry.

## Managers inherit the calibration burden

A rubric cannot remove managerial judgment. It can make disagreements inspectable.

One reviewer may value a workflow because it saves drafting time. Another may discount it because a separate team
performs the verification. A third may worry that only employees with premium tools can produce the evidence. These are
calibration questions about cost ownership, credit and access. They should be discussed before ratings are finalized,
not hidden inside a single adoption score.

Managers also need time to observe the work. HiBob's finding that only 36% of surveyed decision-makers viewed direct
managers as highly prepared is important because those managers are expected to judge both job performance and a new
technical practice. Adding a competency without adding review capacity turns the rubric into another form field.

Organizations can reduce that load with a small evidence packet. Each promoted example should name the task, prior
baseline, tool and data boundary, acceptance test, reviewer, failure or correction, measured outcome and reuse. Finance
can challenge savings. Security can review data handling. A peer can test whether the workflow transfers. The manager
still owns the rating, but no longer has to reconstruct the work from a token dashboard and a polished self-review.

A practical calibration meeting can divide the questions instead of pretending the manager has every answer.

| Calibration question                          | Primary reviewer                           | Evidence to bring                                                                   |
| --------------------------------------------- | ------------------------------------------ | ----------------------------------------------------------------------------------- |
| Did the work improve the job outcome?         | Role manager and customer or process owner | Baseline, accepted result, quality measure and affected volume                      |
| Did the workflow save usable capacity?        | Finance or operations partner              | Full-cycle time, model cost, review cost, rework and destination of released time   |
| Was the data and tool use appropriate?        | Security, privacy or risk owner            | Approved system, data classification, access record, exception and incident history |
| Can another person repeat it?                 | Peer reviewer or receiving team            | Versioned instructions, test cases, failure notes, training and maintenance owner   |
| Does the evidence support the proposed level? | Manager calibration group                  | Job rubric, comparable opportunity set, counterexample and employee response        |

These assignments identify questions, not vetoes. A small company may have one person covering several roles. A large
company may already route them through existing finance, security and people processes. The important feature is visible
ownership. When every issue returns to the line manager, the quickest narrative usually wins.

Access differences need a separate calibration pass. Teams should compare employees within realistic opportunity sets,
including approved tools, eligible tasks, working time and support. They should also examine who received the invisible
review labor. A promotion system that rewards generation while ignoring correction will push careful employees toward
silence.

Regulation raises the stakes when AI moves from an employee's tool into the evaluator. The EU AI Act lists systems used
to make promotion decisions or monitor and evaluate worker performance among employment uses classified as high risk in
[Annex III](https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng). Under the current
[European Commission timeline](https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-high-risk-systems), those
high-risk rules apply from December 2, 2027. The Act calls for human oversight and, for covered workplace systems,
information to affected workers and their representatives. Other employment, privacy and discrimination rules can apply
before that date.

The classification depends on intended use and facts, so a company should obtain its own legal analysis. The operational
lesson is narrower. An AI assistant that helps an employee draft a self-review is different from a system that scores
employees for promotion. Moving from support to decision changes the evidence, oversight and challenge path a company
needs.

Employees need to know that path before ratings arrive. They should be able to correct a missing project, explain a
protected absence, identify review work assigned to them and challenge a tool-generated inference. Human oversight is
weak when the human sees only the final score.

## One review packet can settle the argument

Consider a hypothetical promotion packet, not a reported Meta employee or meeting. A product operations employee has
used an AI assistant fewer times than several peers. Her packet shows why.

One high-volume task failed its accuracy test, so she stopped it and retained the manual process. A second workflow cut
the time required to classify support issues, but only after she added a rejection option and weekly sampling. She
documented two errors, changed the instructions and trained three colleagues. The team's response time improved, while
the finance line includes inference and review cost. A separate column names the analyst who handles exceptions.

The manager can dispute the baseline, the sample or the allocation of credit. A peer can rerun the test. The employee
can answer with work rather than an identity label.

That packet will never be as easy to rank as a token counter. It is far more useful in a promotion meeting. Under Meta's
September guidance, the decisive evidence should resemble what its own engineering system preserves: the contribution,
the correction that improved it and an accountable approval.

## Continue reading

- [Judgment Gets a $100 Million Bonus Pool at EY](https://digidai.github.io/2026/09/05/ey-human-skills-bonus-pool/): See how one employer attached a large bonus pool to judgment, leadership and other human skills.
- [Junior Roles Lost the Work That Taught Judgment](https://digidai.github.io/2026/06/23/junior-roles-training-ground/): Follow the career-ladder problem created when AI removes the routine work that once trained judgment.
- [AI Skill Premiums Put Pay Bands on Trial](https://digidai.github.io/2026/06/19/ai-skill-premiums-pay-bands/): Compare advertised AI skill premiums with the harder task of changing internal pay bands.
- [After Hiring, HR AI Moves Into Performance Management](https://digidai.github.io/2026/04/21/next-hr-ai-fight-performance-management-post-hire-decisions/): Review the controls needed when AI enters performance, promotion and compensation decisions.
