# Mechanical Turk Reaches Its Final Day. AI Still Hires Humans

> Amazon set September 30 as Mechanical Turk's final task deadline. Workers, researchers and AI teams now face a different market for paid human judgment, with new costs and unresolved protections.

- Published: 2026-09-30
- Author: Gene Dai
- Canonical: [https://digidai.github.io/2026/09/30/mechanical-turk-shutdown-human-ai-work/](https://digidai.github.io/2026/09/30/mechanical-turk-shutdown-human-ai-work/)
- Topics: Artificial Intelligence, Crowdwork, Data Labor, Research, Future of Work, Deep Investigation

---

![A stack of blank task cards ends as one ochre human figure continues along an open path beneath the words HUMAN WORK REMAINS.](/images/articles/mechanical-turk-shutdown-human-ai-work/cover-v1.jpg)

_AI-generated editorial illustration._

September 30 is the last scheduled day to submit a Human Intelligence Task on Amazon Mechanical Turk. A worker who has
already completed one must check a payment schedule. A requester has until October 30 to approve or reject a submitted
task, and can still award a bonus through that date. Amazon says an unsubmitted task expires, while a submitted task
left without a decision for 30 days is automatically approved. The company's [closure FAQ](https://www.mturk.com/help)
also leaves transaction history accessible until January 28, 2027.

The help page can close an account, but it cannot place a worker in another queue. It cannot tell a university whether a
study run with a new participant pool will mean the same thing. It cannot supply the next AI lab with people able to
recognize a difficult model error. At publication time, Amazon's September 30 closure is a stated schedule, not an
independently observed completion event.

For two decades, Mechanical Turk let a buyer put a small task in front of an anonymous remote workforce and pay for each
answer. The same arrangement supported behavioral research, retail data cleanup, image labels and the human checks
folded into AI workflows. Its exit leaves several markets that look similar on a purchasing screen but require different
people. A researcher needs an appropriate sample. A model team may need an expert who can identify a subtle failure. A
worker needs a rate that includes the time spent finding, qualifying for and disputing tasks.

Amazon has given each group a deadline. It has not offered an automatic path between the old market and the new ones.

## September 30 puts unfinished tasks on a clock

Amazon's first instruction to workers is to check a payment method and transfer frequency in the worker portal.
Completed and approved tasks are to be paid on the existing disbursement schedule. The FAQ promises tax documents
through ordinary channels. Requesters must check their own payment information so remaining prepaid balances can be
refunded within 30 days. The final bill arrives in a later AWS billing cycle. The instructions are precise about
accounts and quiet about work after the cutoff.

Submission ends September 30. Approval and bonuses can continue until October 30. Transaction history remains visible
until January 28. A research lab with unapproved responses cannot treat the first date as the end of its obligations. A
worker who finished a task still has to watch for the final transfer. The [Amazon help page](https://www.mturk.com/help)
assigns the approval decision to the requester before the 30-day auto-approval rule takes over.

A project may also be attached to institutional rules. The
[University of Maine's research compliance office](https://umaine.edu/research-compliance/2026/08/31/amazon-mechanical-turk-mturk-to-shut-down-notice-to-human-subjects-researchers-august-2026/)
warned on August 31 that a researcher moving to another crowdsourcing platform must modify the relevant institutional
review board application before collecting data there. That is a concrete interruption, even if the new platform can
recruit participants quickly. Consent language, the sampling frame and the location of participant data can change with
the provider. A researcher cannot simply paste an old survey link into a new dashboard and call the resulting responses
the same experiment.

For a longitudinal study, the disruption is sharper. The research question may depend on asking the same participants a
second or third time. A different provider's pool may contain people with similar demographics but no link to the first
wave's anonymous IDs. Recruitment can restart, yet the panel cannot necessarily be reconstructed. The buyer must decide
whether to close the old cohort, seek consent and a permitted contact route, or amend the study design. None of those
decisions appears in a headline comparison of platform fees. The Maine notice is one institution's requirement;
researchers elsewhere must follow their own review procedures rather than assume an automatic exemption.

The September deadline also separates work already submitted from tasks left open. Amazon says an unsubmitted HIT will
expire, while a requester has 30 days after closure to decide on submitted HITs. The distinction matters to both sides.
A task visible in a requester's budget is not necessarily a payable submission. A worker's finished submission may
remain pending after the public marketplace stops operating. Keeping a local export of assignment IDs, submission times,
approvals, rejections, bonuses and correspondence is a sensible handoff control; Amazon's FAQ confirms access to
transaction history through January 28 but does not promise that a later employer or panel will import that record.

There is a temptation to read the shutdown as proof that machines finished the work. Amazon has not published such a
causal account. The notice says the company made the closure decision after an assessment. It does not give a current
worker count, a revenue line, a platform margin, a task mix or a breakdown of tasks displaced by AI. The rest of the
story has to start with what the marketplace actually sold.

## A twenty-year market sold access to judgment by the task

Amazon's own [description of Mechanical Turk](https://docs.aws.amazon.com/mturk/) calls it an on-demand human workforce
for jobs people could perform better than computers, such as recognizing objects in photographs. The service launched in
2005, when web services were turning infrastructure into something a developer could buy through an API. Mechanical Turk
applied the same form to small pieces of human effort. The requester wrote a Human Intelligence Task, or HIT, specified
an assignment and reward, and accepted or rejected the result. Workers chose from the available tasks rather than
receiving a salary from the buyer.

The purchase looked simple on an invoice. Amazon's [published pricing](https://www.mturk.com/pricing) says the buyer
sets the worker reward and pays a 20% platform fee on rewards and bonuses. HITs with ten or more assignments attract
another 20% fee on the reward. Premium and Masters qualifications add charges.

That schedule says little about hourly earnings. Ten cents for a minute of actual task work has one implied rate; ten
cents after five minutes searching for a qualified task has another. Rejected work and unpaid screening change it again.

A [study by Kotaro Hara and colleagues](https://arxiv.org/abs/1712.05796) recorded 2,676 Mechanical Turk workers
performing 3.8 million tasks. Its task-level estimate put the median hourly wage near
$2, with only 4% of sampled workers above $7.25 an hour. This is a historical study, first posted in 2017, not a
measurement of pay in September 2026. It included time lost searching, rejected tasks and other unpaid work that a
per-HIT reward hides.

Another [study of 100 workers and 40,903 tasks](https://arxiv.org/abs/2110.00169) estimated that invisible labor used
33% of daily working time in its sample, pulling the measured median hourly wage to $2.83. Neither sample is a current
census of Turkers. Both show why a headline task price is a poor account of earnings.

Requesters had reasons to like the model. A small team could collect many labels or survey responses without building a
payroll department. An academic could turn a course schedule measured in months into a short fieldwork period. A
machine-learning group could send ambiguous objects to people instead of asking an automated classifier to guess. In
[SageMaker's public-workforce documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/sms-workforce-management-public.html),
Amazon says the Mechanical Turk pool offered workers around the clock and typically the fastest turnaround for Ground
Truth labeling or Augmented AI human review. Speed had a procurement value that did not appear in the reward per task.

It also widened who could commission a small study. A graduate student could recruit adults outside a campus subject
pool without contracting with a traditional survey firm. That was a material advantage even when the sample was
imperfect. In the later five-platform comparison, the unpaid undergraduate SONA pool took 14 weeks to collect, while
paid online recruitment took days. The design and period of that study limit the comparison, but they make the buyer's
tradeoff clear: the market supplied time as well as answers. A successor that improves screening but cannot deliver the
required population by a grant or product deadline is not automatically a better substitute.

An image label, an opinion in a research study and an expert's critique of a model are different goods. The label may
need a detailed guide and adjudication. The opinion needs recruitment, consent and a sample matched to the question.
Expert critique may require calibration and a record of disagreement. Mechanical Turk could carry all three kinds of
task; the common purchasing screen did not erase their differences.

The shutdown may send each type of task to a different place: a research panel, an internal review team, a specialist
vendor or no human queue at all. Amazon's announcement gives no basis to estimate the shares. It says nothing about the
number of tasks that AI has replaced.

## Ground Truth loses the crowd option, not every reviewer

SageMaker customers could also send work to Mechanical Turk without visiting its standalone requester site. Ground Truth
routed labeling jobs, and Amazon Augmented AI routed human-review tasks, to the public workforce. The closure FAQ says
that worker type will no longer be available when creating those jobs or workflows from September 30. It does not say
every existing Ground Truth or Augmented AI review team closes that day.

AWS documents [private workforces](https://docs.aws.amazon.com/sagemaker/latest/dg/sms-workforce-private.html), in which
a customer names its own employees or subject-matter experts, and
[vendor-managed workforces](https://docs.aws.amazon.com/sagemaker/latest/dg/sms-workforce-management-vendor.html), where
a third party supplies labelers under a separate subscription. The vendor agreement can specify price, schedule and
refunds. In the private option, a buyer takes responsibility for selecting workers, managing identity and staffing the
queue. In the vendor option, it must examine the provider's workforce and contract. Neither is a one-click transfer of
Mechanical Turk worker accounts, qualifications or historical task reputation.

There is an additional product boundary. AWS's
[workforce security documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/sms-security-workforce-authentication.html)
says Ground Truth is no longer open to new customers, while existing customers can continue to use it and AWS plans
security and availability improvements without new features. So the alternative-workforce description is useful for an
existing Ground Truth user. It would mislead a new buyer if presented as a fresh general-purpose route into the same
service. An existing customer can change a work team; a new customer must check the service's present availability
before designing around it.

The crowd route also carried a data rule. AWS
[instructs customers](https://docs.aws.amazon.com/sagemaker/latest/dg/sms-workforce-management-public.html) not to send
confidential information, personal information or protected health information to the public Mechanical Turk workforce.
Ground Truth and Augmented AI required a declaration that input data was free of personally identifiable information for
that route. A private or vendor team can change the access design, but the buyer still has to decide who may view a
record, where it is stored, how errors are corrected and how approvals are retained. A migration that solves the
September deadline without reviewing data permissions can be an expensive shortcut.

An engineer trying to clear a labeling queue by Friday may still care most about throughput. A safety lead has a
different question: who saw the examples, what instructions did they follow, and how was disagreement resolved? If those
answers are missing, the resulting labels may be hard to defend as training or evaluation data. Mechanical Turk did not
make that review impossible. It left much of the method to the buyer.

![A blank task card forks into a survey participant path and a separate annotation review path, each ending with a human token.](/images/articles/mechanical-turk-shutdown-human-ai-work/interior-v1.jpg)

_AI-generated editorial illustration. The two paths represent different kinds of paid human work, not an Amazon
workflow._

## A cheap response can be the expensive one

Researchers have tried to count the replies they could actually use. Benjamin Douglas, Patrick Ewell and Markus Brauer
[compared five recruitment channels](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0279720) in a
study published in 2023. Their adult US samples included 500 Mechanical Turk participants, 505 on CloudResearch, 496 on
Prolific, 575 through Qualtrics and an undergraduate SONA sample. They paid $0.96 to participants on Mechanical Turk,
CloudResearch and Prolific, then applied attention, identity, timing and response-quality criteria.

In that design, the cost per high-quality respondent was $4.36 on Mechanical Turk, $2.00 on CloudResearch and
$1.90 on Prolific. Qualtrics was $8.17; SONA carried no participant payment but took a semester to collect. The
researchers did not randomly assign people to platforms. The sample was US adults doing a particular survey, with a
particular definition of high quality. Those figures do not establish 2026 prices, and a valid survey answer is not the
same thing as a correct annotation or useful model critique. They do puncture the simple assumption that a lower price
per completed task means a lower cost per accepted result.

A more recent [Nature Human Behaviour comparison](https://www.nature.com/articles/s41562-026-02438-z), published in June
2026, studied nine opt-in online samples with 13,053 respondents. Its authors measured response validity and
representativeness rather than declaring one marketplace universally best. Open Mechanical Turk and one other panel
scored in the lower-validity group; two CloudResearch Connect samples and two Prolific samples were in the
highest-validity group. The paper found that simple early attention checks can improve response validity without
substantially reducing representativeness, but it also described tradeoffs between demographic quotas and validity. A
buyer needs both the ability to reject bad responses and a sample capable of answering the actual research question.

Valid answers and a representative pool deserve separate lines in a purchase order. If a company surveys workers about
pay, a panel of people eager to take online tasks may answer attentively and still fail to describe the company's
workforce. A quota can make a sample look more like a population on age or region without fixing every difference in
experience. The Nature authors measured both dimensions because improving one does not certify the other.

An employer using a panel to decide a staffing policy should name the population it intended to reach and the selection
limits that remain. An AI team evaluating a model for ordinary users faces the same sampling question, even if it calls
the task a preference test rather than a survey.

Platform comparisons are not a verdict on individual workers. A researcher's task design, reward, eligibility rules,
fraud checks and rejection policy affect the result. So does the match between a task and the participant pool. A
Prolific research scientist's
[September 29 retrospective](https://www.prolific.com/resources/the-new-era-of-data-quality-in-online-research) argues
that fraud and AI-assisted responses helped make the old open-crowd model less useful. That is an informed vendor
perspective from a direct competitor, not Amazon's explanation for its decision. The independent studies support a
narrower conclusion: quality can vary sharply across samples and screening methods, making accepted-output cost more
informative than the sticker price alone.

Consider a buyer with ten thousand short text classifications. A model may now produce a cheap first pass. Someone still
has to decide which errors matter and whether a held-out test set can detect them. Other projects ask people to compare
model answers or supply expert examples from the start. The first-pass price and the price of a defensible human
assessment measure different work. Neither includes the cost of finding a bad label after it has entered a training or
evaluation set.

That does not make every human evaluation inherently superior. A rushed or poorly instructed person can miss an error. A
closed panel can exclude the population a study is meant to represent. A model can handle some routine classification at
lower cost. The useful comparison is by task: accepted response or label, reviewer time, disagreement rate, correction
time and the consequences of a mistake. The published cross-platform studies measured particular research responses.
Model buyers need their own acceptance test for their own data.

## Research panels and AI labs buy different people

CloudResearch Connect sells a research participant pool with a stated pay floor. Its
[project-cost guide](https://connect-researcher-help.cloudresearch.com/hc/en-us/articles/5046181555732-Project-Cost)
says a requester sets participant payment, but the site refuses a project below a $7.50 hourly rate. It calls $10 an
hour a baseline for faster collection of basic survey tasks. The platform fee is 25% for academic and nonprofit
researchers and 40% for other buyers. Those are posted terms, not a measurement of worker earnings after screening and
waiting.

[Prolific's MTurk migration page](https://www.prolific.com/prolific-vs-mturk) advertises an
$8 hourly participant minimum, a 43% fee for commercial buyers and 33% for academics and nonprofits. The company says specialist roles can cost more. Its [AI offering](https://www.prolific.com/ai) describes model evaluation, safety testing, preference data and post-training work, and recommends at least $12
an hour for ordinary participants while paying more for specialist skills. The
[skilled-participant program](https://researcher-help.prolific.com/en/articles/445229-participants-skilled-at-ai-tasks)
lists reasoning, fact-checking, image and video annotation and structured writing among assessable capabilities. These
are vendor descriptions of its product and prices. They do not prove any customer's model improved or that a former
Turker can pass the new qualifications.

Take a simple budget line: one thousand tasks, ten minutes each, about 167 hours of completed work. At Prolific's stated
$8 minimum and 43% commercial fee, the set starts around $1,907 before specialist premiums, failed recruitment, rework
or buyer-side review. At Connect's
$7.50 minimum and 40% fee for nonacademic buyers, the arithmetic starts around $1,750.

The two figures do not buy the same participants, quality or waiting time. They show why a finance team needs assumed
minutes, target population, acceptance rate and platform fee on one line. A per-HIT price without those fields cannot
settle the budget.

AI laboratories may need a separate procurement path altogether. Expert evaluation of a coding model, a legal argument
or a hazardous instruction requires test design, calibration and sometimes restricted data access. A broad survey panel
is useful for measuring ordinary user experience, but a representative population is not automatically qualified to
grade a specialized failure. The opposite holds too: a group of expert annotators may be an excellent source of
technical judgments and a poor substitute for a public-opinion sample. Buying 'humans' without specifying which question
those humans are supposed to answer conceals the most important cost driver.

A worker with a strong Mechanical Turk approval history has no published guarantee that another platform will accept the
account, preserve qualifications or offer similar hours. A higher posted minimum may improve the rate for tasks
completed. It says little about admission, task availability, unpaid screening, review disputes or portability. Amazon's
FAQ covers old earnings and records. It does not announce a transfer program.

The
[World Bank's study of online gig work](https://www.worldbank.org/en/news/press-release/2023/09/07/demand-for-online-gig-work-rapidly-rising-in-developing-countries)
warned that the sector can open access to work while leaving large gaps in protection. That is a general labor-market
finding, not a headcount of people displaced by this shutdown.

The worker's comparison is therefore different from the requester's. A commercial buyer might see a 43% platform fee and
look for a cheaper channel. A worker sees the reward, the probability of getting a task, the time spent proving
eligibility and whether a rejection can be appealed. The older Mechanical Turk earnings studies put numbers on the
unpaid time; they do not tell us how a former worker will fare on Connect or Prolific now. A new floor can improve a
completed task and still leave total weekly income lower if the pool is closed, invitations are scarce or specialist
work requires credentials the worker does not have. No public migration data yet measures those transitions.

## The migration file must name the worker, the pay and the proof

A buyer moving a study or annotation queue needs more than a vendor comparison slide. The useful artifact is a migration
acceptance file that can be checked after the first batch. It should describe the work, the people doing it, what counts
as an accepted answer and how a worker can challenge a mistake. It should also preserve the old task and payment records
while Amazon still exposes them.

| Decision in the migration file | Evidence to retain                                                                                                      | Accountable owner                          |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------------ |
| Task and intended use          | Original HIT instructions, consent or label guide, output schema and whether results train, test or only inform a model | Research or data lead                      |
| Worker pool                    | Required geography, language, domain skills, recruitment method and any new eligibility checks                          | Research lead or vendor manager            |
| Compensation                   | Expected minutes, reward, fee, screening time, bonus and rejected-work appeal path                                      | Buyer and vendor                           |
| Quality                        | Blind test items, disagreement and adjudication rules, acceptance threshold and rework log                              | Method or model-evaluation owner           |
| Data access                    | Allowed data classes, identity control, storage location, retention and subcontractors                                  | Privacy and security owner                 |
| Continuity                     | Submission IDs, pending approvals, refunds, tax or billing records and the October 30 / January 28 dates                | Requester finance and worker-support owner |

This is an editorial test, not a claim that Amazon or another platform supplies these fields in one export. It
deliberately connects three budgets that are often separated: the buyer's invoice, the worker's actual time and the cost
of validating the result. If a university changes participant recruitment, its institutional review board may need to
approve the new path. If a SageMaker customer moves from the public pool to private or vendor workers, its security team
needs to review access and data handling. If an AI lab switches from routine labels to specialist assessments, its model
team has to define the expertise it is buying and the threshold for a useful judgment.

A university asking adults about workplace AI use needs a consent path, recruitment criteria and an account of how its
sample differs from the population it wants to discuss. A model developer asking expert reviewers to find coding errors
needs examples with known failure modes, adjudication when reviewers disagree and a record of the model version they
saw.

The university may pay for broad reach; the developer may pay for scarce expertise. Calling both orders 'human feedback'
hides the denominator. The useful price is cost per response or label the project can defend under its own method, with
worker pay and correction time beside it.

No score in the published studies can complete that file on the buyer's behalf. The 2023 comparison measures survey
quality under one design. The 2026 nine-sample study measures response validity and representativeness across defined
panels. Neither measures the error rate of a particular model's safety labels or the fairness of a particular worker's
appeal. A pilot batch with held-out examples can test output quality, but the buyer should also observe time to
approval, work that is rejected and how fast corrections reach the dataset. A contract that names only 'accuracy' can
leave the error denominator and the people bearing rework costs undefined.

Nor can a worker solve portability alone. A downloaded task history is useful evidence of completed work, but another
marketplace may use its own identity checks, geographic restrictions and qualification tests. A platform that advertises
better vetting will have to decide how to recognize experienced people from the older marketplace without simply
importing its fraud and quality problems. Those decisions will determine whether the market creates a route into
better-paid expert work or merely changes the logo over a crowded queue.

## October 30 remains on the payment calendar

A submitted HIT can outlive the marketplace that hosted it. Requesters still need to review work; bonuses can be awarded
through October 30. Any submitted HIT not decided within 30 days will be auto-approved under Amazon's stated policy.
Workers need a functioning payment preference, and requesters need a correct refund account. Transaction history remains
available only until January 28, 2027, according to the [closure FAQ](https://www.mturk.com/help). These are more
immediate than any forecast about AI labor.

After the last payment, the work will be harder to count. Some Mechanical Turk tasks may be automated; some will go to
research panels, specialist evaluation providers or internal teams. Current public evidence cannot say how much will go
to each, how many workers will be admitted, or whether effective pay and protection will improve. Amazon has documented
the exit. Rival vendors have documented their offers. Researchers have shown that response quality and total cost can
differ sharply from a task's sticker price.

A worker looking at a final approved task and a lab looking at a replacement study have different next steps, but they
share one requirement: the record must survive the platform. For the worker, it is the submission, approval and payment.
For the lab, it is the population, instructions and evidence that an answer can be used. The last Mechanical Turk task
closes only one part of that file.

## Continue reading

- [AI Training Work Splits the Pay Band](https://digidai.github.io/2026/06/24/ai-training-work-pay-band/): Compare routine annotation with the specialist work and pay bands now sold to AI teams.
- [Three Measures of Outsourcing's AI Growth](https://digidai.github.io/2026/08/09/ai-outsourcing-revenue-workforce-split/): Follow the gap between AI services revenue and the people who still deliver or check the work.
- [Before an Hourly Shift Starts, Workers Price Gas and Childcare](https://digidai.github.io/2026/09/20/frontline-shift-cost-pay-schedule-ai/): See why a posted rate can miss the worker's actual time and costs.
- [Workers Fear Skill Erosion as CHROs Ask for More Judgment](https://digidai.github.io/2026/09/23/ai-skill-erosion-critical-thinking-work-redesign/): Trace what happens to human judgment when machines take more first-pass tasks.
