A stack of blank task cards ends as one ochre human figure continues along an open path beneath the words HUMAN WORK REMAINS.

AI-generated editorial illustration.

September 30 is the last scheduled day to submit a Human Intelligence Task on Amazon Mechanical Turk. A worker who has already completed one must check a payment schedule. A requester has until October 30 to approve or reject a submitted task, and can still award a bonus through that date. Amazon says an unsubmitted task expires, while a submitted task left without a decision for 30 days is automatically approved. The company’s closure FAQ also leaves transaction history accessible until January 28, 2027.

The help page can close an account, but it cannot place a worker in another queue. It cannot tell a university whether a study run with a new participant pool will mean the same thing. It cannot supply the next AI lab with people able to recognize a difficult model error. At publication time, Amazon’s September 30 closure is a stated schedule, not an independently observed completion event.

For two decades, Mechanical Turk let a buyer put a small task in front of an anonymous remote workforce and pay for each answer. The same arrangement supported behavioral research, retail data cleanup, image labels and the human checks folded into AI workflows. Its exit leaves several markets that look similar on a purchasing screen but require different people. A researcher needs an appropriate sample. A model team may need an expert who can identify a subtle failure. A worker needs a rate that includes the time spent finding, qualifying for and disputing tasks.

Amazon has given each group a deadline. It has not offered an automatic path between the old market and the new ones.

September 30 puts unfinished tasks on a clock

Amazon’s first instruction to workers is to check a payment method and transfer frequency in the worker portal. Completed and approved tasks are to be paid on the existing disbursement schedule. The FAQ promises tax documents through ordinary channels. Requesters must check their own payment information so remaining prepaid balances can be refunded within 30 days. The final bill arrives in a later AWS billing cycle. The instructions are precise about accounts and quiet about work after the cutoff.

Submission ends September 30. Approval and bonuses can continue until October 30. Transaction history remains visible until January 28. A research lab with unapproved responses cannot treat the first date as the end of its obligations. A worker who finished a task still has to watch for the final transfer. The Amazon help page assigns the approval decision to the requester before the 30-day auto-approval rule takes over.

A project may also be attached to institutional rules. The University of Maine’s research compliance office warned on August 31 that a researcher moving to another crowdsourcing platform must modify the relevant institutional review board application before collecting data there. That is a concrete interruption, even if the new platform can recruit participants quickly. Consent language, the sampling frame and the location of participant data can change with the provider. A researcher cannot simply paste an old survey link into a new dashboard and call the resulting responses the same experiment.

For a longitudinal study, the disruption is sharper. The research question may depend on asking the same participants a second or third time. A different provider’s pool may contain people with similar demographics but no link to the first wave’s anonymous IDs. Recruitment can restart, yet the panel cannot necessarily be reconstructed. The buyer must decide whether to close the old cohort, seek consent and a permitted contact route, or amend the study design. None of those decisions appears in a headline comparison of platform fees. The Maine notice is one institution’s requirement; researchers elsewhere must follow their own review procedures rather than assume an automatic exemption.

The September deadline also separates work already submitted from tasks left open. Amazon says an unsubmitted HIT will expire, while a requester has 30 days after closure to decide on submitted HITs. The distinction matters to both sides. A task visible in a requester’s budget is not necessarily a payable submission. A worker’s finished submission may remain pending after the public marketplace stops operating. Keeping a local export of assignment IDs, submission times, approvals, rejections, bonuses and correspondence is a sensible handoff control; Amazon’s FAQ confirms access to transaction history through January 28 but does not promise that a later employer or panel will import that record.

There is a temptation to read the shutdown as proof that machines finished the work. Amazon has not published such a causal account. The notice says the company made the closure decision after an assessment. It does not give a current worker count, a revenue line, a platform margin, a task mix or a breakdown of tasks displaced by AI. The rest of the story has to start with what the marketplace actually sold.

A twenty-year market sold access to judgment by the task

Amazon’s own description of Mechanical Turk calls it an on-demand human workforce for jobs people could perform better than computers, such as recognizing objects in photographs. The service launched in 2005, when web services were turning infrastructure into something a developer could buy through an API. Mechanical Turk applied the same form to small pieces of human effort. The requester wrote a Human Intelligence Task, or HIT, specified an assignment and reward, and accepted or rejected the result. Workers chose from the available tasks rather than receiving a salary from the buyer.

The purchase looked simple on an invoice. Amazon’s published pricing says the buyer sets the worker reward and pays a 20% platform fee on rewards and bonuses. HITs with ten or more assignments attract another 20% fee on the reward. Premium and Masters qualifications add charges.

That schedule says little about hourly earnings. Ten cents for a minute of actual task work has one implied rate; ten cents after five minutes searching for a qualified task has another. Rejected work and unpaid screening change it again.

A study by Kotaro Hara and colleagues recorded 2,676 Mechanical Turk workers performing 3.8 million tasks. Its task-level estimate put the median hourly wage near $2, with only 4% of sampled workers above $7.25 an hour. This is a historical study, first posted in 2017, not a measurement of pay in September 2026. It included time lost searching, rejected tasks and other unpaid work that a per-HIT reward hides.

Another study of 100 workers and 40,903 tasks estimated that invisible labor used 33% of daily working time in its sample, pulling the measured median hourly wage to $2.83. Neither sample is a current census of Turkers. Both show why a headline task price is a poor account of earnings.

Requesters had reasons to like the model. A small team could collect many labels or survey responses without building a payroll department. An academic could turn a course schedule measured in months into a short fieldwork period. A machine-learning group could send ambiguous objects to people instead of asking an automated classifier to guess. In SageMaker’s public-workforce documentation, Amazon says the Mechanical Turk pool offered workers around the clock and typically the fastest turnaround for Ground Truth labeling or Augmented AI human review. Speed had a procurement value that did not appear in the reward per task.

It also widened who could commission a small study. A graduate student could recruit adults outside a campus subject pool without contracting with a traditional survey firm. That was a material advantage even when the sample was imperfect. In the later five-platform comparison, the unpaid undergraduate SONA pool took 14 weeks to collect, while paid online recruitment took days. The design and period of that study limit the comparison, but they make the buyer’s tradeoff clear: the market supplied time as well as answers. A successor that improves screening but cannot deliver the required population by a grant or product deadline is not automatically a better substitute.

An image label, an opinion in a research study and an expert’s critique of a model are different goods. The label may need a detailed guide and adjudication. The opinion needs recruitment, consent and a sample matched to the question. Expert critique may require calibration and a record of disagreement. Mechanical Turk could carry all three kinds of task; the common purchasing screen did not erase their differences.

The shutdown may send each type of task to a different place: a research panel, an internal review team, a specialist vendor or no human queue at all. Amazon’s announcement gives no basis to estimate the shares. It says nothing about the number of tasks that AI has replaced.

Ground Truth loses the crowd option, not every reviewer

SageMaker customers could also send work to Mechanical Turk without visiting its standalone requester site. Ground Truth routed labeling jobs, and Amazon Augmented AI routed human-review tasks, to the public workforce. The closure FAQ says that worker type will no longer be available when creating those jobs or workflows from September 30. It does not say every existing Ground Truth or Augmented AI review team closes that day.

AWS documents private workforces, in which a customer names its own employees or subject-matter experts, and vendor-managed workforces, where a third party supplies labelers under a separate subscription. The vendor agreement can specify price, schedule and refunds. In the private option, a buyer takes responsibility for selecting workers, managing identity and staffing the queue. In the vendor option, it must examine the provider’s workforce and contract. Neither is a one-click transfer of Mechanical Turk worker accounts, qualifications or historical task reputation.

There is an additional product boundary. AWS’s workforce security documentation says Ground Truth is no longer open to new customers, while existing customers can continue to use it and AWS plans security and availability improvements without new features. So the alternative-workforce description is useful for an existing Ground Truth user. It would mislead a new buyer if presented as a fresh general-purpose route into the same service. An existing customer can change a work team; a new customer must check the service’s present availability before designing around it.

The crowd route also carried a data rule. AWS instructs customers not to send confidential information, personal information or protected health information to the public Mechanical Turk workforce. Ground Truth and Augmented AI required a declaration that input data was free of personally identifiable information for that route. A private or vendor team can change the access design, but the buyer still has to decide who may view a record, where it is stored, how errors are corrected and how approvals are retained. A migration that solves the September deadline without reviewing data permissions can be an expensive shortcut.

An engineer trying to clear a labeling queue by Friday may still care most about throughput. A safety lead has a different question: who saw the examples, what instructions did they follow, and how was disagreement resolved? If those answers are missing, the resulting labels may be hard to defend as training or evaluation data. Mechanical Turk did not make that review impossible. It left much of the method to the buyer.

A blank task card forks into a survey participant path and a separate annotation review path, each ending with a human token.

AI-generated editorial illustration. The two paths represent different kinds of paid human work, not an Amazon workflow.

A cheap response can be the expensive one

Researchers have tried to count the replies they could actually use. Benjamin Douglas, Patrick Ewell and Markus Brauer compared five recruitment channels in a study published in 2023. Their adult US samples included 500 Mechanical Turk participants, 505 on CloudResearch, 496 on Prolific, 575 through Qualtrics and an undergraduate SONA sample. They paid $0.96 to participants on Mechanical Turk, CloudResearch and Prolific, then applied attention, identity, timing and response-quality criteria.

In that design, the cost per high-quality respondent was $4.36 on Mechanical Turk, $2.00 on CloudResearch and $1.90 on Prolific. Qualtrics was $8.17; SONA carried no participant payment but took a semester to collect. The researchers did not randomly assign people to platforms. The sample was US adults doing a particular survey, with a particular definition of high quality. Those figures do not establish 2026 prices, and a valid survey answer is not the same thing as a correct annotation or useful model critique. They do puncture the simple assumption that a lower price per completed task means a lower cost per accepted result.

A more recent Nature Human Behaviour comparison, published in June 2026, studied nine opt-in online samples with 13,053 respondents. Its authors measured response validity and representativeness rather than declaring one marketplace universally best. Open Mechanical Turk and one other panel scored in the lower-validity group; two CloudResearch Connect samples and two Prolific samples were in the highest-validity group. The paper found that simple early attention checks can improve response validity without substantially reducing representativeness, but it also described tradeoffs between demographic quotas and validity. A buyer needs both the ability to reject bad responses and a sample capable of answering the actual research question.

Valid answers and a representative pool deserve separate lines in a purchase order. If a company surveys workers about pay, a panel of people eager to take online tasks may answer attentively and still fail to describe the company’s workforce. A quota can make a sample look more like a population on age or region without fixing every difference in experience. The Nature authors measured both dimensions because improving one does not certify the other.

An employer using a panel to decide a staffing policy should name the population it intended to reach and the selection limits that remain. An AI team evaluating a model for ordinary users faces the same sampling question, even if it calls the task a preference test rather than a survey.

Platform comparisons are not a verdict on individual workers. A researcher’s task design, reward, eligibility rules, fraud checks and rejection policy affect the result. So does the match between a task and the participant pool. A Prolific research scientist’s September 29 retrospective argues that fraud and AI-assisted responses helped make the old open-crowd model less useful. That is an informed vendor perspective from a direct competitor, not Amazon’s explanation for its decision. The independent studies support a narrower conclusion: quality can vary sharply across samples and screening methods, making accepted-output cost more informative than the sticker price alone.

Consider a buyer with ten thousand short text classifications. A model may now produce a cheap first pass. Someone still has to decide which errors matter and whether a held-out test set can detect them. Other projects ask people to compare model answers or supply expert examples from the start. The first-pass price and the price of a defensible human assessment measure different work. Neither includes the cost of finding a bad label after it has entered a training or evaluation set.

That does not make every human evaluation inherently superior. A rushed or poorly instructed person can miss an error. A closed panel can exclude the population a study is meant to represent. A model can handle some routine classification at lower cost. The useful comparison is by task: accepted response or label, reviewer time, disagreement rate, correction time and the consequences of a mistake. The published cross-platform studies measured particular research responses. Model buyers need their own acceptance test for their own data.

Research panels and AI labs buy different people

CloudResearch Connect sells a research participant pool with a stated pay floor. Its project-cost guide says a requester sets participant payment, but the site refuses a project below a $7.50 hourly rate. It calls $10 an hour a baseline for faster collection of basic survey tasks. The platform fee is 25% for academic and nonprofit researchers and 40% for other buyers. Those are posted terms, not a measurement of worker earnings after screening and waiting.

Prolific’s MTurk migration page advertises an $8 hourly participant minimum, a 43% fee for commercial buyers and 33% for academics and nonprofits. The company says specialist roles can cost more. Its AI offering describes model evaluation, safety testing, preference data and post-training work, and recommends at least $12 an hour for ordinary participants while paying more for specialist skills. The skilled-participant program lists reasoning, fact-checking, image and video annotation and structured writing among assessable capabilities. These are vendor descriptions of its product and prices. They do not prove any customer’s model improved or that a former Turker can pass the new qualifications.

Take a simple budget line: one thousand tasks, ten minutes each, about 167 hours of completed work. At Prolific’s stated $8 minimum and 43% commercial fee, the set starts around $1,907 before specialist premiums, failed recruitment, rework or buyer-side review. At Connect’s $7.50 minimum and 40% fee for nonacademic buyers, the arithmetic starts around $1,750.

The two figures do not buy the same participants, quality or waiting time. They show why a finance team needs assumed minutes, target population, acceptance rate and platform fee on one line. A per-HIT price without those fields cannot settle the budget.

AI laboratories may need a separate procurement path altogether. Expert evaluation of a coding model, a legal argument or a hazardous instruction requires test design, calibration and sometimes restricted data access. A broad survey panel is useful for measuring ordinary user experience, but a representative population is not automatically qualified to grade a specialized failure. The opposite holds too: a group of expert annotators may be an excellent source of technical judgments and a poor substitute for a public-opinion sample. Buying ‘humans’ without specifying which question those humans are supposed to answer conceals the most important cost driver.

A worker with a strong Mechanical Turk approval history has no published guarantee that another platform will accept the account, preserve qualifications or offer similar hours. A higher posted minimum may improve the rate for tasks completed. It says little about admission, task availability, unpaid screening, review disputes or portability. Amazon’s FAQ covers old earnings and records. It does not announce a transfer program.

The World Bank’s study of online gig work warned that the sector can open access to work while leaving large gaps in protection. That is a general labor-market finding, not a headcount of people displaced by this shutdown.

The worker’s comparison is therefore different from the requester’s. A commercial buyer might see a 43% platform fee and look for a cheaper channel. A worker sees the reward, the probability of getting a task, the time spent proving eligibility and whether a rejection can be appealed. The older Mechanical Turk earnings studies put numbers on the unpaid time; they do not tell us how a former worker will fare on Connect or Prolific now. A new floor can improve a completed task and still leave total weekly income lower if the pool is closed, invitations are scarce or specialist work requires credentials the worker does not have. No public migration data yet measures those transitions.

The migration file must name the worker, the pay and the proof

A buyer moving a study or annotation queue needs more than a vendor comparison slide. The useful artifact is a migration acceptance file that can be checked after the first batch. It should describe the work, the people doing it, what counts as an accepted answer and how a worker can challenge a mistake. It should also preserve the old task and payment records while Amazon still exposes them.

Decision in the migration fileEvidence to retainAccountable owner
Task and intended useOriginal HIT instructions, consent or label guide, output schema and whether results train, test or only inform a modelResearch or data lead
Worker poolRequired geography, language, domain skills, recruitment method and any new eligibility checksResearch lead or vendor manager
CompensationExpected minutes, reward, fee, screening time, bonus and rejected-work appeal pathBuyer and vendor
QualityBlind test items, disagreement and adjudication rules, acceptance threshold and rework logMethod or model-evaluation owner
Data accessAllowed data classes, identity control, storage location, retention and subcontractorsPrivacy and security owner
ContinuitySubmission IDs, pending approvals, refunds, tax or billing records and the October 30 / January 28 datesRequester finance and worker-support owner

This is an editorial test, not a claim that Amazon or another platform supplies these fields in one export. It deliberately connects three budgets that are often separated: the buyer’s invoice, the worker’s actual time and the cost of validating the result. If a university changes participant recruitment, its institutional review board may need to approve the new path. If a SageMaker customer moves from the public pool to private or vendor workers, its security team needs to review access and data handling. If an AI lab switches from routine labels to specialist assessments, its model team has to define the expertise it is buying and the threshold for a useful judgment.

A university asking adults about workplace AI use needs a consent path, recruitment criteria and an account of how its sample differs from the population it wants to discuss. A model developer asking expert reviewers to find coding errors needs examples with known failure modes, adjudication when reviewers disagree and a record of the model version they saw.

The university may pay for broad reach; the developer may pay for scarce expertise. Calling both orders ‘human feedback’ hides the denominator. The useful price is cost per response or label the project can defend under its own method, with worker pay and correction time beside it.

No score in the published studies can complete that file on the buyer’s behalf. The 2023 comparison measures survey quality under one design. The 2026 nine-sample study measures response validity and representativeness across defined panels. Neither measures the error rate of a particular model’s safety labels or the fairness of a particular worker’s appeal. A pilot batch with held-out examples can test output quality, but the buyer should also observe time to approval, work that is rejected and how fast corrections reach the dataset. A contract that names only ‘accuracy’ can leave the error denominator and the people bearing rework costs undefined.

Nor can a worker solve portability alone. A downloaded task history is useful evidence of completed work, but another marketplace may use its own identity checks, geographic restrictions and qualification tests. A platform that advertises better vetting will have to decide how to recognize experienced people from the older marketplace without simply importing its fraud and quality problems. Those decisions will determine whether the market creates a route into better-paid expert work or merely changes the logo over a crowded queue.

October 30 remains on the payment calendar

A submitted HIT can outlive the marketplace that hosted it. Requesters still need to review work; bonuses can be awarded through October 30. Any submitted HIT not decided within 30 days will be auto-approved under Amazon’s stated policy. Workers need a functioning payment preference, and requesters need a correct refund account. Transaction history remains available only until January 28, 2027, according to the closure FAQ. These are more immediate than any forecast about AI labor.

After the last payment, the work will be harder to count. Some Mechanical Turk tasks may be automated; some will go to research panels, specialist evaluation providers or internal teams. Current public evidence cannot say how much will go to each, how many workers will be admitted, or whether effective pay and protection will improve. Amazon has documented the exit. Rival vendors have documented their offers. Researchers have shown that response quality and total cost can differ sharply from a task’s sticker price.

A worker looking at a final approved task and a lab looking at a replacement study have different next steps, but they share one requirement: the record must survive the platform. For the worker, it is the submission, approval and payment. For the lab, it is the population, instructions and evidence that an answer can be used. The last Mechanical Turk task closes only one part of that file.