Where AI Products Turn Into Jobs
At 7 p.m. Pacific time on August 24, two public labor snapshots froze four companies in voice AI and four in legal AI. ElevenLabs mapped to 622 current profiles and 219 open jobs. Harvey mapped to 867 and 335. Twenty-one hours earlier, a separate inference-serving table had put Anyscale at 489 current profiles while Baseten, with 189, carried the largest visible job book at 166.
Placed beside one another, those figures invite a fast story about who is hiring hardest. They cannot support it. The pages count public profiles and visible job postings associated with company names. They do not count hires, departures, accepted offers, payroll employees, contractors, or work completed. Some company names collide with unrelated organizations. A job can disappear because it was filled, paused, moved to another system, or simply missed by the mapping.
A better comparison begins where each AI product meets expensive human work.
Forward deployment reaches that point in a customer’s operating environment. Voice reaches it during a live exchange where latency, pronunciation, safety, and escalation are audible. Legal software reaches it when a professional must rely on, correct, and take responsibility for an answer. Inference serving reaches it when a model has to perform under a production load at an acceptable cost.
I am a co-founder and the chief product officer of Metix AI, which produced the four reports used here. This is a founder’s analysis of vendor-produced research, not an independent review. The counts are mapped floors drawn from public evidence, not official company headcount. That conflict gives me more reason to state what the data cannot prove before arguing for what it can do.
Read together, the maps offer a practical way to design a role. Find the place where the product promise can fail and the work that absorbs that failure. That evidence should determine the title, target companies, interview exercise, and budget owner.
Four reports stop at four operating boundaries
All four pages observe current public profiles and visible open jobs, but they classify the evidence differently. One is organized around forward-deployed title phrases. Another separates voice-related titles from broader role categories. Legal roles and lawyer titles receive separate treatment. The inference page compares company-level stocks.
| Metix report | Mapped current profiles | Mapped open jobs | Boundary exposed by the comparison | Claim the table cannot support |
|---|---|---|---|---|
| FDE Employed vs JD Map 2026 | 229 across the five-name FDE-title union | 199 across the five-name open-title union | Customer deployment work can spread across several titles | The table is not an official count of all forward-deployed engineers |
| Voice & Speech AI Talent Map 2026 | 764 across four company names | 243 | A voice product joins research and audio engineering to revenue and service work | The current-to-jobs pair is not a hiring or growth rate |
| Legal AI Talent Map 2026 | 2,237 across four company names | 670 | Domain labor appears in role classification, product work, delivery, and sales | A legal-role count and a lawyer-title count cannot be added into a workforce composition |
| Inference-Serving Talent Map 2026 | 1,508 across five company names | 280 | Production infrastructure creates a job book that ranks differently from the visible bench | A small visible job book does not establish a freeze, contraction, or weak demand |
Those sums are arithmetic over the rows Metix selected. They do not describe market size. Salesforce dominates the FDE title table, while Harvey supplies the largest legal current-profile row and ElevenLabs supplies the largest voice row. Those companies differ in age, product scope, naming conventions, geography, distribution, and data visibility. Ranking the totals would flatten the differences that make the maps useful.
Timing introduces another break. A current-profile count is a stock of visible people associated with a company at collection time. A visible job count is a stock of postings. Neither is the flow of people hired during a period. The U.S. Bureau of Labor Statistics makes the same distinction in JOLTS: job openings are counted on the last business day of a month, while hires cover additions to payroll throughout the month. Metix is not JOLTS, but the stock-versus-flow discipline still applies.
A ratio can therefore move without any new hiring. A company may remove duplicate postings, consolidate locations, open one evergreen requisition for several seats, or leave filled roles visible. Employees may update profiles late or use titles that the query misses. For several names in these reports, unrelated profiles can enter the result. A precise quotient would add confidence without fixing the denominator.
There is a reasonable objection to using these maps at all: public profile and posting data is too noisy for workforce planning. That objection wins if a team wants a census, a growth rate, or a vendor ranking. It is less persuasive when the task is narrower. A recruiting leader often needs to know which work families to inspect before running a verified search. A product leader may need to see that two companies with similar products expose different role vocabularies. A finance partner may want to ask why a proposed “AI engineer” belongs in services, product, research, or infrastructure.
Job seekers experience the data gap from the other side. Someone who keeps a profile private, writes in another language, takes a career break, or describes customer work without an AI keyword becomes harder to see. That is a visibility problem, not evidence of weaker ability. A market map should widen a recruiter’s search logic before any person enters an assessment. It should never become the assessment.
Visible limits make the reports usable for those questions. Each page declares its company set, as-of time, numbers, and collision warnings. Test the pattern against official career pages, role descriptions, recruiter intake notes, interviews, and internal operating data. Public evidence supplies a hypothesis. A hiring decision still requires a second source and a named owner.
Forward deployment hides inside unstable titles
The FDE employed-versus-job-description map is the clearest warning against treating a title as a stable occupation. Across Salesforce, Stripe, Palantir, Anthropic, and Tomoro, the employed FDE-title union is 229. The exact phrase “Forward Deployed Engineer” produces 163. The open-job title union is 199.
Salesforce accounts for 223 of the broader employed-title union and 123 of the open-title union. Palantir, the company most associated with forward deployment, maps to one employed profile under the report’s exact company and title logic. Its broader open-title union is 69, yet the exact FDE job count is one. “Forward Deployed Software Engineer,” infrastructure variants, reliability roles, and other forms change the result.
Metix explicitly warns that its Palantir profile coverage is incomplete. An exact company-name search for any current title returns only 56 profiles, and Palantir Technologies variants may sit outside the matched company name. That single match cannot be translated into “Palantir has almost no forward-deployed engineers.” It describes what one set of public strings captured.
Palantir’s own current Forward Deployed Software Engineer posting explains the work better than the acronym. The company describes engineers embedded with customers who learn hard operational problems, work with business-critical data, and build solutions with Palantir technology. The job sits where product engineering, data work, implementation, and customer trust overlap.
That boundary can live under many titles. An enterprise software company might use solutions architect, deployment strategist, applied AI engineer, field engineer, resident engineer, professional services engineer, or technical account lead. A startup may give the same work to a founding engineer or product manager because the org chart has not caught up. A consultancy may separate discovery, build, change management, and support across several people.
Searching only for “FDE” will miss candidates who have already done the work. Expanding to every “forward deployed” string can pull in roles with very different technical depth. For hiring, use an evidence pattern. Did the person enter a customer environment and translate an operating problem? Did they build or configure a production solution, handle constraints that the core product did not solve, and leave behind a supportable handoff?
That evidence changes an interview because a coding exercise alone tests too little, while a polished customer story can hide who wrote the software and who carried the escalation. Use one deployment where the available data was incomplete, access was constrained, or the workflow changed after users touched it. The candidate should separate reusable product from customer-specific work, name who supported the system at 2 a.m., and explain how the team knew it was safe to step away.
This also changes who gets through the screen. A solutions architect who wrote production code and forced three recurring customer exceptions into a platform release may fit the work better than an FDE whose role ended at a proof of concept. The reverse can be true when a solutions title covered mostly pre-sales demos. Recruiters need a short evidence screen that asks about the environment, artifact, owner, and handoff before relying on either label.
Funding shapes the role. Product-funded forward deployment should feed recurring problems back into the roadmap. Services-funded work may optimize for project delivery and billable utilization. Sales-funded work may concentrate before a contract and thin out after signature. Customer-success funding may emphasize adoption and renewal. A title will not reveal which economic logic governs the role.
A company can easily hire an impressive FDE and then remove the conditions that made the person effective. If access to product engineers takes weeks, customer exceptions become bespoke code. If sales controls priorities, the engineer may inherit commitments that the platform cannot support. If support owns every post-launch issue, the feedback loop closes too late. Role design has to specify authority, handoff, and the destination for repeated customer problems.
Put the public map on the screen during the first intake meeting. Even the category’s defining employer can look almost absent under the wrong title and company match. A brief for “ten Palantir-style FDEs” is not ready until someone defines the customer problem, technical scope, travel expectation, production ownership, and exit condition.
Voice turns latency into an organization problem
The voice and speech talent map counts 622 current profiles and 219 open jobs for ElevenLabs, 93 and 24 for Cartesia, 48 and zero for Hume AI, and one and zero for Sesame AI. The four rows sum to 764 current profiles and 243 visible jobs.
Its title table adds a second view. ElevenLabs maps to 13 research titles, 15 audio-engineering titles, and 152 go-to-market titles. Hume AI maps to 13 research titles, one audio-engineering title, and four GTM titles. Cartesia has five research titles and six GTM titles. These categories do not partition the current-profile totals, and their absence does not prove a function has no employees.
A role table broadens the view again. ElevenLabs maps to 17 research roles, 110 engineering and technical roles, 21 sales roles, 43 marketing roles, and 31 customer-success roles. The distinction between a title table and a role table matters. “Research engineer” can match both a research concept and a technical function. A GTM title is not identical to a sales classification. Summing across the two tables would count classification choices as people.
Commercial voice AI requires much more than speech researchers. ElevenLabs’ current careers page displays openings across engineering and product, research, growth, operations, and revenue. Its live total will not match the Metix snapshot because the time, source, grouping, and deduplication differ. Inspect the method instead of choosing whichever number looks newer.
Voice has an unusually unforgiving failure surface. Text can wait behind a spinner. A live caller hears the delay. They notice when the system interrupts, mispronounces a name, changes language poorly, misses emotion, or fails to transfer to a person. The product is judged across a chain of components rather than one model score.
ElevenLabs’ latency documentation separates model inference from time to first audio. Its Flash models can produce about 75 milliseconds of model inference latency for representative short inputs, excluding network and application overhead. A complete voice-agent path may include speech recognition, endpoint detection, an LLM, tool calls, text-to-speech generation, network travel, buffering, and playback. Each stage spends part of the listener’s patience.
That chain creates distinct jobs. Speech researchers improve generation and recognition, while audio engineers deal with codecs, streaming, signal quality, and playback. Distributed-systems engineers manage regional capacity and tail latency. Product engineers connect tools to business logic; safety specialists test abuse and impersonation; linguists and local teams inspect pronunciation, accent, and cultural fit. Solutions staff integrate telephony and customer systems before revenue and customer-success teams turn a technically working call into a purchased, monitored service.
Some of that labor may sit with the customer or a systems integrator. A bank deploying a voice agent can retain responsibility for identity checks, call recording, escalation, and complaint handling. A contact-center vendor may operate the telephony layer. The model provider may supply an API and reference architecture. Counting one vendor’s staff will never show the complete workforce behind the caller’s experience.
That distribution can be sensible. A smaller buyer does not need to reproduce a voice lab, a telephony carrier, and a round-the-clock safety operation inside one payroll. It can buy those layers. Procurement still has to name who hears failed transfers, who can stop a bad call flow, how quickly a person takes over, and which conversation evidence reaches the supplier. Outsourcing moves the boundary across a contract; it does not make the listener’s bad call disappear.
Keep the zero rows intact. Hume AI’s zero mapped open jobs and Sesame AI’s zero are real values in the Metix table, not blanks to smooth away. They do not reveal whether either company is hiring through referrals, an unobserved system, a parent organization, or no channel at all. They also say nothing about contractors or roles filled just before the snapshot.
For a hiring plan, the market boundary starts with a latency and quality budget. Which delay matters: model inference, time to first audio, or full task completion? Which languages and channels must work? What happens when a guardrail blocks a response or a caller asks for a human? Who reviews conversation logs, and which errors trigger product work rather than prompt changes?
An “audio ML engineer” brief can be correct for model quality and useless for a telephony outage. A backend engineer can keep an API available while failing to recognize that a 500-millisecond buffering change ruins turn-taking. A strong sales team can win an enterprise account before the company has enough deployment and support capacity. The map’s research-to-GTM spread helps a leader see that product maturity is an organizational balance, not a single scarce title.
Metrics should preserve that balance. Research quality, p50 and p95 time to first audio, interruption rate, tool completion, escalation rate, reviewer workload, support incidents, and renewal answer different questions. A hiring decision should state which measure the role can move and which neighboring team absorbs the exceptions.
Legal software still ends with professional judgment
The legal AI talent map contains the largest current-profile sum of the four reports: 2,237 across Harvey, EvenUp, Luminance, and Legora, with 670 mapped open jobs. Harvey maps to 867 current profiles and 335 jobs. EvenUp maps to 683 and 59, Luminance 405 and 27, and Legora 282 and 249.
Those totals do not show one uniform legal software market. EvenUp focuses on personal-injury case work, while the other companies span different legal research, drafting, contract, due-diligence, and workflow surfaces. Company-name collisions affect Harvey and Luminance. Geography, business model, customers, and career-page architecture differ.
Classification exposes a more useful tension. EvenUp maps to 195 people in the legal role category but only three with a lawyer title. Harvey maps to 47 legal roles and 15 lawyer titles. Legora maps to 38 and six. Luminance maps to 15 and four. A company can employ substantial domain labor without putting “lawyer” in the visible title.
Product work also reaches far outside the legal category. The report’s role table maps Harvey to 150 engineering and technical roles, 173 sales roles, 24 marketing roles, and 69 customer-success roles. EvenUp maps to 106 engineering and technical, 120 sales, 20 marketing, and 29 customer-success roles. These are separate mapped categories rather than a complete org chart, but they show where to look for the work that surrounds professional judgment.
Legal responsibility does not require a lawyer to rewrite every product answer. Someone still has to select authoritative sources, protect confidential information, understand jurisdiction, supervise tool use, check citations, communicate limits, and decide when an output can enter client work.
ABA Formal Opinion 512 places those duties on lawyers who use generative AI. It covers competence, confidentiality, client communication, supervision, candor, meritorious claims, and reasonable fees. A vendor can reduce the work involved in research or review, but the lawyer remains accountable for professional conduct.
That accountability creates jobs at several distances from the model. Legal subject-matter experts build evaluation sets and inspect errors, while knowledge engineers structure sources and retrieval. Product managers translate practice into workflows. Engineers implement permissions, citations, document handling, and audit trails. Implementation and customer teams then adapt the product to a firm’s knowledge, security, review process, pricing, staffing, and client communication.
Review work also has a career structure. A junior lawyer who once learned by reading the first draft may now receive a polished answer with a hidden source error. A senior lawyer can catch it, but that review time competes with client work and mentoring. Product evaluation cannot be designed only around final-answer accuracy. It needs to show who sees the error, how much context the reviewer gets, and whether less experienced professionals still encounter the reasoning that builds judgment.
My article on owned legal models and frontier APIs examined one technical choice inside that system. The talent map adds a wider point. Even a vendor that controls more model weights still needs legal judgment, product engineering, deployment, revenue, and support. Model ownership can relocate work. It does not remove the operating boundary.
Thomson Reuters’ 2026 report for law-firm leaders shows why that boundary now reaches pricing and talent. Its analysis draws on 736 law-firm professionals across 46 countries, plus corporate legal respondents. Seventy-one percent of in-house legal professionals expected outside firms to change commercial models as AI use increased, while 28% of law firms reported making pricing changes.
That survey is vendor-sponsored and does not measure the four companies in the Metix table. It does identify a real buyer conflict. If AI compresses time on a task, a firm must decide how to price the result, who checks it, how junior lawyers learn, and whether the client shares in the efficiency. A product team cannot solve those choices with model quality alone.
Legora’s 282 mapped current profiles and 249 open jobs form the closest same-order pair in the table. That does not mean the company plans to almost double. A single posting can represent more than one seat, remain open for months, or target several locations. Current profiles and jobs can cover different geographies and naming rules. The number is useful as a prompt to inspect official roles and expansion plans, not as a forecast.
A legal AI hiring brief should specify the consequence of error. Contract review, litigation research, personal-injury demand preparation, and firm-knowledge search carry different sources, deadlines, review paths, and client harms. The team needs to know which professional signs off, which evidence must accompany an answer, and how corrections reach downstream work.
Titles offer a poor shortcut here. A licensed lawyer may work as product counsel, legal engineer, solutions lead, knowledge director, or domain expert. A nonlawyer may have deep workflow knowledge while lacking authority to give legal advice. Recruiters need to verify jurisdiction, practice experience, product contribution, and the candidate’s actual decision rights rather than treating a legal keyword as a credential.
Inference serving reverses the bench and job-book ranking
The inference-serving talent map produces two different rankings. Anyscale has the largest mapped current-profile row at 489, followed by Groq at 393, Together AI at 280, Baseten at 189, and Fireworks AI at 157. The visible job book starts with Baseten at 166, then Together AI at 52, Fireworks AI at 47, Anyscale at 14, and Groq at one.
Baseten’s 189 current profiles and 166 jobs sit close together. Groq’s 393 and one create the longest contrast. The report keeps the one rather than treating it as a missing-data flag. It also warns that Anyscale, Groq, and Together AI can collide with other company names.
Those pairs are tempting because inference demand is a current strategic topic. The figures still cannot show whether Baseten is growing fastest or Groq has stopped hiring. These companies sell different combinations of hardware, cloud capacity, model APIs, software, and managed infrastructure. They publish jobs through different systems. A snapshot near a careers-site migration could change the job book without changing a workforce plan.
Product requirements make the labor boundary easier to establish. Ray Serve’s official documentation describes production LLM serving across multiple nodes and models, with autoscaling, load balancing, streaming, custom routing, observability, fault tolerance, and several forms of parallelism. Those requirements persist after a model passes an offline benchmark.
Anyscale builds around Ray, so its visible bench includes work connected to a broad distributed-computing platform as well as serving. Groq’s stack includes purpose-built inference hardware and cloud operations. Together AI, Fireworks AI, and Baseten expose their own mixes of model access, optimization, deployment, and enterprise service. The five names belong in a comparison set, but their employees are not interchangeable units of one occupation.
Baseten’s careers page calls inference an interdisciplinary problem spanning performance, reliability, latency, and economics. That description names the work better than “AI infrastructure.” Runtime and compiler engineers improve how models execute. Distributed-systems engineers manage placement, scheduling, and failure, while site-reliability and capacity teams keep workloads available. Performance engineers trace tail latency and throughput. Developer-experience and customer engineers make the system usable for actual models and traffic.
Commercial roles remain part of the boundary. Inference is purchased through capacity commitments, usage, support, deployment options, and service guarantees. A technically efficient runtime can lose if customers cannot migrate a model, secure a region, predict a bill, or get help during an incident. Sales engineers, account teams, finance, and support translate low-level performance into an operating contract.
Most application companies should not build every serving layer. Managed infrastructure can turn scarce systems work into a bill and service agreement. The buyer still owns traffic forecasts, workload tests, fallback behavior, data constraints, and the product response when a provider misses. If nobody inside can distinguish a model regression from capacity exhaustion, the contract has removed headcount without creating operating control.
Hiring one generic machine-learning engineer will not cover this stack. Model training experience can help, but serving changes the objective. The team cares about time to first token, tokens per second, batch efficiency, cache behavior, cold starts, utilization, error recovery, regional capacity, and unit cost under a real traffic distribution. The best benchmark at batch size one may be an expensive production choice.
Interview evidence should match the failure surface. Ask a runtime candidate to explain a performance regression with trace data. Ask a platform engineer how they would isolate a noisy tenant or roll out a new engine without breaking API behavior. Ask a capacity lead how a forecast handles bursty traffic and scarce accelerators. Ask a customer engineer which problem should become platform capability and which should remain a deployment exception.
Use the job-book reversal without turning it into a growth claim. It tells a researcher to inspect why Baseten exposes many roles relative to its mapped current profiles and why Groq exposes one. Official career systems, funding, product changes, geography, and direct company confirmation can test the pattern. Until then, the map marks a difference in visible recruiting posture on August 23.
Infrastructure labor is easy to hide behind model prices. A buyer sees an API rate or committed-capacity quote. The provider carries people who tune kernels, replace failed machines, forecast regions, investigate latency, negotiate capacity, and answer the pager. A low token price can coexist with a labor-intensive service. The map points toward that workforce, even though it cannot count all of it.
A four-boundary hiring map
Turn the observation unit from company size to product failure, and the four reports become actionable. A hiring team can use a boundary map before approving a title or sourcing plan.
| Product boundary | Customer promise | Failure surface that reaches people | Role families to inspect | Interview evidence | Operating measures | Likely budget owners |
|---|---|---|---|---|---|---|
| Forward deployment | The product works inside a specific customer’s data and workflow | Integration gaps, access constraints, bespoke code, weak handoff, adoption failure | Applied or forward-deployed engineering, solutions architecture, deployment strategy, product, customer success | A production deployment with a changed requirement, reusable product decision, and explicit exit | Time to first working use, exception volume, rework, support load, reusable product contribution | Product, services, sales, customer success |
| Voice and speech | A live interaction sounds natural and completes the user’s task | Delay, interruption, transcription error, pronunciation, unsafe response, failed transfer | Speech research, audio, distributed systems, telephony, language quality, safety, solutions, revenue, support | An end-to-end latency or quality incident traced across components and resolved for users | Time to first audio, tail latency, task completion, escalation, error review, renewal | Research, product, infrastructure, operations, revenue |
| Legal AI | A professional can rely on a source-linked output in a consequential workflow | Wrong authority, confidentiality breach, missing citation, review burden, pricing conflict | Legal experts, knowledge engineering, product, security, implementation, customer success, sales | A matter-specific workflow with authority, review, correction, and client communication | Citation and answer quality, correction time, reviewer effort, task economics, adoption, client outcome | Product, practice leadership, knowledge, risk, IT, client teams |
| Inference serving | A model performs reliably under production traffic at a predictable cost | Tail latency, capacity shortage, poor utilization, runtime failure, migration friction, bill surprise | Runtime, compilers, distributed systems, SRE, performance, capacity, developer experience, customer engineering | A production performance or reliability problem connected to traces, tradeoffs, rollout, and recovery | Throughput, tail latency, availability, utilization, error rate, time to recovery, cost per completed workload | Infrastructure, engineering, finance, product, enterprise operations |
This is a planning artifact, not a taxonomy standard. One person can cross several columns. A forward-deployed engineer may investigate inference latency. A voice company may employ legal and safety specialists. A legal product may run its own inference platform. Keep the primary failure surface visible while the team decides where a role belongs.
The boundary can also move. A vendor may absorb runtime work this quarter and ask the customer to operate a private deployment next quarter. A professional-services partner may own implementation during launch, then leave the customer’s team with support. Record both the launch owner and the steady-state owner. Otherwise a hiring plan can look complete on opening day and fail at handoff.
Write the work unit in plain language during intake. “Deploy an agent” is too broad. “Connect a claims workflow to three customer data systems, move it into production, and reduce manual exception handling without exposing protected data” gives a candidate and manager something testable. For voice, the sentence might be: “Keep a multilingual phone agent below an agreed p95 response time while preserving a reliable human transfer.”
The same intake needs a consequence. A delayed demo may belong to product, while a broken customer production workflow brings support and account ownership into scope. A wrong legal citation calls for professional review and risk. A capacity miss can reach finance through emergency compute and service credits.
Public market evidence comes after the work definition. Use multiple title variants, separate current profiles from open postings, and record the date, company aliases, geography, and classification rule. A wide query can reveal adjacent work, while a narrow query tests whether the market uses the proposed title. Neither result should become an automatic candidate list.
Bring the public pattern back inside. Product incident logs, support queues, implementation plans, customer escalations, usage patterns, and roadmap debt show where labor is already being spent. Finance can expose contractor, cloud, travel, service-credit, and rework costs that do not appear in headcount. Employees doing the work can explain which tasks lack an owner.
Now the team can choose a role family. A repeated customer integration may justify a forward-deployed engineer, a platform feature, a partner program, or better implementation documentation. Voice latency may require a speech researcher, a network change, a telephony specialist, or a different service-level promise. Legal review burden may call for a knowledge engineer rather than another model engineer. Inference cost may be a runtime problem or a purchasing problem.
Candidate evaluation should reproduce the boundary. Generic AI trivia rewards familiarity with fashionable tools. A work sample should show how the person gathers missing context, chooses a metric, handles an exception, documents a handoff, and decides what belongs in the product. For regulated or customer-facing work, the assessment should include communication and escalation, not only technical output.
Set a validation date before opening the role. Later, compare applications, qualified interviews, offer acceptance, time to productive work, exception reduction, and movement in the operating measure named at intake. Change the title if it attracts the wrong evidence. Stop adding seats if the problem belongs to product design or procurement.
Metix maps can help at the public-evidence stage. They cannot replace the rest of the file. Their best use is to prevent an organization from treating one fashionable title as the universal labor unit of AI.
Monday’s requisition begins at the product boundary
Imagine the same request reaching four hiring managers on Monday morning: “We need another AI engineer.”
For the enterprise software manager, a customer’s production data cannot move into the standard environment. The role needs enough engineering depth to build safely, enough product judgment to resist a permanent fork, and enough customer authority to agree on a handoff. Searching only for an FDE acronym will discard relevant deployment histories and admit people whose title never carried production ownership.
Across the hall, a voice manager has a caller waiting too long for a reply. Model inference is only one part of the delay; endpoint detection, an LLM, tool calls, synthesis, the network, and playback share the budget. The manager needs the trace before deciding whether the next hire belongs in speech research, distributed systems, telephony, or customer integration.
The legal product manager brings a document-review workflow that returns plausible answers with uneven authority. Its next hire may be a legal subject-matter expert who can design evaluations, a knowledge engineer who can improve retrieval, or a product engineer who can make corrections and citations travel through the workflow. “Legal AI engineer” hides the decision.
One floor down, an inference workload passes its demo and misses the production cost target at peak traffic. Its owner may need runtime optimization, scheduling, capacity, reliability, or commercial architecture. Training another model could add expense without touching the bottleneck.
Each requisition enters the same labor market only at a high altitude. On the ground, the work, evidence, budget, and consequence differ. That is why the four public maps should not be collapsed into a league table. Their value comes from the mismatch.
Salesforce can dominate a forward-deployed title query while Palantir nearly disappears under exact matching. ElevenLabs can show a large visible bench and a broad revenue job surface because a voice product has to cross research, infrastructure, adoption, and service. EvenUp can map many legal roles and few lawyer titles because domain work does not always advertise a license in the title. Baseten can lead a visible inference job book while Anyscale leads the mapped current bench.
These are starting signals. A defensible hiring plan records the population, checks the aliases, reads the official roles, talks to the people carrying the exceptions, and measures the failure the hire is meant to reduce.
By Monday afternoon, the requisition should have changed. It should name a production boundary, an owner, evidence a candidate can show, and an operating measure that will move if the hire works. The title comes last.