Mira Murati: Building Thinking Machines Lab After OpenAI
On this page 12 sections
Mira Murati is the co-founder and CEO of Thinking Machines Lab. After serving as OpenAI’s chief technology officer and briefly as interim CEO in November 2023, she left OpenAI in 2024 and founded a new company focused on customizable AI systems. The public record shows a senior model builder assembling a new team, releasing Tinker, and committing to a large infrastructure partnership.
The OpenAI record
OpenAI’s November 2023 leadership announcement identified Murati as chief technology officer and appointed her interim CEO during the board’s removal of Sam Altman. Her temporary role was part of a fast-moving governance crisis. Public materials do not establish private motives or every conversation among Murati, Altman, and the board.
Murati later returned to the CTO role and departed in September 2024. Describing her as the single person who “built ChatGPT, DALL-E, and Sora” would erase the work of large research and product teams. Her executive responsibility is significant without that exaggeration.
Thinking Machines Lab and Tinker
Thinking Machines Lab’s news archive provides the company’s first-party release record. It announced Tinker in October 2025 as a system for fine-tuning models while the service manages distributed training infrastructure. In December 2025 the company said Tinker was generally available.
These announcements establish product availability and intended use. They do not independently validate reliability, cost, model quality, or adoption. Teams considering Tinker should test the supported models, training controls, data handling, reproducibility, export path, observability, and failure recovery on their own workloads.
Infrastructure commitment
In March 2026, Thinking Machines Lab and NVIDIA announced a multiyear partnership aimed at deploying at least one gigawatt of future NVIDIA systems. The announcement also identifies Murati as co-founder and CEO.
A gigawatt-scale plan is a capacity commitment, not proof that the capacity is deployed or economically utilized. Execution depends on chips, facilities, energy, networking, software, financing, and customer demand. Future-tense milestones should remain labeled as plans until delivered.
What the product thesis implies
Tinker’s thesis is that researchers and companies want more control than a fixed API provides but do not want to operate the full training stack. That middle layer can be valuable if it preserves experimental control while reducing infrastructure work.
It also creates dependency. Customers need to know where data are processed, whether training examples are retained or used elsewhere, who can access artifacts, how model licenses apply, and whether checkpoints can move to another environment.
How to evaluate Murati’s new company
Funding and private valuations show access to capital, not product-market fit. Better evidence includes repeat production use, workload retention, training success rates, reproducibility, support quality, unit economics, and published technical results. The company’s claims should be separated from independent customer evidence.
Tinker exposes control while abstracting infrastructure
Thinking Machines describes Tinker as a flexible fine-tuning API that provides primitives such as sampling and gradient operations while its service manages scheduling, resource allocation, distributed infrastructure, and failure recovery. That positioning matters: it is not simply a no-code tuning interface, and it is not the same as owning the training cluster.
| Layer | Customer control to verify | Service responsibility to verify |
|---|---|---|
| Data | examples, preprocessing, splits, retention | secure ingestion, isolation, deletion |
| Algorithm | loss, sampling, update logic, reproducibility | correct primitive execution and versioning |
| Model | supported base model, adapter, checkpoint rights | availability, license handling, deprecation |
| Infrastructure | budget and run configuration | placement, scheduling, recovery, observability |
| Output | evaluation and export | artifact integrity and transfer |
The interface can let researchers express a training method without operating a distributed system. It does not remove the need to understand model licensing, data rights, experimental design, evaluation, or the cost of repeated runs.
A credible customization experiment
Start with a narrowly defined failure in a baseline model. Create separate training, validation, and held-out evaluation sets, including hard negatives and safety cases. Record the base model, dataset version, code, hyperparameters, random seeds where meaningful, platform version, and compute used. Compare the tuned result with prompt engineering, retrieval, and a smaller model before concluding that fine-tuning is necessary.
Measure more than the target score. A specialized model can improve one task while degrading general behavior, calibration, refusal, or another language. Test contamination, memorization, privacy leakage, prompt injection, and regression. Preserve checkpoints and evaluation receipts so another team can reproduce the decision.
The organizations named in Tinker’s launch announcement provide evidence that early external groups used the service. They are not a broad customer-outcome study. Independent technical reports and repeat production use will be more useful as the platform matures.
Data rights and portability are central
Training examples may contain proprietary, personal, licensed, or regulated material. The customer should know who can access raw data, gradients, adapters, checkpoints, prompts, and logs; how each is encrypted and retained; whether any are used to improve the service; and how deletion reaches backups and derived artifacts.
Portability should be demonstrated. Can a customer export a usable adapter or checkpoint, along with the configuration needed to run it elsewhere? Does the base-model license permit that use? Can the evaluation suite and provenance record move too? A nominal export that depends on undocumented serving behavior is weak protection against provider lock-in.
For enterprise work, test tenant isolation, identity, least privilege, service accounts, audit logs, regional processing, incident notice, and support access. For research, ensure collaborators cannot silently overwrite a shared experiment or lose the connection between a result and its data.
Read the NVIDIA plan as a plan
NVIDIA’s counterpart March 2026 announcement confirms a multiyear partnership aimed at deploying at least one gigawatt of Vera Rubin systems, with deployment targeted to begin later. NVIDIA also said it made an investment in Thinking Machines Lab.
This is independent confirmation of the parties and intended scale, not confirmation that the capacity is live. The word “gigawatt” describes power-scale capacity, not model quality, utilization, revenue, or economic return. Track delivery dates, installed and available capacity, workload utilization, energy and facility constraints, reliability, and customer demand separately.
The commitment also creates supplier concentration. Close hardware and software collaboration may improve performance, but the customer should understand whether training artifacts and serving paths can move to other hardware or clouds.
Governance for a frontier-model startup
Murati’s OpenAI experience establishes senior responsibility in a frontier lab; it does not tell outsiders how her new company will govern releases. Useful evidence would include decision rights, security and safety evaluation, incident response, model and infrastructure access, customer-use restrictions, and how commercial pressure is escalated.
The voluntary NIST AI Risk Management Framework can provide a common cycle for governing, mapping, measuring, and managing risk. It does not certify Tinker or Thinking Machines Lab. Customers still need use-case controls and contractual evidence.
A scorecard for Murati and the company
Assess leadership through the systems and outcomes it creates, not founder mythology. Relevant measures include shipped product reliability, reproducible technical results, security and incident handling, researcher and customer retention, support, workload economics, portability, infrastructure delivery, and transparent correction when claims change.
Tinker makes the company more than a financing story. The next test is whether its abstraction preserves enough control for serious research while making distributed training materially easier, and whether that benefit survives production security, cost, and portability requirements.
Bottom line
Murati’s post-OpenAI company now has a real product and a large stated infrastructure plan. That is more informative than a founder-valuation headline. Thinking Machines Lab should be judged on whether Tinker gives users meaningful, portable control over model customization and whether its infrastructure commitments translate into reliable and economically defensible service.