Jonathan Ross founded Groq after working on Google’s first Tensor Processing Unit, then led a different approach to AI inference built around compiler scheduling and predictable execution. The current story has an important boundary: in December 2025, Groq and Nvidia signed a non-exclusive technology license, and Ross joined Nvidia with other Groq employees. Groq said the company remained independent and that GroqCloud would continue operating.

As of September 13, 2026, Groq’s company page names Adam Winter as chief executive. Ross should therefore be described as Groq’s founder and former CEO, not its current operating chief. The public record does not support the older narrative that he left Google carrying a secret, nor does it show that Nvidia acquired Groq.

From Google’s TPU team to Groq

A TIME profile of Ross describes his work on Google’s early TPU and his later decision to found Groq. An Axios profile independently covers Groq’s inference focus and Ross’s effort to compete in a market dominated by Nvidia hardware.

These accounts support a clear but limited career claim: Ross had direct experience developing specialized machine-learning hardware at Google and used it to frame a new architecture. They do not make him the sole creator of the TPU. Chips, compilers, systems software, fabrication, deployment, and operations are team efforts, and public profiles rarely expose every contributor.

Groq was founded in 2016. Instead of optimizing a general-purpose GPU design, the company built what it calls a Language Processing Unit, or LPU. The label is Groq’s trademark and product category, not an industry-standard conclusion that the hardware processes language in isolation.

What the LPU claim means

Groq’s technical explainer says its compiler schedules operations and data movement in advance, reducing the need for the dynamic scheduling and cache hierarchy common in other accelerators. The company argues that this makes token generation fast and predictable.

This is a vendor explanation of its architecture. It is plausible engineering context, but it does not establish superiority for every model or workload. End-to-end inference includes model loading, prompt processing, networking, batching, sampling, safety filters, and output delivery. A chip-level number can look excellent while a user’s total request latency is dominated elsewhere.

Four measures are commonly confused:

  • Time to first token measures how long a user waits before output begins.
  • Inter-token latency measures the cadence of generated tokens.
  • Throughput measures work completed across concurrent requests.
  • Total cost includes hardware or API price, utilization, power, networking, and operations.

Optimizing one can hurt another. Aggressive batching may raise throughput while increasing the wait for an individual request. A low per-token price may not help if the required model, context length, reliability, or regional capacity is unavailable.

Groq’s production-readiness documentation provides its own guidance for measuring latency. Buyers should use that documentation to configure a fair test, then collect their own traces rather than quote a showcase result.

Funding showed demand, not a settled winner

Groq announced a $640 million round at a $2.8 billion valuation in August 2024. In September 2025 it announced another $750 million round at a $6.9 billion valuation.

The releases establish the terms the company disclosed. The valuations were private financing prices, not exchange-traded market values. Statements about demand, capacity, customers, or deployment targets in the releases are company claims distributed through a press-release service. They do not reveal audited revenue, gross margin, utilization, supply commitments, or customer concentration.

The financing nevertheless identifies the strategic problem investors were backing. As model use moved from training experiments to production applications, inference latency and cost became a major part of AI economics. Groq attempted to address that problem with both its architecture and GroqCloud, allowing developers to consume the hardware through an API rather than operate accelerators themselves.

The Nvidia agreement was not an acquisition

On December 24, 2025, Groq announced a non-exclusive inference-technology licensing agreement with Nvidia. The same announcement said Ross, president Sunny Madra, and other team members would join Nvidia, while Groq would remain an independent company and continue GroqCloud.

Nvidia’s subsequent quarterly filing provides unusually strong corroboration. Nvidia described a non-exclusive license and the hiring of certain Groq employees. It specifically said Nvidia did not purchase Groq’s equity, customer contracts, or products.

That distinction matters. An acquisition transfers ownership of a company or assets under agreed terms. This transaction combined a technology license with employee moves while leaving Groq independent. Calling it a Groq acquisition would misstate the legal and operational structure disclosed by both parties.

The transaction also changes how Ross’s original competitive narrative should be read. He moved from leading an Nvidia challenger to working inside Nvidia under a license arrangement. Public documents do not disclose his current responsibilities, the licensed intellectual property in technical detail, the duration of the agreement, or how the consideration is allocated among licensing and employment arrangements.

Groq’s current company page names Adam Winter as CEO. This supports the leadership transition but does not reveal board control or the company’s post-license financial position.

A fair inference benchmark

A buyer comparing Groq with GPUs, other accelerators, or hosted APIs should freeze the workload before measuring. Use the same model revision, quantization where possible, prompt set, output-length distribution, sampling settings, concurrency, region, and service-level target.

The evaluation should report at least:

  1. Median and tail time to first token.
  2. Median and tail inter-token latency.
  3. Successful requests per second at several concurrency levels.
  4. Error, timeout, and rate-limit behavior.
  5. Quality on the application’s task, including structured-output validity.
  6. Cost per successful task, not only cost per token.
  7. Model availability, context window, and time to support new models.

Tail performance matters for interactive products. Aggregate throughput matters for batch systems. Neither should be substituted for answer quality. If a faster deployment uses a different quantization or model, the comparison must disclose that difference.

Operational diligence is equally important. Teams should test regional capacity, failover, quota increases, observability, data retention, and how quickly a model or API version can change. A benchmark run during a lightly loaded demo does not establish production reliability.

Strategy after the license

The agreement leaves Groq with a different strategic test. Its public announcement says the company remains independent, but some senior leaders and licensed technology moved to the dominant accelerator supplier. GroqCloud can still compete on developer experience, model availability, latency, capacity, and price. The company may also benefit if the license broadens use of its architectural ideas.

There are unresolved risks. Retaining engineering talent after a major team transfer can be difficult. A non-exclusive license may reduce scarcity around the licensed technology. Nvidia’s scale can accelerate adoption of ideas while making differentiation harder for the original company. None of those outcomes is predetermined, and the public record as of this revision is too early to measure them.

Ross’s durable contribution is the architectural argument that inference deserves hardware and software designed around predictable execution. Whether Groq captures the long-term economic value is a separate question from whether the argument influences the market.

Source and correction note

This revision uses Groq’s technical and transaction materials, Nvidia’s regulatory filing, named reporting, and dated financing releases available through September 13, 2026. Vendor benchmarks and financing metrics are labeled as such. The previous version used a fictionalized departure scene, attributed the TPU too narrowly, implied unverified infrastructure share, and treated Groq as a straightforward Nvidia challenger without accounting for the December 2025 license and leadership change. Those claims have been corrected.