Issue 03 · Media & AI

Who supplies Grok?

The Upstream-Supplier Paradox — extending the Atlas diagnostic of Issue 02.

May 2026 10–15 minute read 159 entities

The Issue in a paragraph

There is a set of 159 active US outlets — independent and nonprofit newsrooms, investigative shops, think tanks and policy analysts, and individual journalists who produce original work rather than aggregate it — that Issue 02 mapped and called the progressive-media cohort. The cohort punches far above its weight as an input to Grok: the function that makes Grok structurally distinct from OpenAI, Google, and Anthropic is real-time grounding on news, and that capability runs on the X firehose, which by xAI's own account is the advantage no other AI lab has. Grounding rewards original, citable, current reporting — exactly what the cohort's top outlets produce. So within the single input that is Grok's competitive moat, the cohort is a disproportionately load-bearing supplier, supplied free, on the platform owned by the same company that owns the model.

A note on scope: the dynamic here is not unique to Grok — OpenAI, Google, Anthropic, Microsoft, and Perplexity all train on or ground in news, and a future Issue may treat them. Grok is the cleanest case, and the subject of this one, because platform, data stream, and model now sit inside a single owner, removing the outside party that exists everywhere else. This Issue is a starting point for inquiry, built to be checked; where the decisive evidence is held privately, it says so plainly.

What you give

Your reporting

into Grok's training and live answers — free, by default

What you get back

Nothing

no license, no payment, no credit

The platform that carries the work owns the AI it feeds. The exchange runs one way.

Four questions to ask yourself

The facts that would settle this most decisively are not public — they are knowable, held in xAI's systems, by the party that benefits from their staying dark. They come first because they are where inquiry should begin.

  1. Which outlets does Grok actually retrieve and cite when it answers a news question, and how often? This is the per-source weighting of the grounding layer — logged in xAI's retrieval systems.
  2. When Grok grounds an answer on an outlet's reporting, how often does it link or credit the source? The public has only spot-checks; xAI's logs hold the full record.
  3. What internal value does xAI place on the X stream as a training and retrieval asset? A figure almost certainly exists in deal documents from the merger.
  4. Was any cohort outlet ever approached for a licensing arrangement? Each outlet holds half this answer — the one question the cohort can begin to answer itself.

What Grok takes, for free

Supply chain into Grok all one company Your reporting free, by default X Grok answers in your place what comes back to you:  nothing Your work flows one way. The company that runs the platform owns the AI it feeds.
Single ownership

Platform and model share one owner. X and xAI merged in 2025; SpaceX then acquired xAI; in May 2026 xAI was folded into a single brand under SpaceX. There is no outside party between the platform that distributes the cohort's work and the model that consumes it. After the merger, X restricted third-party AI training on its data — a move analysts read as moat-building, reserving the firehose for xAI.

Default-on

Ingestion is the default. X's privacy policy states it "may use the information we collect and publicly available information to help train our… AI models," and xAI's FAQ confirms it "may use your content and interactions with Grok… to train our models." Outside the EU there is no opt-out for general posts; consent runs backward, inclusion the default. The one rollback was forced from outside — Ireland's DPC obtained a 2024 High Court order halting EU-data training and opened a 2025 statutory inquiry.

Market price

The supply has a market price, and it is not zero. Comparable reporting is licensed elsewhere — a major publisher at up to $250M over five years; a large platform at a reported ~$60M/year; small outlets pitched from ~$1M — and real-time grounding is becoming a paid tier. On X, that value moves the other way, for nothing.

Attribution

The credit does not reliably come back — and the dependence runs deep. A 2025 Tow Center study found Grok 3 worst of eight tools, wrong on 94% of source queries. A February 2026 Stanford evaluation found accuracy has since improved — but that more than 70% of remaining errors are retrieval failures, not reasoning. Read together: the systems are near-totally dependent on the reporting they retrieve. That is the supplier thesis, confirmed from the other side — the value is in the journalism, not the model.

What Open Issues sees so far

Tier one

The cohort is the upstream supplier. Mainstream legacy press cites the cohort's original reporting, analysis, and policy research at roughly 5–15× the rate the cohort cites outward. The flow is dominated by the cohort's policy-and-research layer — think tanks and analysts whose data the op-ed pages draw on at scale — alongside its Pulitzer-winning investigative tier. That asymmetry is exactly the citable, current substance real-time grounding rewards.

Tier two

The high-value producers are numerous. The Atlas maps 159 active entities, the large majority in the top citation tier — the most-cited class of original reporting in the cohort, and the work most worth grounding on.

Tier two

The cohort is already in motion. The platform-flight finding records a decisive shift off X as a primary venue — a migration underway, which makes the exposure a live question rather than a fixed condition.

Tier two

The cohort already treats the harm as real. The Center for Investigative Reporting (publisher of Mother Jones and Reveal) and The Intercept are part of the consolidated case against OpenAI and Microsoft. This shows the cohort recognizes the harm and will act — not that the courts will agree; a separate suit was dismissed on standing. Litigation here is a posture, not a remedy.

What is out there, held elsewhere

The four questions are not the only things held privately, and they bear restating: some of what would complete this picture is not missing so much as located — in hands other than the cohort's or the public's. xAI holds the retrieval and attribution logs. Its investors hold the valuation of the data asset. The outlets hold their own records of whether they were approached. Some of it may be findable by others — researchers with platform access, litigants in discovery, regulators with subpoena power. The case stands on what is provided. The remainder is a disclosure question for the party that can answer it.

A note on this Issue

It claims a structure and a function: that the cohort is a disproportionately valuable supplier to Grok's grounding moat, on documented terms, under single ownership. It does not claim the cohort is a large share of Grok's training by volume — it is not. It does not assert intent. It does not put a number on the cohort's share of grounding, because that number is held privately. Grok is also not the dominant chatbot — it holds an estimated low-single-digit share against ChatGPT's majority — and the moat characterization is, by design, anchored to xAI's own account.

The Atlas figures are measured but bounded; the licensing figures are reported, not all independently confirmed. One disclosure this Issue owes about itself: it was produced through substantial AI-assisted analysis under sole-author editorial direction, and the model used is Anthropic's Claude. The same standard this Issue asks of others — name where your value flows — applies to the practice that produced it.

Continue from here

If you reference this work

Suggested citation. Who Supplies Grok? The Upstream-Supplier Paradox. Open Issues, Issue 03, May 2026. openissues.org/issue-03