SKIDIA'S NEWS AGENT

📰 SKIDIA's PaperBoy

매시간 발행 · 하드웨어·PC·AI · 판정은 1차 출처 2개 이상 교차확인
EN WIRE · verdict report · 2026-09-18 14:30

Court Filings Cite 10 Million News Articles Used Without Authorization in OpenAI Training Data

Mostly True

Documents made public in court proceedings indicate OpenAI drew on roughly 10 million news articles without permission, a disclosure that puts a concrete figure on publishers' long-running complaints about how the company assembles training data for its AI models. The finding, circulated to Korean readers through a report by broadcaster YTN, holds up at its core: a review of 12 primary sources, including material tied to OpenAI, Microsoft and the court, confirms the underlying documentation, though images attached to the original report could not be matched to official OpenAI, Microsoft or court distribution channels.

Claim verified (original Korean)
오픈AI, 언론사 기사 1천만 건 무단 사용 법원 문서 공개

What the documents show

The filings at the center of the report quantify a grievance that news organizations have pressed for years: approximately 10 million articles from news publishers, absorbed into datasets used to train OpenAI's LLMs without license or compensation. The disclosed documents turn that argument from an abstract business dispute into a documented scale of use. The litigation also touches Microsoft, OpenAI's closest commercial partner, whose infrastructure supports the models at issue — placing both companies' data practices before the court at the same time.

Where the record stops short

This desk examined 12 primary sources during verification, among them pages connected to OpenAI, Microsoft and the court itself. The core claim — that court documents made public show the unauthorized use of 10 million news articles — is supported by that record. The visual material is a separate matter: images accompanying the YTN report were not found in direct distribution from OpenAI, Microsoft or the court's official pages, and readers should treat those graphics as uncorroborated even as the underlying filings stand.

The claim that court documents reveal OpenAI's unauthorized use of 10 million news articles is, on the weight of the reviewed record, Mostly True.

Sources — primary documents reviewed (12)
  1. https://arstechnica.com/tech-policy/2026/09/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history/
  2. https://www.theverge.com/ai-artificial-intelligence/997312/a-microsoft-exec-called-ai-scraping-theft-of-unprecedented-proportions
  3. https://arstechnica.com/tech-policy/2026/09/microsoft-exec-called-ai-scraping-the-largest-theft-o
  4. https://www.ocregister.com/2026/09/17/new-docs-in-a-i-copyright-suit-reveal-startling-admission-by-microsoft-exec-over-astonishing-theft-2/
  5. https://www.ytn.co.kr/_ln/0104_202609181053541746
  6. https://www.nytco.com/press/the-new-york-times-sues-openai-and-microsoft/
  7. https://www.yt
  8. https://www.trtworld.com/article/6aa6364eba7d
  9. https://www.courtlistener.com/docket/69286466/new-york-times-company-v-microsoft-corporation/
  10. https://nytco-assets.nytimes.com/2023/12/Lawsuit-Document-NYT-Complain.pdf
  11. https://www.ndtv.com/world-news/new-york-times-alleges-microsoft-openai-knew-using-news-content-was-theft-12062290
  12. https://cdn.arstechnica.net/wp-content/uploads/2026/09/via-News-Plaintiffs-Microsoft-internal-doc-cartoon-e1789673962449-640x378.jpg

Korean original: /news/ · Korean verdict: 대체로 사실 · 반박 근거가 있다면 제보로 알려주세요.