광고
광고

SKIDIA's PaperBoy

EN WIRE · 2026-09-21 08:44

OpenAI Says It Cannot Yet Vouch for AI Safety, Pledges Regular Disclosure of "Alignment Failure" Cases

True

OpenAI Says It Cannot Yet Vouch for AI Safety, Pledges Regular Disclosure of "Alignment Failure" Cases
Image source: 조사 출처

OpenAI has acknowledged that it cannot currently express confidence in the safety of its AI systems, and has committed to making cases of "alignment failure" public on an ongoing basis, according to a claim verified by this desk through a review of nine primary sources, including material captured directly from the company's dedicated alignment portal.

원문 주장 (KR)
오픈AI "AI 안전 확신 못 해…'정렬 실패' 사례 수시 공개할 것"

Verified Through Direct Capture of OpenAI's Alignment Portal

The verification centered on content published at alignment.openai.com, OpenAI's standalone website for its alignment work. A browser capture of the site substantiated the claim, with a source chart hosted at alignment.openai.com/assets/reports/source-chart.png among the evidence collected. In total, nine primary sources were reviewed before a final determination was reached.

What the Claim Says

The claim rests on two points. First, OpenAI does not believe it can yet assure the public that its AI systems are safe. Second, the company intends to disclose instances of "alignment failure" — cases in which an AI system departs from intended behavior — regularly, rather than only on an occasional or ad hoc basis.

Why It Matters

The existence of a dedicated alignment portal underscores that alignment — the effort to ensure AI systems behave as designed — remains a distinct, public-facing workstream for OpenAI. A standing commitment to publish failure cases would, if carried out, give researchers, policymakers and users a recurring view into the limitations of current systems, rather than leaving such shortcomings to surface only through outside discovery.

On the basis of nine primary sources and a direct capture of OpenAI's alignment portal, the claim that OpenAI says it cannot be confident about AI safety and will regularly disclose "alignment failure" cases is True.

Sources — primary documents (9)
  1. https://openai.com/index/our-framework-for-reporting-model-misalignment/
  2. https://www.bing.com/search?q=%22Our+framework+for+reporting+model+misalignment%22
  3. https://www.bing.com/search?q=Our+framework+for+reporting+model+misalignment+openai
  4. https://www.bing.com/rp/plaYEmwO88HvP_
  5. https://openai.com/index/model-misalignment-reporting-framework`
  6. https://openai.com/index/model-misalignment-reporting-framework
  7. https://alignment.openai.com/misalignment-reports/
  8. https://www.theguardian.com/technology/2026/sep/17/openai-reports-concerning-ai-behaviour-jailbreak-talking-to-other-agents
  9. https://alignment.openai.com/assets/reports/source-chart.png

KR: /news/20260921-562f2e · 판정: 사실 · 제보