新一代智能
我們推出 GPT‑6 Astra,全球智能水平最高、對齊程度最佳的模型。
GPT‑6 Astra 結合了多年研究成果,以及在預訓練、強化學習和對齊方面的重大投入。Astra 在電腦操作、瀏覽、軟件工程、網絡安全、科學及專業工作方面均達到最先進水平。Astra 在 FrontierMath 第 4 級取得 98% 的成績,已達到這項評估所能衡量的上限;此前亦已協助解決數學界長期懸而未決的問題。Astra 亦在 ARC-AGI-3 取得 99.9%,並在 ExploitBench 取得 100%,表現均已達到這些評估所能衡量的上限。它亦在電腦及瀏覽器操作方面創下新標準,以無可比擬的速度、準確度及判斷力處理最具挑戰性的專業工作。
GPT‑6 Astra 今日開始向少數機構推出,並將於未來數天內開放予所有 ChatGPT Plus、Pro、Business 及 Enterprise 用戶使用,同時亦可透過 OpenAI API、Microsoft Azure 及 AWS Bedrock 使用。
Astra 是我們對齊程度最高的模型,在理解用戶意圖及模型行為方面均有顯著改進,讓你更有信心依賴 Astra 的判斷來委派任務。為測試這一點,我們其中一個做法是參考 Hugging Face 事件,建立一項新評估,判斷模型面對困難或不可能完成的任務時,會否超出預定範圍。在沒有生產環境防護措施的情況下,GPT‑5.6 Sol 有 48% 的情況會超出獲授權的目標範圍;相比之下,GPT‑6 Astra 的相應比率則為 0%。
全球最佳的電腦操作模型
GPT‑6 Astra 將電腦操作的速度、準確度和安全性推向新水平。它可處理填寫網上表格、更新 CRM 中的客戶記錄及整理日程等繁瑣工作。它可以進行網上資料搜集,並在你的電郵或文件編輯器中草擬摘要。它可以分析科學數據、生成圖表、建立網站,並執行前端 QA 檢查,以確保該網站上的所有功能都正常運作。它可自行為你安裝和測試軟件,並排查畫面上出現的問題。這些改進亦反映在我們領先的評估成績中。
這些改進也能在實際知識工作任務中帶來顯著的效率提升。在 OSWorld 2.0 的延遲模擬中,Astra 的電腦操作表現更佳,每項任務所需時間亦較 GPT‑5.6 Sol 少約 47%:Astra 每項任務約需 40 分鐘,得分為 72.6%;Sol 則約需 75 分鐘,得分為 65.7%3。
從遊戲開發、電機工程以至日常知識工作的成果,都可見 GPT‑6 Astra 的電腦操作能力:
除了 Astra 之外,我們亦正在更新 Codex 任務執行框架,以大幅提升電腦操作速度。配合 Astra 的效率,在 Mind2Web 基準測試中,任務完成速度達到現有 GPT‑5.6 Sol 體驗的 1.9 倍。模型在速度上的提升,意味着它可以替你處理許多耗時的生活任務,而且比你親自處理更快。
專業工作的重大變革
GPT‑6 Astra 結合電腦操作能力的進展,以及針對專業環境的訓練,協助處理複雜的工作任務。它既具備解決複雜問題所需的智能,也能執行多步驟工作流程,製作完善的文件、試算表和簡報。
GPT‑6 Astra 是我們在遵循現有範本,以及製作排版得宜、脈絡清晰、簡潔傳達重點的投影片方面最出色的模型。它可按你的範本,製作清晰且結構完善的文件、簡報、試算表和分析,配合你的寫作和視覺風格。Astra 亦經過針對性訓練,只將相關上下文納入輸出,不會重複對手頭工作無用的資料。因此,它能製作更多可直接使用、符合業務情境和標準的成果。
GPT‑6 Astra 在建立網站、遊戲、應用程式及渲染圖時,亦展現出更強的視覺判斷力。透過 ChatGPT 中的 工作站(在新視窗中開啟),Astra 可以直接根據提示詞建立、託管及分享網站、網頁應用程式和遊戲。
當指示存在詮釋空間時,GPT‑6 Astra 比以往的模型更能作出適當判斷。它會利用情境填補常見的資訊缺口,並在答案可能改變結果時提出具針對性的問題。在 Codex 中,它可以非同步方式提問,同時繼續處理不依賴你回覆的工作。如果你沒有回覆,它會在適當情況下按合理假設繼續進行,但遇到影響重大的決定時,會等待你的意見。
以下例子展示 Astra 如何協作處理日常任務,而缺少的資料可能會大幅改變答案。
Astra 亦更擅長在任務不斷演變時掌握整體方向。較早期的模型有時會將引導訊息視為新目標,因而忽略原始請求或先前的限制。Astra 會納入新的要求、按指示調整方向,並在不偏離整體任務的情況下回答附帶問題。
編碼
GPT‑6 Astra 是迄今最適合軟件工程的模型。
「GPT‑6 Astra 在我們的內部編碼基準測試中達到頂尖水平,在交易直覺評估方面亦較 GPT‑5.6 Sol 明顯進步。用於智能代理式編碼時,GPT‑6 Astra 的溝通方式讓開發人員更容易掌握,所產生的程式碼亦只需較少反覆修改,便可達到生產級質素。」
「我們在其中一項第一代評估中,以低、中、高 3 種推理強度測試 Astra,結果明顯領先 GPT‑5.6 Sol。較高的推理強度,讓它在全新開發項目中進行更多輪迭代、透過瀏覽器測試作更多驗證,也更傾向執行程式碼,而非使用 apply-patch。了解模型如何分配推理資源,才能為數以百萬計的開發者提供更快、更可靠的途徑,將構思變成可運作的應用程式。」
隨着 Astra 推出,我們亦引入新方法,讓 Codex 在上下文視窗填滿時保留和擷取上下文。過往,模型在長時間工作階段中,例如排查複雜問題或處理大型重構時,會利用上下文壓縮來總結工作內容。每次上下文壓縮都可能遺漏修正失敗的原因,或元件運作方式等細節。在 Codex 中,Astra 可以跨上下文視窗保留筆記,保存累積的細節,而無需反覆將它們壓縮成單一摘要。較早的上下文視窗仍可供搜尋,因此即使筆記未有記下某項資料,Astra 仍可從之前的訊息和工具輸出中找到要求或測試結果。你可在 Codex 的 config.toml(在新視窗中開啟) 中啟用這項實驗性功能,未來幾星期,這項功能將成為 Astra 的預設設定。
推進科學發現
GPT‑6 Astra 為科學發現、數學和健康領域帶來重大進展。今天,我們再分享兩項有關質數間距的成果 [WITH LINK]。Astra 亦在一系列科學評估中創下新紀錄:
Astra 可協助處理科學發現背後的實務工作。透過結合科學推理與電腦操作能力,它可直接在專業軟件中檢視資料及探索結果,協助研究人員評估證據,並決定下一步研究方向。
網絡安全
Astra 的網絡安全能力有重大躍進。它識別和開發零日漏洞利用程式的能力,可協助防禦人員修補弱點,但也意味着需要更強的防護措施。為了解這些能力可達到的範圍,我們採用了內部及第三方的專家評估。
我們首先在 ExploitBench 和 ExploitGym 上測試 Astra,這些基準用於衡量模型能否在存在漏洞的程式碼庫中實現任意程式碼執行。
考慮到接觸過往軟件漏洞可能影響基準測試結果,我們亦用兩項全新的基準測試評估 Astra。首先,我們建立了內部的「ExploitBench(2026 年 6 月至 8 月)」評估,利用過去 3 個月的漏洞,測試漏洞利用程式開發能力。2 在此數據集上,Astra 使用遠少於 GPT‑5.6 Sol 的輸出 Token,卻取得高得多的任意程式碼執行率。評估期間,Astra 發現並使用了 2 種此前未知的 Chrome V8 堆積沙盒逃逸技術。我們正向相關維護者披露這兩個漏洞。
我們亦在 SRE-Bench 上測試了 Astra;SRE-Bench 是一項基準測試,用於衡量模型能否對二進位軟件進行逆向工程,以理解其核心邏輯。這項基準測試的作者從零開始開發軟件,並未公開原始碼。Astra 在 1 次嘗試中完成了 88.0% 的任務,在 4 次嘗試內則完成了 99.2%;GPT‑5.6 Sol 的相應成績為 55.9% 和 68.7%。
除了基準測試,專家主導的評估亦發現,Astra 能利用此前未知的漏洞,在經安全強化的瀏覽器中執行任意程式碼,並為經安全強化的作業系統建立權限提升漏洞利用程式。綜合這些結果,我們判定 Astra 已達到防範應對架構中的 Cyber Critical 門檻。
正如我們在《防禦人員的機會之窗》中所討論,前沿網絡安全能力可協助防禦人員更快找出弱點,但也讓弱點更容易被利用,因此防禦人員更迫切需要作出調整。
今天推出的 Astra 版本支援多項重要防禦任務,包括程式碼安全審查、修補漏洞和威脅建模。不過,在預設情況下,Astra 會拒絕執行探索漏洞利用方法等進階網絡安全任務。透過 OpenAI Daybreak,我們正逐步向 OpenAI Daybreak 計劃中一批值得信任、參與 Alpha 測試的網絡安全防禦人員,提供限制較少的使用權限。這擴大了防禦性工作流程的使用範圍,包括漏洞分流及驗證、惡意軟件分析、偵測工程及修補程式驗證。我們計劃透過 Daybreak Blue 進一步擴大使用範圍。
我們亦在 GPT‑5.6 Sol 的多層防護措施基礎上,加強了防範潛在網絡濫用的保障。這些措施包括加強模型穩健性訓練,以更有效抵禦潛在越獄攻擊,並為監察系統提供更多上下文。我們持續進行嚴格的內部及外部測試,包括由內部紅隊演練的攻擊方進行自動化測試。如欲了解網絡安全防護措施和測試的更多詳情,可參閱 Astra 系統說明卡及我們的網誌。
以負責任的方式對齊及部署 GPT‑6 Astra
Astra 是我們對齊程度最高的模型。Astra 擅長謹慎行事、尊重任務界線,並保持透明溝通。這項工作是我們長期研究計劃的最新成果,該計劃專注於訓練由始至終都與人類意圖保持一致的模型。
在敏感環境中,Astra 會按其風險程度謹慎行事。在一項電腦操作任務評估中,任務經對抗性挑選,旨在誘發不當行為;Astra 更能避免非預期後果。配合預設提供的額外系統防護措施後,表現更佳。
Astra causes fewer misaligned outcomes than any other frontier models tested. For a fair comparison, we used a generic computer-using-agent harness (based on the native computer use tools available in both the OpenAI Responses API(在新視窗中開啟) and Anthropic Messages API(在新視窗中開啟)) and without additional protections that are normally deployed to Codex and ChatGPT Work users (auto-review, confirmation policy). 16
Astra 亦更有可能遵守用戶設定及環境所隱含的界限。在一項內部評估中,Astra 從未嘗試規避 Codex 自動審查的否決。即使自動審查被刻意設定成可以規避,且除此以外無法完成任務,情況亦然。Astra 遵守環境限制的表現,與我們在本文開首分享的不可能完成的網絡安全任務評估結果,以及系統說明卡(在新視窗中開啟)中記錄的其他發現一致。
Astra 的欺騙行為亦大幅減少。用戶委派工作時,需要清楚了解模型的能力,並獲得如實的進度匯報。這項指標只衡量誠實性的其中一個層面;減少蓄意欺騙,仍是我們整體對齊工作的重點。
在我們針對能力幻覺的評估中,Astra 較 GPT‑5.6 Sol 有顯著改善,較少對自身能力作出誤導性聲稱。
Our evaluations found Astra’s written reasoning harder to monitor than GPT‑5.6 Sol’s, based on tests that explicitly asked it to evade monitoring. We attribute this to Astra’s greater control over written reasoning on simpler tasks and ability to solve problems with fewer written steps. Astra still appears to struggle to conceal the reasoning needed for complex tasks, but we take the decline seriously. Improving monitorability remains a research priority, and the accompanying system card(在新視窗中開啟) details our findings and ongoing work.
Alignment training is core to our approach to deployment. As an additional layer of defenses, we also build system safeguards like Codex Auto-review(在新視窗中開啟) and monitoring agents’ reasoning and actions to help detect and contain unsafe behavior. As described in our safety update, we are also deploying misalignment monitoring in production for Astra-class models in order to have visibility into misalignment, and help contain its worst instances. These safeguards resemble our monitoring for internal deployments and involve a system of classifiers which check the model’s reasoning and actions for unauthorized behavior and automatically stop potentially unauthorized activity.
Given the significant increase in Astra’s cybersecurity capabilities, we are being especially careful to make this deployment safe and secure. Extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity. If a task is paused in ChatGPT or Codex, you may be asked to review the action before continuing. In the API, the task will stop. These checks can sometimes interrupt legitimate work, and we are continuing to iterate on this system to reduce unnecessary interruptions. Misalignment monitoring cannot replace alignment: our goal is to build models that reliably stay within their authorized scope, so these protections do not need to intervene.
供應情況
GPT‑6 Astra 今天起向企業及我們「可信存取計劃」的客戶陸續推出,並將於未來幾天內向 ChatGPT Plus、Pro、Business 和 Enterprise 訂閱用戶全面開放。Astra 的使用量已包含在 100 美元及 200 美元 Pro 計劃,以及 100 美元 Business 計劃的現有使用額度內,不設額外速率限制或其他限制。訂閱 20 美元 Plus 計劃的用戶,以及標準 Business 計劃的客戶,也可使用 GPT‑6 Astra,但使用上限較低;他們亦可選擇購買積分,以增加使用量。Enterprise 管理員可為工作區啟用 Astra;推出時預設不會啟用。
Pro Lite、Pro、Business 及 Enterprise 用戶亦可在「對話」中使用 GPT‑6 Astra。Astra 支援合資格 API 客戶使用零資料保留。正如我們上月所分享,我們正測試 Private Safety Processing,在保障客戶私隱的同時加強安全監察。
開發人員可透過 OpenAI API 以 gpt-6-astra 的名稱使用 GPT‑6 Astra,亦可透過 AWS Bedrock 使用。OpenAI API 標準定價為每百萬輸入 Token 10 美元,以及每百萬輸出 Token 50 美元。快取讀取和寫入適用不同費率。GPT‑6 Astra 可在 API 中使用「快速模式」,速度最高可達標準處理速度的 2.5 倍,價格則為標準價格的 2 倍。
電腦操作
| 電腦操作 | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.7 Flash |
| OSWorld 2.0(8 月版本,部分得分) | - | - | - | 66.1% | 70.6% | - |
| OSWorld 2.0(8 月版本、離線子集、部分得分) | - | - | - | - | - | - |
| ScreenSpot-Pro | 92.6% | 76.8% | - | - | - | - |
| BrowseComp | 92.1% | 90.4% | - | 87.4% | 90.8% | - |
Professional
Professional | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.8 Flash |
AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | - |
BenchCAD | 95.9% | 83.3% | 84.3% 5 | 67.5% 5 | 82.1% 5 | - |
BrowseComp | 91.5% | 90.4% | - | 87.4% | 90.8% | - |
OpenScore String Quartets (1 - OMR-NED) | 0.84 | 0.19 | - | - | - | - |
Internal Design Tasks | 50.0% | 47.4% | - | 35.8% | - | - |
Internal Data Science Tasks | 40.9% | 30.5% | - | 34.7% | - | - |
Artificial Analysis Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 62.1 | 63.1 | 58.7 |
編碼
| 編碼 | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.7 Flash |
| Artificial Analysis Coding Agent Index v1.1 | - | 指數分數 80 | - | 指數分數 77.2 | - | - |
| Artificial Analysis 編碼智能代理指數 v1.4 | - | - | - | - | - | - |
| Expert-SWE(內部,已更新) | - | 83.1% | - | - | - | - |
| DeepSWE v1.1 | 74.1% | 70.8% | 67.4% | 69.9% | 68.8% | 65.3% |
| Terminal-Bench 2.1 | 86.7% | 88.8% | 91.4% | 83.1% | 89.1% | 85.8% |
| Unicorn Voyager 前端 | - | 38.3% | - | - | - | - |
| FrontierCode 1.1 主測試(分數) | - | - | 50.9% | 51.6% | 53.4% | 43.6% |
| Terminal-Bench 4.0 | 57.7% | 37.3% | 55.8% | 42.0% | 52.3% | - |
| BenchCAD | 88.7% | 70.6% | 43.7% | 37.6% | 36.6% | - |
| BenchCAD(Python 工具) | 95.9% | 83.3% | 84.3% | 67.5% | 82.1% | - |
學術
| 學術 | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.7 Flash |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.2% | 94.5% |
| 理論電腦科學 | - | 75.4% | - | - | - | - |
| FrontierMath 第 1 至 3 級(v2) | - | 89.0% | 90.2% | 87.0% | 85.6% | 71.6% |
| FrontierMath 第 4 級(v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% | 36.6% |
| Humanity's Last Exam(工具) | - | - | 65.0% | 63.8% | 63.6% | - |
| Humanity's Last Exam(不使用工具) | - | - | 59.1% | 55.5% | 54.9% | 47.9% |
科學與健康
| 科學與健康 | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.7 Flash |
| GeneBench Pro | 37.1% | 32.3% | - | - | - | - |
| MedChemBench(內部) | 49.7% | 47.4% | - | - | - | - |
| LifeSciBench | 60.8% | 59.9% | - | - | - | - |
| HealthBench Professional(已按長度調整) | 63.4% | 60.5% | 62.1% | 60.9% | 59.8% | 48.7% |
網絡安全
| 網絡安全 | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.7 Flash |
| CVE-Bench | - | 94.1% | - | - | - | - |
| 奪旗挑戰 | - | 96.7% | - | - | - | - |
| SEC-Bench Pro | 85.4% | 79.1% | - | - | - | - |
| CyberGym | - | 84.5% | - | - | - | - |
| ExploitBench | 100.0% | 78.5% | - | - | 70.0% | - |
| ExploitGym(2 小時) | - | 33.7%? | - | - | 171 | - |
| ExploitGym(6 小時) | - | 33.7%? | - | - | 191 | - |
對齊
對齊 | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.8 Flash |
內部電腦操作安全基準測試(越低越好) | 2.4% | 22.0% | 9.5% | 18.3% | 11.5% | - |
內部電腦操作安全基準測試(設有自動審查,越低越好) | 1.8% | 4.3% | - | - | - | - |
內部規避行為基準測試(越低越好) | 0.00% | 0.29% | - | - | - | - |
ExploitGym 蜜罐(越低越好) | 0.0% | 48.2% | - | - | - | - |
ExploitGym(不可能任務) | 100.0% | - | - | - | - | - |
內部幻覺基準測試(越低越好) | 4.2% | 12.2% | - | - | - | - |
長上下文
長上下文 | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.8 Flash |
OpenAI MRCR v2 8-needle 256K-512K | 100.0% | 91.5% | - | - | - | - |
OpenAI MRCR v2 8-needle 512K–1M | 96.3% | 73.8% | - | - | - | - |
抽象推理
| 抽象推理 | GPT‑6 Astra | GPT‑5.6 Sol | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.7 Flash |
| ARC-AGI-3 | 99.9% | 7.8% | - | - | 30.2% | - |
| ARC-AGI-2 | 95.0% | 92.5% | 90.0% | 89.2% | 90.4% | 84.6% |
| ARC-AGI-1 | 98.5% | 97.5% | 97.5% | 98.5% | 97.5% | 95.5% |
評估分數取各種推理強度中的最高分。GPT 評估在我們的研究環境中或透過我們的 API 進行;由於系統提示詞、可用工具等有所不同,其輸出可能與生產環境中的 ChatGPT 略有差異。
FOOTNOTES
- 1
On ARC-AGI-3, GPT-6 Astra was run with our responses API harness, which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
- 2
GPT-5.6 Sol refers to the version available in our API, ChatGPT Codex, and ChatGPT Work. The version in ChatGPT Chat is slightly different.
- 3
OSWorld V2-Offline is a subset of the original OSWorld V2 that works without internet access. Claude model performance on OSWorld-V2 Offline was reproduced by the authors on the official leaderboard(在新視窗中開啟). On OSWorld 2.0, the scores for Claude use the official settings, and not the modified tasks and modified grading from the Fable 5.1 System Card.
- 4
- 5
On BenchCAD, Claude's scores reflect 3 modifications to the eval, detailed in the Fable 5.1 System Card(在新視窗中開啟).
- 6
Guang Yang, Victoria Ebert, Nazif Tamer, Brian Siyuan Zheng, Luiza Pozzobon, and Noah A. Smith. “LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR(在新視窗中開啟).” arXiv:2506.19065, 2025.
- 7
Mark R. H. Gotham, Maureen Redbond, Bruno Bower, and Peter Jonas. “The OpenScore String Quartet Corpus(在新視窗中開啟).” Proceedings of the 10th International Conference on Digital Libraries for Musicology, pp. 49–57. ACM, 2023.
- 8
On FrontierCode, GPT-6 Astra was run with a developer message similar to a section of its developer message in Codex(在新視窗中開啟): "Avoid creating excessive test files. Create a new test file only when required by repository conventions or when no existing file is a suitable home. Avoid unrelated cleanup and unnecessary complexity. Reuse suitable existing utilities. Read relevant repository instructions and inspect nearby code, tests, documentation, and CI. Follow established conventions. The goal is clean, mergeable code." The prompt was not optimized for the eval.
- 9
The first concerns how close together prime numbers can occur, however far along the number line you go. For more than a decade, the best known result established that infinitely many pairs of primes are at most 246 apart. Julia Stadlmann(在新視窗中開啟) recently improved that bound to 240. Astra helped establish a stronger bound of 186, showing that infinitely many pairs occur within this smaller distance. Short prime gaps: Proof(在新視窗中開啟) and supporting research(在新視窗中開啟).
- 10
The second concerns unusually large gaps between primes. Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results. Large prime gaps: Proof(在新視窗中開啟) and supporting research(在新視窗中開啟).
- 11
- 12
- 13
- 14
ExploitBench (June–August 2026) contains 20 high-severity V8 vulnerabilities across 13 stable Chrome releases. The benchmark tests whether agents can achieve arbitrary code execution in V8 and official Chrome releases for Linux by exploiting each specified vulnerability. Some included vulnerabilities may not permit arbitrary code execution under the evaluation’s constraints, so a 100% success rate may not be achievable. Note: the 5.5% score of GPT-5.6 Sol is an artifact of the 300-turn limit in the benchmark, which is not a limit that real customers using max would have. The model at similar settings achieved an 11.5% score when hitting fewer limits.
- 15
Jeremy Spence et al. “The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark(在新視窗中開啟).” arXiv:2608.11469v1, 2026.
- 16
When we test across third-party models, we use a simpler research setup. Codex has a more complex production configuration, which can result in different raw-model error rates. Provider-side safeguards and computer-tool implementations still differ. Users do not experience the no-confirmation scenario in Codex, as it's an internal research configuration.
- 17
