自己対戦を使って AI の安全性、アラインメント、プロンプトインジェクションへの堅牢性を高める、OpenAI の自動レッドチーミングシステム GPT-Red を紹介します。
GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model. The safeguards we have built for this launch—our most robust yet—are built to deliver these models safely and at scale, around the world.
OpenAI の新たな分析により、人気のコーディングベンチマーク SWE-Bench Pro に問題があることが明らかになり、AI モデル評価の信頼性と正確性に懸念が生じています。
GPT-Live-1 and GPT-Live-1 mini are a new generation of voice models designed to make conversations with AI feel more natural and intelligent.
複雑な実世界のデータセットを用いて、ゲノミクス、生物学、科学研究におけるAIの性能を評価する新しいベンチマーク、GeneBench-Pro をご紹介します。
GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model. The safeguards we have built for this launch – our most robust yet – are built to deliver these models safely and at scale, around the world.
OpenAI と Molecule.one は、GPT-5.4 を活用する準自律型AI化学者が、医薬品合成における重要な反応を改善し、医薬化学研究を前進させた事例を紹介します。
LifeSciBench のご紹介。専門家が作成・レビューした、AI システムが現実のライフサイエンス研究タスクと意思決定にどう対応するかを評価するベンチマークです。