मुख्य मजकूराकडे जा
OpenAI

१३ ऑगस्ट, २०२६

अनुप्रयुक्त AI

The builder’s guide to GPT‑5.6

Technical lessons from startups in production

लोड होत आहे...

GPT‑5.6 sets a new standard for price-performance

The GPT‑5.6 model family makes frontier-level agent performance dramatically more affordable, while also advancing the frontier of what is possible.

In this guide, we show how startups are using smarter model selection and new API controls that help with reasoning continuity, multi-agent orchestration, and programmatic tool calling to build faster, more capable agents at a fraction of the cost.

A better out-of-the-box experience

Since GPT‑5, each model generation has sought to tackle longer-horizon tasks with fewer tokens. GPT‑5.6 continues that trajectory: stronger agent performance, lower costs, with minimal changes to the underlying harness.

The improvements in top-line cost efficiency are compounded with increased accuracy at lower reasoning efforts. For example, on Agents’ Last Exam, GPT‑5.6 Sol at “low” reasoning outperformed GPT‑5.5 at “high” reasoning when the harness was kept constant. We’ve seen similar success stories in production testing where startups report seeing significant cost improvements across a range of workflows by reducing the reasoning effort from the prior defaults.

आम्ही GPT‑5.6 आमच्या हार्नेसमध्ये समाविष्ट केले आणि कमी रीझनिंग प्रयत्नांत आम्हाला सर्वोत्तम परिणाम मिळाले. त्याने ओळखले की त्याच्याकडे केव्हा डेटा कमी होता, चुकीच्या दिशेने शोध घेतला नाही आणि कमी टोकनमध्ये योग्य उत्तर मिळवले.
— इझी मिलर, AI संशोधन प्रमुख, Hex(नवीन विंडोमध्ये उघडेल)

Model Selection

Historically, upgrading to a flagship model at the highest reasoning available has been the best option for long-horizon use cases. This has been in large part due to these models being significantly more capable than cost-optimized models at handling longer contexts and tool calling. This has changed with the 5.6-family: with more test-time compute, Luna and Terra can often perform similar to GPT‑5.4 and 5.5 while being significantly cheaper.

3पैकी 1
Luna हे GPT‑5.5 च्या एक-अठरांश खर्चात त्याची 98% माहिती-काढण्याची अचूकता कायम राखते. त्यामुळे आमच्या एजंटना उच्च गुणवत्तेचे दस्तऐवज आकलन अशा किमतीत मिळते की ते आणखी अनेक कार्यप्रवाहांमध्ये वापरणे व्यवहार्य ठरते.
— सरहाय श्कोहोलिएव्ह, एजंट इंजिनिअरिंग प्रमुख, Hypha(नवीन विंडोमध्ये उघडेल)

Consider tasks in BrowseComp: a search-based benchmark that tests a model’s ability to search for obscure facts. Three months ago, GPT‑5.5 (Extra High) scored 84.36% on this benchmark for a total cost of $33.27. At launch, GPT‑5.6 Luna (Extra High) delivers essentially the same performance, scoring 84.04% at a cost of $1.33. We’ve since reduced prices further. Read more on our latest price cuts.

The smaller 5.6-family models are a strong fit for high-volume workloads, latency-sensitive interactions, and repeated steps within agentic workflows. For example, if you’re operating a legal-tech startup that parses handwritten memos prior to agentic analysis, instead of using a frontier model for the entire use case, you can now use Terra or Luna for extraction and register significant cost savings.

Evolving the Responses API to architect more efficient agents

In addition to making GPT‑5.6 more performant out of the box, we also shipped new primitives to the Responses API to unlock further gains. We trained GPT‑5.6 end-to-end with three complementary architectural interventions that enable agents to operate more efficiently:

  1. Reuse work already performed: by allowing reasoning to be persisted(नवीन विंडोमध्ये उघडेल) across model turns and using native compaction(नवीन विंडोमध्ये उघडेल) to compress long-running conversations, the model can maintain coherence in its work across longer task horizons without getting confused or having to reconstruct prior context.
  2. Parallel decomposition where appropriate: using native multi-agent orchestration(नवीन विंडोमध्ये उघडेल) allows coordinating multiple agents across parallel workstreams to finish complex tasks faster.
  3. Move deterministic work into code: using programmatic tool calling(नवीन विंडोमध्ये उघडेल) to filter, aggregate, and orchestrate tool outputs outside the model’s context window, reserving model tokens for judgment and reducing cost, latency, and context rot.

हे एकत्र वापरल्यास परिणामातील फरक विलक्षण असू शकतो. उदाहरणार्थ, स्टँडर्ड हार्नेससह ARC-AGI-3 वर GPT‑5.6 Sol ने 13.3% गुण मिळवले. मात्र, जतन केलेले रीझनिंग आणि कॉम्पॅक्शन सुरू केल्यानंतर सुमारे 6 पट कमी आउटपुट टोकन वापरूनही गुण 38.3%पर्यंत वाढले. मॉडेलमध्ये कोणताही बदल नाही, पण कामगिरी जवळपास तिप्पट झाली. आमच्या ARC-AGI-3 हार्नेस तपासाबद्दल येथे अधिक माहिती मिळेल.

प्रोग्रामद्वारे टूल्सना कॉल करणे

एजंट-आधारित कार्यप्रवाहांमध्ये अनेकदा दोन प्रकारची कामे असतात:

  1. निर्णयक्षमतेची गरज असलेली कामे
  2. मुख्यतः माहिती हलवणे, फिल्टर करणे आणि एकत्रित करण्याची कामे

एजंट 100 फाइल केलेली कागदपत्रे मिळवून तारखेनुसार फिल्टर करतो आणि संबंधित व्यवहार ओळखतो, तर मॉडेलला त्याच्या कॉन्टेक्स्ट विंडोमधील प्रत्येक मधल्या निकालावर रीझनिंग करण्याची गरज नसावी. प्रोग्रामद्वारे टूल्सना कॉल करताना GPT‑5.6 टूल्सचे संयोजन करण्यासाठी JavaScript लिहू शकते, स्वतंत्र कॉल्स समांतरपणे चालवू शकते आणि त्यांच्या आउटपुटवर कॉन्टेक्स्ट विंडोबाहेर प्रक्रिया करू शकते. त्यामुळे मॉडेल बुद्धिमत्ता आवश्यक असलेल्या गोष्टीवर, म्हणजे निर्णयक्षमतेच्या वापरावर, लक्ष केंद्रित करू शकते.

आर्थिक संशोधनात दाखल कागदपत्रे विश्वासार्हपणे मिळवणे, टूल्सचा समन्वय साधणे आणि आकडे तपासणे हे कठीण काम असते. आमच्या मूल्यमापनांत, प्रोग्रामद्वारे टूल्सना कॉल करणाऱ्या GPT‑5.6 ने 21% कमी इनपुट टोकन वापरूनही आमच्या निकषांनुसार समान गुणवत्ता साधली. आर्थिक संशोधनावर चर्चा करू शकणारा एजंट आणि ते काम प्रत्यक्षात करू शकणारा एजंट यांच्यात हाच फरक आहे.
— अलेक्स वांग, अप्लाइड AI, Rogo(नवीन विंडोमध्ये उघडेल)

बहु-एजंट

गुंतागुंतीच्या आणि समांतरपणे करता येणाऱ्या कामांमध्ये अनेक एजंट कार्यप्रवाहांदरम्यान कृती व रीझनिंगचे वितरण केल्याने काम जलद पूर्ण होते आणि बुद्धिमत्ताही वाढते. अशा व्यवस्थांमध्ये उप-एजंटचे संयोजन करणे आणि त्यांना कामे सोपवणे ही मुख्य एजंटची जबाबदारी असते. उप-एजंट आपापली उद्दिष्टे समांतरपणे पूर्ण करतात आणि शेवटी अंतिम संकलनासाठी आपले आउटपुट मुख्य एजंटकडे परत पाठवतात. Responses API मध्ये बहु-एजंट सक्षम करून(नवीन विंडोमध्ये उघडेल) टीम्स थेट या क्षमतेचा लाभ घेऊ शकतात. ChatGPT मधील Ultra क्षमता सेटिंगही अशाच प्रकारे काम करते.

2पैकी 1
Qualia मुक्त स्वरूपाच्या संशोधन समस्यांवर एजंटची एक टीम चालवते आणि GPT‑5.6 Sol त्यासाठी अगदी योग्य ठरले. GPT‑5.5 च्या तुलनेत त्यात लक्षणीय सुधारणा दिसली, आम्ही तपासलेल्या जवळपास प्रत्येक मॉडेलपेक्षा त्याने काम लवकर पूर्ण केले आणि ते झटपट आमचे पसंतीचे OpenAI मॉडेल बनले.
— ई ची, संस्थापक, Quadrillion(नवीन विंडोमध्ये उघडेल)

उप-एजंटची योग्य संख्या आणि ते केव्हा तयार करावेत याची GPT‑5.6 ला चांगली जाण असली, तरी बहु-एजंट वर्तनाला सहज दिशा देता येते. उप-एजंटना केव्हा कॉल करायचा याबद्दल मॉडेलला सूचना दिल्यास, अतिरिक्त टोकन खर्च केल्यामुळे चांगली कामगिरी मिळेल अशाच परिस्थितीत एजंट तयार होण्याची शक्यता वाढते.

प्रॉम्प्ट कॅशिंग

संपूर्ण मॉडेल कुटुंबासाठी प्रॉम्प्ट कॅशचा TTL किमान 30 मिनिटांपर्यंत वाढवला आहे आणि आता मॉडेलच्या कॉन्टेक्स्ट विंडोमध्ये निश्चितपणे कॅश ब्रेकपॉइंट ठरवता येतात. यामुळे स्टार्टअपना त्यांच्या कॅश हिटचे प्रमाण लक्षणीयरीत्या वाढवता आले आहे.

आम्ही शेअर केलेल्या 29,000 टोकनच्या प्रॉम्प्टमध्ये कॅश ब्रेकपॉइंट्स आणि प्रत्येक वर्कस्पेससाठी स्वतंत्र की जोडले आणि कॅश न केलेले इनपुट 28% कमी केले. 30 मिनिटांच्या कॅश कालावधीमुळेही मोठी प्रगती साधली: प्रत्येक वेळी सुरुवातीपासून काम करण्याऐवजी आमचे एजंट वेगवेगळ्या फेऱ्यांमध्ये तोच संदर्भ पुन्हा वापरू शकले.
— लॉरेंझो जेंटाइल, AI इंजिनिअर, Ploy(नवीन विंडोमध्ये उघडेल)

कॅश एंडपॉइंट ठरवण्याबरोबरच योग्य prompt_cache_key(नवीन विंडोमध्ये उघडेल) वापरत राहिल्याने, तोच उपसर्ग यापूर्वी हाताळलेल्या निष्कर्ष इंजिनवर विनंत्या पोहोचण्याची शक्यता वाढते आणि त्यामुळे लॅटन्सी कमी होते.

निष्कर्ष

या सर्व उदाहरणांमध्ये एजंट तयार करण्याचे अर्थकारण किती बदलले आहे, ही बाब विशेष ठळकपणे दिसते.

पूर्वी प्रत्येक टप्प्यावर अत्याधुनिक मॉडेल आवश्यक असलेल्या वापर प्रकरणांमध्ये आता लहान मॉडेल वापरून, रीझनिंगचे कमी करून आणि कार्यक्षम आर्किटेक्चर पर्याय निवडून अत्यल्प खर्चात समान किंवा अधिक चांगले परिणाम साधता येतात.

तुम्ही सर्वजण काय तयार करता, हे पाहण्यासाठी आम्ही खूप उत्सुक आहोत!

  • 2026
  • API प्लॅटफॉर्म

लेखकांविषयी

सुरुवातीच्या चाचण्यांपासून प्रत्यक्ष वापरापर्यंत GPT‑5.6 वर उत्पादने उभारणाऱ्या स्टार्टअपसोबत जवळून काम करण्याच्या अनुभवावर आधारित ही मार्गदर्शिका समर्थ मुदुरू(नवीन विंडोमध्ये उघडेल), प्रशांत मित्तल(नवीन विंडोमध्ये उघडेल), डेव्ह लियो(नवीन विंडोमध्ये उघडेल) आणि जूलियन रायमन(नवीन विंडोमध्ये उघडेल) यांनी तयार केले आहे.