مرکزی مواد پر جائیں
OpenAI

API صارفین کے لئے ترجیحی پراسیسنگ

Fast mode offers reliable, high-speed performance with the flexibility to pay-as-you-go. For our latest frontier model, gpt-5.6-sol, you can access up to 2.5x faster speeds with Fast mode.

By choosing Fast mode, you can unlock:

  • Predictably low latency: Fast mode generates tokens faster and at a more consistent speed than the Standard processing service, even during peak demand.
  • Easy-to-use flexibility: Like Standard processing, Fast mode can be accessed on a flexible, pay-as-you-go basis instead of requiring advance provisioning.

Note: Priority processing was renamed Fast mode on July 30, 2026. You can use either service_tier: priority or service_tier: fast in your API requests.

1 ملین ان پٹ ٹوکن کی قیمت1M ان پٹ ٹوکنز (کیچڈ) کی قیمتفی 1 ملین آؤٹ پٹ ٹوکنز کی قیمتUptime SLA3تاخیر SLA3
GPT-5.6 Sol
طویل سیاق و سباق کو خارج کرتا ہے۱
$۱۰٫۰۰$۱٫۰۰$۶۰٫۰۰۹۹٫۹%۹۹% > ۸۰ ٹوکنز فی سیکنڈ۲
GPT-5.6 Terra
طویل سیاق و سباق کو خارج کرتا ہے۱
$۴٫۰۰$۰٫۴۰$۲۴٫۰۰۹۹٫۹%۹۹% > ۷۰ ٹوکنز فی سیکنڈ۲
GPT-5.6 Luna
طویل سیاق و سباق کو خارج کرتا ہے۱
$۰٫۴۰$۰٫۰۴$۲٫۴۰۹۹٫۹%۹۹% > ۱۰۰ ٹوکنز فی سیکنڈ۲
GPT-5.5
طویل سیاق و سباق کو خارج کرتا ہے۱
$۱۲٫۵۰$۱٫۲۵۰$۷۵٫۰۰۹۹٫۹%۹۹% > ۵۰ ٹوکنز فی سیکنڈ۲
GPT-5.4 mini
طویل سیاق و سباق کو خارج کرتا ہے۱
$۱٫۵۰$۰٫۱۵۰$۹٫۰۰۹۹٫۹%۹۹% > ۱۰۰ ٹوکنز فی سیکنڈ۲
GPT-5.4
طویل سیاق و سباق کو خارج کرتا ہے۱
$۵٫۰۰$۰٫۵۰۰$۳۰٫۰۰۹۹٫۹%۹۹% > ۵۰ ٹوکنز فی سیکنڈ۲
GPT-5.2
$۳٫۵۰$۰٫۳۵۰$۲۸٫۰۰۹۹٫۹%۹۹% > ۵۰ ٹوکنز فی سیکنڈ۲
GPT-5.1
$۲٫۵۰$۰٫۲۵۰$۲۰٫۰۰۹۹٫۹%۹۹% > ۵۰ ٹوکنز فی سیکنڈ۲
GPT-5
$۲٫۵۰$۰٫۲۵۰$۲۰٫۰۰۹۹٫۹%۹۹% > ۵۰ ٹوکنز فی سیکنڈ۲
GPT-5 mini
$۰٫۴۵$۰٫۰۴۵$۳٫۶۰۹۹٫۹%۹۹% > ۸۰ ٹوکنز فی سیکنڈ۲
GPT-5.1 codex
$۲٫۵۰$۰٫۲۵۰$۲۰٫۰۰۹۹٫۹%۹۹% > ۵۰ ٹوکنز فی سیکنڈ۲
GPT-5 codex
$۲٫۵۰$۰٫۲۵۰$۲۰٫۰۰۹۹٫۹%۹۹% > ۵۰ ٹوکنز فی سیکنڈ۲
GPT-4.1
$۳٫۵۰$۰٫۸۷۵$۱۴٫۰۰۹۹٫۹%۹۹% > ۸۰ ٹوکنز فی سیکنڈ۲
GPT-4.1 mini
$۰٫۷۰$۰٫۱۷۵$۲٫۸۰۹۹٫۹%۹۹% > ۹۰ ٹوکنز فی سیکنڈ۲
GPT-4.1 nano
$۰٫۲۰$۰٫۰۵۰$۰٫۸۰۹۹٫۹%۹۹% > ۱۰۰ ٹوکنز فی سیکنڈ۲
GPT-4o
gpt-4o-2024-11-20
gpt-4o-2024-08-06
$۴٫۲۵$۲٫۱۲۵$۱۷٫۰۰۹۹٫۹%۹۹% > ۸۰ ٹوکنز فی سیکنڈ۲
gpt-4o-2024-05-13
$۸٫۷۵$۲۶٫۲۵۹۹٫۹%۹۹% > ۸۰ ٹوکنز فی سیکنڈ۲
GPT-4o mini
$۰٫۲۵$۰٫۱۲۵$۱٫۰۰۹۹٫۹%۹۹% > ۹۰ ٹوکنز فی سیکنڈ۲
o3
$۳٫۵۰$۰٫۸۷۵$۱۴٫۰۰۹۹٫۹%۹۹% > ۸۰ ٹوکنز فی سیکنڈ۲
o4-mini
$۲٫۰۰$۰٫۵۰۰$۸٫۰۰۹۹٫۹%۹۹% > ۹۰ ٹوکنز فی سیکنڈ۲
۱درخواستیں اندازاً 272K پرومپٹ ٹوکنز ہیں
۲پانچ منٹ کے وقفے پر p50 درخواست کی تاخیر کے طور پر حساب کیا جاتا ہے۔ ان صارفین کے لیے جن کے پاس پہلے سے ایسے انٹرپرائز معاہدے موجود ہیں جن میں لیٹنسی SLA فی منٹ بنیاد پر p50 ریکویسٹ لیٹنسی کے طور پر حساب کیے جاتے ہیں، سابقہ SLA اب بھی قابلِ اطلاق ہیں۔
۳یہ صرف Enterprise صارفین کے لیے لاگو ہوتا ہے

یہ کیسے کام کرتا ہے

Customers can direct traffic to Fast mode on a per request basis using the existing service_tier parameter, with the option service_tier = "fast".

Tokens served by Fast mode will be billed on a per-token basis, priced at a premium relative to Standard processing rates.

In addition to being configured at the request level, you can also default a project to Fast mode in Project settings > Default Service Tier: Fast. You can still override per request. Selecting Fast in your project settings is equivalent to selecting Priority.

حدود

  • Fast mode rate limits are shared with other service tiers.
  • In rare cases, rapid increases to your Fast mode Tokens per Minute can lead to hitting ramp rate limits. If you exceed the ramp rate limit, then additional traffic may be sent to Standard processing instead.

Frequently asked questions

قیمتوں کا تعین کرنا

ماڈلز

ریٹ کی حدود

قابلیت اعتماد

پالیسیاں