The API that powers most of the AI apps you use. GPT-5, Whisper, embeddings, and more — pay per token.
The OpenAI API is the most widely-used LLM endpoint in the world. If you've talked to a chatbot, summarized a doc, or generated an image on the web in 2026, there's a better-than-even chance the request hit OpenAI's servers. The API exposes GPT-5 (flagship), GPT-5-mini (fast/cheap), GPT-5-nano (cheapest), the o-series reasoning models, Whisper (speech-to-text), TTS (text-to-speech), DALL·E and the image models, and embeddings — all behind a single OpenAI-compatible REST surface.
Who it's for: Developers building AI features into products. Startups prototyping. Enterprises needing SOC 2 / HIPAA / zero-data-retention tiers. If you're calling an LLM from code, this is the default starting point — even if you later migrate to Anthropic, open-weights, or a self-hosted stack.
The workhorse endpoint. Send a list of messages, get a streamed or non-streamed completion. Tool calling (function calling), structured outputs (JSON schema enforcement), and vision are all baked in. GPT-5 brings 256K context and a real reasoning mode toggle.
o4 and o5 reason step-by-step before answering. Slower and more expensive per call, but dramatically better on math, code, and multi-step planning. Set a reasoning.effort parameter to trade depth for speed.
Whisper-large-v4 transcribes audio at scale, supports 99 languages, and costs ~$0.004/minute. The single endpoint that killed most paid transcription services below 99% accuracy requirements.
1536-dim (or 3072-dim) vectors for semantic search, RAG, and clustering. $0.02 per 1M tokens. Pairs with pgvector, Pinecone, or any vector DB to build retrieval pipelines in an afternoon.
| Model | Input / 1M tok | Output / 1M tok | Context |
|---|---|---|---|
| GPT-5 | $5.00 | $15.00 | 256K |
| GPT-5-mini | $1.25 | $10.00 | 256K |
| GPT-5-nano | $0.10 | $0.40 | 128K |
| o5 | $15.00 | $60.00 | 200K |
| text-embedding-4 | $0.02 | — | 8K |
| Whisper-large-v4 | $0.004/min | — | — |
Prices are list rates. Volume discounts kick in via the Batch API (50% off, 24-hour turnaround) and committed-use discounts for enterprise. Cached prompt prefixes cut input cost up to 90% on long, repeated system prompts.
If you're shipping an AI feature in 2026 and you have to pick exactly one provider to start with, pick OpenAI. The breadth of models, the maturity of the SDK, and the sheer volume of community code mean you'll hit the ground running. The moment cost or latency starts hurting, route the cheap calls to GPT-5-nano or evaluate Anthropic / open-weights for the heavy lifting — but the OpenAI API is the right default.
Claude API. Best-in-class long context, prompt caching, and tool use. The strong #2.
Run open-source models (Flux, Llama, Whisper) behind a simple API. Pay per second.
Serverless GPU containers. Self-host any open model and pay per second of compute.
Open-source transcription model. Run locally or via the OpenAI API.