OpenAI's cheap, fast, multimodal workhorse โ the default model for high-volume and real-time apps.
GPT-4o mini is OpenAI's small, fast, and absurdly cheap multimodal model. It handles text and images, supports a 128K context, and costs roughly 60x less than GPT-4o while beating the old GPT-3.5 on most benchmarks. It's the model most developers actually ship behind production features โ search, classification, extraction, chatbots โ where latency and cost matter more than max reasoning.
Who it's for: Production engineers, startups, and anyone routing high-volume or latency-sensitive traffic that doesn't need the flagship.
Sub-second latency at ~$0.15/M input tokens โ built for scale.
Text and image input in one model.
Long documents and conversations without truncation.
OpenAI SDK compatible, easy to swap with GPT-4o.
If you're shipping at scale, GPT-4o mini is the model you route 80% of traffic to. Save the flagship for the hard 20%.