— Open-Weights LLM

GPT-OSS

Last updated July 21, 2026 · Reviewed by ToolForge Editorial

OpenAI's first open-weights model — 120B parameters, Apache 2.0, runs on a single H100. The model that finally made OpenAI take "open" seriously.

★ 4.4/5 · Apache 2.0 · Since August 2025 · 120B params (MoE)

OpenAI finally ships an open model — and it's good

GPT-OSS dropped in August 2025 with surprisingly little warning and shocked the open-source LLM community. It's a 120B-parameter Mixture-of-Experts model (only ~5B active per token), released under Apache 2.0 — fully commercial-use, no strings. On standard benchmarks it lands between GPT-4o and GPT-4.5, well behind GPT-5 but ahead of every comparably-sized open model at launch. The headline number: it runs inference on a single H100 at ~30 tokens/sec, putting frontier-class quality within reach of anyone with one GPU.

Who it's for: Teams that need frontier-quality reasoning without sending data to OpenAI — healthcare, finance, legal, defense, and anyone with strict data residency. Also the foundation for the current wave of fine-tunes (coding, agents, roleplay) that have flooded Hugging Face since release.

Key features

MoE Mixture-of-Experts

120B total parameters but only ~5B active per token, thanks to 64-expert MoE. This is the architecture trick that makes a "120B model" run on consumer-grade hardware — inference cost is closer to a 7B dense model.

License True Apache 2.0

Not "open weights with caveats" — actual Apache 2.0. Commercial use, modification, redistribution, fine-tuning, all explicitly allowed. No OpenAI Commercial Use Addendum, no "you can't use this to train a competing model" clause.

128K 128K context window

Same 128K context as GPT-4o. Handles long documents, full codebases, and multi-turn conversations without catastrophic forgetting at the tail. RoPE scaling under the hood.

Tool-use Native function calling

Trained with OpenAI's function-calling format, so it drops into existing tool-use pipelines (LangChain, LlamaIndex, custom JSON-schema loops) with zero prompt engineering.

Quantized Runs on 24GB VRAM

The official 4-bit quantized build fits in 24GB of VRAM (single RTX 4090 / A10G). Quality drop is minimal — within 1-2 points on MMLU. This is what unlocked the explosion of self-hosted deployments.

The honest take

✓ What works

  • Best open-weights model at release on most benchmarks — beats Llama 4 Scout and Mistral Large 2 on reasoning
  • Apache 2.0 with no asterisks — actually commercial-use friendly
  • Runs on a single H100 (or 4090 with quantization) — no cluster required
  • Native function calling drops into existing tool-use stacks
  • Huge fine-tune ecosystem already on Hugging Face — coding, roleplay, agent variants

✗ What doesn't

  • Not frontier — clearly behind GPT-5, Claude Opus 4.5, Gemini 3 Pro on hard reasoning
  • Mixture-of-Experts means 120B on disk — not a "laptop LLM" by any stretch
  • No official multimodal version — text only, no vision/audio
  • Limited non-English performance compared to Qwen 3 or DeepSeek V3.2
  • OpenAI's safety tuning baked in — some prompts refused that Llama 4 will answer

Verdict

GPT-OSS is the model that finally ended the "open vs closed" quality gap for most use cases. If you're building a product that needs frontier-ish quality, can't send data to OpenAI, and have one H100 (or a 4090 and tolerance for 4-bit), this is your default. If you're a casual user — just use the GPT-5 API, it's cheaper and better. The biggest impact is downstream: every other open-weights lab now has to beat it, and that's a win for everyone.

💡 Transparency: This review contains affiliate links. If you sign up through our link, we may earn a commission at no cost to you. We only recommend tools we use ourselves. Full disclosure.

Related Tools

Get GPT-OSS

Free · Apache 2.0 · 120B MoE · Runs on a single H100

Download weights →