OpenAI gpt-oss (120B and 20B): What You Need to Know

IREN – 8/20/2025

OpenAI gpt-oss (120B and 20B): What You Need to Know

New open-weight reasoning models


After several years without new open-weight releases, OpenAI has introduced two new models, gpt-oss-20b and gpt-oss-120b. According to OpenAI, these models reflect a renewed effort to give developers and organizations more flexibility in how advanced AI systems are deployed and operated. 


Unlike API-based models, gpt-oss-20b and gpt-oss-120b are released as open-weight models under the Apache 2.0 license. OpenAI states that this license permits commercial use, modification, and self-hosted deployment, subject to standard attribution and notice requirements.  


For organizations that need more direct involvement in infrastructure decisions, data handling, and deployment timelines, this release expands the architectural choices available. 


As AI workloads grow in size and duration, many organizations are reassessing how much control they want over the systems running their models. Longer training runs, always-on inference, and stricter governance requirements have increased interest in infrastructure that teams can plan, operate, and evolve internally. 


Why OpenAI’s gpt-oss release matters 


OpenAI’s open-weight release changes how organizations can design and operate AI systems over time. 


Infrastructure costs can become more predictable 


According to OpenAI documentation, deploying gpt-oss models outside of managed APIs allows organizations to avoid per-request or per-token pricing. Instead, costs are driven by infrastructure capacity and operational choices made by the organization. For many teams, this structure can support more consistent budgeting and long-term cost planning. 


Greater control over deployment and data handling 

 

By hosting models within their own environments, organizations can align deployments with internal security policies and regional data requirements. OpenAI notes that self-hosted deployments allow teams to manage access controls, data residency, and lifecycle decisions directly, rather than relying on external service configurations. 


Model customization on organizational timelines 

 

OpenAI states that open-weight models allow teams to fine tune, update, and maintain models on schedules defined by their own operational needs. This approach can reduce disruption caused by externally driven model updates and support longer term application stability. 


What is new in gpt-oss-20b and gpt-oss-120b 


OpenAI positions these models as reasoning-focused systems designed for practical deployment. 


Open-weight reasoning models 


According to OpenAI, gpt-oss-20b and gpt-oss-120b are the company’s first open-weight reasoning models released since earlier GPT generations. Both models are distributed under the Apache 2.0 license, which OpenAI states supports commercial deployment and modification.


GPU memory requirements 


OpenAI documentation indicates that gpt-oss-120b can be deployed within approximately 80 GB of GPU memory when using MXFP4 quantization. The gpt-oss-20b model requires approximately 16 GB of GPU memory, making it suitable for development environments and smaller scale inference workloads. 

For use cases that require higher throughput or lower latency, OpenAI notes that bfloat16 precision can be used, with corresponding increases in memory and compute requirements. 


Inference framework and tooling support 


OpenAI states that gpt-oss models are compatible with commonly used inference and deployment frameworks, including vLLM, Ollama, and Hugging Face. Microsoft has also released a Windows optimized version of gpt-oss-20b using the ONNX Runtime. 


What organizations need to run gpt-oss models 


Deployment requirements vary depending on workload characteristics and performance objectives. 


OpenAI documentation suggests that gpt-oss-120b is typically deployed on high-memory GPUs such as NVIDIA H100 class systems for sustained production workloads. The gpt-oss-20b model can be deployed on systems with lower memory capacity, making it suitable for development and lightweight inference use cases. 


These configurations are examples rather than strict requirements. Organizations may choose alternative architectures based on availability, performance goals, and operational constraints. 


Why IREN Cloud is architecturally aligned for gpt-oss workloads 


Open-weight models place greater responsibility on the infrastructure that supports them. Sustained training and inference workloads depend on consistent power delivery, predictable networking behavior, and operational stability over extended periods. 


IREN Cloud™ is designed to support dense compute environments through vertically integrated data center infrastructure and grid connected power. IREN Cloud™ operates using NVIDIA’s reference architecture and high-bandwidth InfiniBand networking designed to support large-scale distributed workloads. 


These infrastructure characteristics align with the deployment requirements described by OpenAI for gpt-oss class models, particularly for organizations running long-duration or performance-sensitive workloads. 


What gpt-oss means for AI infrastructure planning 


As open-weight AI adoption grows, organizations face new decisions around infrastructure ownership, operational responsibility, and long-term planning. Models such as gpt-oss-20b and gpt-oss-120b expand the deployment options available to builders, while increasing the importance of stable, well planned infrastructure. 


IREN Cloud™ provides infrastructure designed for high performance compute workloads that prioritize consistency, scalability, and disciplined execution over time. 


Have questions about this post?

Reach out and our team will be happy to help.