IREN – 8/20/2025

After several years without new open-weight releases, OpenAI has introduced two new models, gpt-oss-20b and gpt-oss-120b. According to OpenAI, these models reflect a renewed effort to give developers and organizations more flexibility in how advanced AI systems are deployed and operated.
Unlike API-based models, gpt-oss-20b and gpt-oss-120b are released as open-weight models under the Apache 2.0 license. OpenAI states that this license permits commercial use, modification, and self-hosted deployment, subject to standard attribution and notice requirements.
For organizations that need more direct involvement in infrastructure decisions, data handling, and deployment timelines, this release expands the architectural choices available.
As AI workloads grow in size and duration, many organizations are reassessing how much control they want over the systems running their models. Longer training runs, always-on inference, and stricter governance requirements have increased interest in infrastructure that teams can plan, operate, and evolve internally.
OpenAI’s open-weight release changes how organizations can design and operate AI systems over time.
According to OpenAI documentation, deploying gpt-oss models outside of managed APIs allows organizations to avoid per-request or per-token pricing. Instead, costs are driven by infrastructure capacity and operational choices made by the organization. For many teams, this structure can support more consistent budgeting and long-term cost planning.
By hosting models within their own environments, organizations can align deployments with internal security policies and regional data requirements. OpenAI notes that self-hosted deployments allow teams to manage access controls, data residency, and lifecycle decisions directly, rather than relying on external service configurations.
OpenAI states that open-weight models allow teams to fine tune, update, and maintain models on schedules defined by their own operational needs. This approach can reduce disruption caused by externally driven model updates and support longer term application stability.
OpenAI positions these models as reasoning-focused systems designed for practical deployment.
According to OpenAI, gpt-oss-20b and gpt-oss-120b are the company’s first open-weight reasoning models released since earlier GPT generations. Both models are distributed under the Apache 2.0 license, which OpenAI states supports commercial deployment and modification.
OpenAI documentation indicates that gpt-oss-120b can be deployed within approximately 80 GB of GPU memory when using MXFP4 quantization. The gpt-oss-20b model requires approximately 16 GB of GPU memory, making it suitable for development environments and smaller scale inference workloads.
For use cases that require higher throughput or lower latency, OpenAI notes that bfloat16 precision can be used, with corresponding increases in memory and compute requirements.
OpenAI states that gpt-oss models are compatible with commonly used inference and deployment frameworks, including vLLM, Ollama, and Hugging Face. Microsoft has also released a Windows optimized version of gpt-oss-20b using the ONNX Runtime.
Deployment requirements vary depending on workload characteristics and performance objectives.
OpenAI documentation suggests that gpt-oss-120b is typically deployed on high-memory GPUs such as NVIDIA H100 class systems for sustained production workloads. The gpt-oss-20b model can be deployed on systems with lower memory capacity, making it suitable for development and lightweight inference use cases.
These configurations are examples rather than strict requirements. Organizations may choose alternative architectures based on availability, performance goals, and operational constraints.
Open-weight models place greater responsibility on the infrastructure that supports them. Sustained training and inference workloads depend on consistent power delivery, predictable networking behavior, and operational stability over extended periods.
IREN Cloud™ is designed to support dense compute environments through vertically integrated data center infrastructure and grid connected power. IREN Cloud™ operates using NVIDIA’s reference architecture and high-bandwidth InfiniBand networking designed to support large-scale distributed workloads.
These infrastructure characteristics align with the deployment requirements described by OpenAI for gpt-oss class models, particularly for organizations running long-duration or performance-sensitive workloads.
As open-weight AI adoption grows, organizations face new decisions around infrastructure ownership, operational responsibility, and long-term planning. Models such as gpt-oss-20b and gpt-oss-120b expand the deployment options available to builders, while increasing the importance of stable, well planned infrastructure.
IREN Cloud™ provides infrastructure designed for high performance compute workloads that prioritize consistency, scalability, and disciplined execution over time.
Reach out and our team will be happy to help.