Previous All Posts Next

GPT-OSS regulated industries deployments succeed or fail on three decisions: the model you run, the hardware you run it on, and the compliance boundary you draw around both. This guide covers all three for the two open-weight models OpenAI released in the GPT-OSS family. gpt-oss-20b runs within 16 GB of memory, which puts it on a single workstation GPU. gpt-oss-120b runs on a single 80 GB GPU such as an NVIDIA H100. Both ship under the Apache 2.0 license, and both run entirely on infrastructure you own, which is what a compliance lead needs when the data involved is Controlled Unclassified Information, protected health information, or financial records. The sections below cover the model facts, hardware sizing, what the benchmarks do and do not prove, what Section 1532 of the FY2026 NDAA actually restricts, and how to verify the deployment instead of assuming it is fine.

Key takeaways

  • GPT-OSS 20B and 120B are OpenAI's first open-weight models since GPT-2 in 2019, released August 5, 2025 under the Apache 2.0 license.
  • gpt-oss-20b runs within 16 GB of memory; gpt-oss-120b runs on a single 80 GB GPU with its native MXFP4 quantization, and BF16 inference without quantization would need roughly 240 GB.
  • Section 1532 of the FY2026 NDAA (Public Law 119-60) names DeepSeek and High Flyer only, is silent on deployment mode, and never mentions GPT-OSS.
  • No model is HIPAA compliant or CMMC certified on its own; compliance attaches to the system around the model, and every control has to be verified.
  • On SimpleQA the models hallucinate on 78.2 percent (120b) and 91.4 percent (20b) of answers, so retrieval grounding and human review are requirements, not options.

What GPT-OSS 20B and 120B are

OpenAI released gpt-oss-120b and gpt-oss-20b on August 5, 2025 under the Apache 2.0 license. They are OpenAI's first open-weight models since GPT-2 in 2019. Open weight means the weights themselves download to your infrastructure, which is the premise of a private LLM. You can inspect them, pin a specific version, and run them on hardware you control, for as long as you want, without sending a single prompt to an external service.

The two sizes are mixture-of-experts models, which means only a fraction of the total parameters activates on each token. gpt-oss-120b carries 117B total parameters and activates 5.1B per token across 36 layers and 128 experts with 4 active per token. Its context window is 128k tokens. gpt-oss-20b carries 21B total parameters and activates 3.6B per token across 24 layers and 32 experts.

Both models use the harmony chat format and offer three reasoning effort levels: low, medium, and high. They support native agentic tool use and Structured Outputs, which constrains responses to a JSON schema you define. They also expose a fully visible chain of thought, so reviewers can read the reasoning the model produced rather than reverse-engineering it from the answer alone.

The weights download from Hugging Face, and both Ollama and LM Studio run the models locally. NVIDIA publishes NeMo 25.11 fine-tuning recipes for gpt-oss, so a team that needs domain adaptation has a documented path that stays inside its own environment.

Adoption so far supports the deployment pattern this post describes. Early adopters hosted the models on premises for data security reasons, including AI Sweden, Orange Business and Snowflake. Before release, OpenAI's Preparedness Framework adversarial fine-tuning evaluations did not reach the high capability threshold for biological, chemical or cyber misuse, and OpenAI ran a 500,000 dollar Red Teaming Challenge with external experts. That is evidence of a safety process, not a compliance certificate, and the distinction matters for everything that follows.

Hardware requirements for GPT-OSS regulated industries deployments

Hardware sizing is where most private AI projects go wrong first: the model that fits a demo does not fit production concurrency. Start from what each size needs at minimum, then add headroom for the number of simultaneous users and the length of the documents you process.

gpt-oss-120b runs on a single 80 GB GPU, for example an NVIDIA H100, using its native MXFP4 quantization. If you skip quantization and run BF16 inference, the same model needs roughly 240 GB of memory, which means 2 to 4 GPUs. A single 80 GB card serves one steady inference stream for the 120B model; batched or multi-user workloads need more memory or more cards, so plan concurrency before you buy.

gpt-oss-20b runs within 16 GB of memory and is sized for edge devices and workstation GPUs. That makes it the practical choice for a team-level deployment: a single workstation that handles document summarization, drafting, and structured extraction without any network egress.

Attributegpt-oss-120bgpt-oss-20b
ReleasedAugust 5, 2025August 5, 2025
LicenseApache 2.0Apache 2.0
Total parameters117B21B
Active per token5.1B3.6B
Layers3624
Experts128 (4 active)32
Context window128k tokensSee model card
MemoryOne 80 GB GPU with MXFP4; BF16 needs roughly 240 GB (2 to 4 GPUs)Within 16 GB
SimpleQA hallucination78.2 percent91.4 percent

For deployments that outgrow a single box, the reference private AI cluster pattern uses NVIDIA GB10 Grace Blackwell developer nodes with 128 GB of unified memory each, clustered over a QSFP112 400G interconnect; two nodes pool 256 GB over that link for larger models. Single-box inference workstations are sized around RTX 5090, RTX 6000, or H200 class GPUs, and larger models generally want 2 to 4 GPU workstations rather than one card. Those are reference points, not a quote: the build that fits your workload comes out of the scoping prototype, then gets matched to the model size, the concurrency profile and the isolation requirements, in that order.

What the benchmarks show, and what they hide

The published benchmark results are strong, and they are also the wrong thing to buy on alone. Our own LLM benchmarks exist for the same reason: measure on your hardware and your workload. On the OpenAI model card, measured at high reasoning effort: AIME 2025 competition math scores 92.5 for gpt-oss-120b and 91.7 for gpt-oss-20b without tools. GPQA Diamond scores 80.1 and 71.5. MMLU scores 90.0 and 85.3. SWE-Bench Verified scores 62.4 and 60.7. Tau-Bench Retail, an agentic tool-use benchmark, scores 67.8 and 54.8, and gpt-oss-120b reaches a Codeforces competitive programming Elo of 2463 without tools. OpenAI positions gpt-oss-120b as near parity with o4-mini and gpt-oss-20b as comparable to o3-mini on these benchmarks.

For clinical text, HealthBench, a benchmark built on physician-written rubrics, scores 57.6 for gpt-oss-120b against 42.5 for gpt-oss-20b, and both outperform OpenAI's o1 and GPT-4o on that benchmark. HealthBench Hard scores 30.0 and 10.8, and HealthBench Consensus scores 89.9 and 82.6. Those numbers describe performance on rubric-graded text generation. They do not mean the models are safe to answer clinical questions unattended, and no benchmark substitutes for validation against your own cases.

The number that matters most for regulated work is the hallucination rate. On SimpleQA, gpt-oss-120b hallucinated on 78.2 percent of answers with 16.8 percent accuracy, and gpt-oss-20b hallucinated on 91.4 percent of answers with 6.7 percent accuracy, per the OpenAI model card. Read those numbers plainly: asked open-domain factual questions without grounding, these models produce confident wrong answers most of the time. Every regulated deployment needs retrieval grounding and human review for factual work, and any vendor who skips that discussion is selling you the benchmark table, not the system.

The practical implication: treat the benchmarks as evidence the models can do the reasoning your workflows need, then measure hallucination and accuracy on your own document set before anything touches production data.

What the law actually restricts: Section 1532 of the FY2026 NDAA

Compliance teams are asking whether open-weight models are legal to run, and the answer requires reading the actual statute instead of repeating headlines. Section 1532 of the FY2026 National Defense Authorization Act (Public Law 119-60, enacted December 18, 2025; contractor prohibition effective January 17, 2026) provides that no contractor may, during the period of performance of a contract with the Department of Defense, use covered artificial intelligence with respect to the performance of that contract.

Covered artificial intelligence is defined solely as AI developed by the Chinese company DeepSeek, or by High Flyer or entities High Flyer owns, funds, supports, or holds at least a 20 percent stake in. That is the complete list. GPT-OSS is not named anywhere in the statute, and neither is any other model family.

Two boundaries of the statute matter for deployment decisions. First, the restriction is scoped to performance of DoD contracts; it is not a ban on a contractor's purely commercial, non-DoD work, though implementers read "performance" broadly. Second, the statute restricts the model's developer, not its deployment mode. The text never mentions cloud versus on premises, API versus downloaded weights, and it contains no local-hosting exemption. The accurate framing: the ban is on the model regardless of how it is accessed. Hosting DeepSeek on your own hardware does not move it outside the statute, and that point is worth emphasizing because a persistent misconception says self-hosting solves the problem. It does not.

The same act adds Section 6604, which requires removal of the DeepSeek application or service from intelligence community systems, including systems operated by contractors to IC elements. The only government directive that explicitly reached local download and installation is narrower still: a U.S. Navy memo of January 24, 2025 instructed Navy personnel to refrain from downloading, installing, or using the DeepSeek model in any capacity. That is a service-level directive for Navy personnel, not a DoD-wide contractor rule.

Be equally precise about what the statute does not do. No enacted law names Qwen, GLM, Kimi, MiniMax, Tencent, Alibaba or Baidu. Alibaba and Baidu were added to the DoD Section 1260H list on June 8, 2026, but the statutory trigger for extending the contractor ban to listed companies is Defense Department guidance that has not been verified as issued, Alibaba is litigating its designation, and the FY2027 NDAA provision that would add those companies by name is still pending. Claims that "Chinese AI models" are banned as a category overstate enacted law.

CMMC does not fill the gap either, because there is nothing to fill. The CMMC program rule at 32 CFR Part 170 and NIST SP 800-171 contain no AI model origin prohibition of any kind. DFARS 252.204-7012 requires adequate security on covered contractor information systems and says nothing about AI origin. The restriction lives in the statute and in contract performance, not in the framework.

Separate from the legal question, the risk evidence is worth reading. NIST's CAISI evaluation of DeepSeek models, published September 30, 2025, found 94 percent jailbreak success versus 8 percent for U.S. reference models, and 12 times higher agent-hijacking susceptibility. And CISA, NSA and FBI joint guidance from May 22, 2025, "AI Data Security: Best Practices for Securing Data Used to Train and Power AI Systems," treats the AI data supply chain as an attack surface and recommends encryption, digital signatures and provenance tracking. Those are risk findings and best practices, not prohibitions, but they are the reason a defense contractor should document model provenance even for models nobody has restricted.

Running inference inside the SSP boundary

The argument for open-weight models in regulated environments is architectural, and it is about the system security plan, not the model license. An SSP defines a boundary: the systems, networks, and people that handle the data, and the controls that protect it. When inference runs on hardware you own, inside a network segment you control, the model's inputs and outputs stay inside the boundary your SSP already describes. When inference runs on a third-party API, you have introduced a new external system into the data flow, and every control that touches that flow now has to be re-examined: where prompts are logged, how long they are retained, who can access them, and what happens on incident.

This is the same logic that drives CUI enclaving generally. Enclaving CUI into a narrowly scoped boundary keeps the 110 controls of NIST SP 800-171 from applying across the entire enterprise network. An on-premises inference node inside the enclave extends that boundary in a way you can document; a hosted API usually extends it in a way you have to negotiate with someone else's compliance team.

The seven-stage deployment method Petronella Technology Group, Inc. uses makes the sequence explicit: define the data boundary (CUI, PHI, financial records and the frameworks that apply); run a paid scoping engagement that builds an MVP or prototype on your own data, often starting with a data ingestion project; size GPU hardware from what the prototype measured, because data volume changes the answer and a company ingesting 40 TB needs a different build than a smaller one; isolate the cluster on a segmented VLAN or full air gap; deploy open-weight models on an inference stack you control; layer role-based access control, encryption at rest and in transit, and audit logging mapped to framework controls; validate against those controls before production. Hardware is the third decision, not the second: until a prototype has run against your real data, nobody can tell you honestly which GPUs you need. The frameworks the boundary gets mapped to are CMMC Levels 1, 2, or 3, HIPAA, and DFARS 252.204-7012.

Two control details deserve special attention. NIST SP 800-171 requires FIPS-validated cryptography to protect the confidentiality of CUI; encryption that is strong but not validated does not meet the requirement as written, and that includes the storage and transport layers around your model weights and logs. And encrypted CUI is still CUI: putting it in a cloud still requires FedRAMP Moderate or equivalency under DFARS 252.204-7012, which is one more reason defense contractors keep CUI on systems they operate.

For HIPAA compliance, the starting point is the Security Rule's risk analysis requirement at 45 CFR 164.308(a)(1)(ii)(A): you must assess the risks to electronic protected health information and implement measures to reduce them. There is no HHS-issued HIPAA certification for software, so the question is never "is this model certified" but "does the system handling PHI have a current risk analysis, and do its controls address what the analysis found." An on-premises model with role-based access control, encryption, and audit logging is easier to analyze than a third-party API whose data handling you cannot inspect.

Regulated workloads: what to verify before production

The use cases below are real patterns for open-weight models in regulated environments. Each one is framed as what to verify, because a model plus a workflow is not compliant until the controls around it are tested.

HIPAA workflows: clinical and operational text

Discharge summaries, prior-authorization letters, and patient-facing explanations are drafting workloads: the model drafts, a clinician reviews, the record of the review is kept. HealthBench scores suggest gpt-oss-120b is capable of clinically relevant text generation, and that is the right claim. What has to be verified before production is everything around the model: that the retrieval layer only reaches systems the user is authorized to see, that PHI never leaves the segment, that outputs are reviewed against your clinical documentation standards, and that the audit log captures who requested what and who signed off. Verify all of it in a staged rollout before the workflow touches live patient data.

CUI workloads for defense contractors

Proposal drafting, contract summarization, and POA&M tracking are document workloads where the data is CUI and the framework is CMMC. Running gpt-oss on an inference node inside the CUI enclave keeps prompts and outputs on systems your SSP already covers, the pattern behind our AI for defense contractors work. What has to be verified: FIPS-validated cryptography on the storage that holds weights and logs, isolation of the inference segment, access control tied to your existing identity system, and logging that captures the prompts and outputs so the annual assessment has evidence to look at. Also verify the origin question against Section 1532 for any model family that will touch DoD contract work: the covered list today is DeepSeek and High Flyer only.

Financial records and fraud operations

Financial institutions face a different pressure: generative AI is now a fraud vector. FinCEN alert FIN-2024-Alert004, issued November 13, 2024, warned financial institutions of deepfake fraud schemes using generative AI. The same structured-output capability that makes GPT-OSS useful for transaction triage, responses constrained to a JSON schema downstream systems can consume, also has to be verified against your model risk management process, because a structured output that is confidently wrong is worse than an unstructured one. IBM's Cost of a Data Breach Report 2025 puts the United States average breach cost at 10.22 million dollars, the highest of any region, and finds breaches involving shadow AI, unsanctioned AI tools, cost on average 670,000 dollars more. It also finds 97 percent of organizations that reported an AI security incident lacked proper AI access controls, and 63 percent had no AI governance policies. An on-premises model under your access controls is the direct answer to the shadow AI problem, provided the access controls actually exist.

The economics, stated once

Our private AI deployment page puts the screening thresholds this way: at 500,000 or more tokens per day, private deployment breaks even within 6 to 12 months, and at 5 million or more tokens daily, costs run 60 to 80 percent below equivalent API spend. Those thresholds are a screening tool, not a promise, and the token-volume audit that produces your real numbers is part of the scoping work, not an afterthought.

How Petronella Technology Group, Inc. deploys open-weight models

Petronella Technology Group, Inc. designs, builds, and operates private AI clusters for regulated businesses end to end, and delivers the blueprint stack turnkey for regulated teams. The same infrastructure underpins the company's own operations: its 24/7 AI-plus-human hybrid threat analysis stack runs ten-plus production AI agents on the enterprise private AI cluster, supporting managed detection and response for defense industrial base and healthcare clients that cannot send CUI or PHI to a public-cloud SOC. The guardrails are the point: the AI never closes a ticket on its own, never touches production systems without human authorization, and every action is logged for CMMC and HIPAA audit.

The free private AI blueprint walks through eight steps, and its hardening checklist covers network isolation, secrets handling, prompt and access logging, and egress control. Those four items map directly onto the failure modes that show up in assessments: an inference node that can reach the internet, a default credential left on a GPU host, prompts logged to a share nobody controls, and no record of what the model was asked.

On the compliance side, Petronella Technology Group, Inc. is a Cyber AB Registered Provider Organization, RPO #1449, offering CMMC consulting, and every engineer assigned to a defense client holds the CMMC-RP credential. An RPO prepares a contractor for assessment but cannot conduct the assessment or issue a certificate; only a C3PAO can issue a CMMC Level 2 certificate. Craig Petronella founded the company in 2002 and has 30+ years of experience. He holds the CMMC-RP credential, CCNA, CWNE, North Carolina Licensed Digital Forensic Examiner license #604180, and an MIT AI certificate. Framework experience spans CMMC Level 1, 2 and 3 programs, HIPAA, and DFARS 252.204-7012.

A verification checklist before production

Run every item below before an open-weight model touches regulated data. Each one produces evidence your assessor or auditor can examine later.

  1. Confirm the model's developer against Section 1532's definition of covered artificial intelligence, and record the finding. For GPT-OSS the developer is OpenAI, which is not named in the statute.
  2. Record the license terms and pin the exact weights you downloaded, with the download date, so the running version is a documented configuration item in the SSP.
  3. Draw the data boundary before deployment: which systems hold CUI, PHI, or financial records, and which frameworks apply to each, whether that is CMMC Levels 1, 2, or 3, HIPAA, or DFARS 252.204-7012.
  4. Run a scoping prototype on your own data before buying hardware. An ingestion project over a few terabytes and one over 40 TB lead to different builds, and only the prototype tells you which one you are.
  5. Then size hardware to the model and the measured workload: within 16 GB for gpt-oss-20b, one 80 GB GPU for gpt-oss-120b with MXFP4, roughly 240 GB if you run BF16, plus headroom for the concurrency the prototype showed.
  6. Isolate the inference cluster on a segmented VLAN or a full air gap, and control egress from the segment so prompts cannot leave.
  7. Require FIPS-validated cryptography for CUI at rest and in transit; encryption that is strong but not validated does not meet the requirement as written.
  8. Ground factual answers in retrieval from a verified document store, force structured outputs where downstream systems consume them, and log prompts, retrieved passages, outputs, and reviewer decisions.
  9. Keep a human in the loop on decisions that matter, the same way the hybrid SOC pattern does: the AI never closes a ticket on its own and never touches production systems without human authorization.
  10. If DFARS 252.204-7012 applies, rehearse incident reporting before you need it: report cyber incidents affecting covered defense information to DoD within 72 hours through dibnet.dod.mil, which requires a DoD-approved medium assurance certificate obtained in advance, and preserve images of affected systems and monitoring data for at least 90 days.
  11. Validate the whole stack against the controls you mapped in stage one before production, then keep the SSP current and review the POA&M at least quarterly; POA&M items are allowed only within the limits of 32 CFR 170.21, with a 180-day closeout.

Related reading

Frequently asked questions

Is GPT-OSS HIPAA compliant or CMMC certified?

No model is. HIPAA has no certification for software, and CMMC assesses a contractor's systems and practices rather than software products, so a model cannot hold a certificate on its own. Compliance attaches to the system around the model: where CUI and PHI live, which controls protect them, and whether those controls are documented in the SSP and verified before production.

What hardware does gpt-oss-120b need?

One 80 GB GPU, for example an NVIDIA H100, runs gpt-oss-120b with its native MXFP4 quantization. Inference in BF16 without quantization would need roughly 240 GB of memory, which means 2 to 4 GPUs. Plan for concurrency before you pick a host.

Can gpt-oss-20b run on a single workstation?

Yes. gpt-oss-20b runs within 16 GB of memory and is sized for edge devices and workstation GPUs, which makes it a fit for team-level deployments where the workload stays small and the data cannot leave the building.

Does Section 1532 of the FY2026 NDAA restrict GPT-OSS?

No. Section 1532 of the FY2026 National Defense Authorization Act (Public Law 119-60) defines covered artificial intelligence solely as AI developed by DeepSeek or by High Flyer and its affiliates. The statute is silent on deployment mode, and GPT-OSS is not named anywhere in it.

How do you control hallucinations when a model handles regulated work?

Start from the model card numbers: gpt-oss-120b hallucinated on 78.2 percent of SimpleQA answers and gpt-oss-20b on 91.4 percent. Ground answers in retrieval from a verified document store, force structured outputs where systems consume them, and require a human review step for anything that feeds a regulated decision.

Plan a GPT-OSS regulated industries deployment

This post is part of a series on running American open-weight AI inside compliance boundaries. The related post on why Gemma is the right local AI model for CMMC compliance covers Google's model family for the same environments, and DeepSeek DoD restrictions and what defense contractors must do covers the Section 1532 prohibition in depth. For a related field look at the smaller model, see the GPT-OSS 20B voice-agent write-up.

Petronella Technology Group, Inc. is based at 5540 Centerview Drive, Suite 200, Raleigh, NC 27606, and serves clients remotely across all 50 states. Start with the AI solutions hub for the full service map, or go directly to private AI solutions, the private AI deployment service, the free private AI blueprint, the CMMC compliance program, the CUI compliance page, and managed detection and response for teams that cannot send CUI or PHI to a public-cloud SOC.

The first call is a free 30-minute consultation. Call Penny at 919-348-4912 or use the contact form to book it.

About the author: Craig Petronella founded Petronella Technology Group, Inc. in 2002 and has 30+ years of experience in cybersecurity and networking. He holds the CMMC-RP credential, CCNA, CWNE, North Carolina Licensed Digital Forensic Examiner license #604180, and an MIT AI certificate.

Get the Private AI Infrastructure Blueprint

Reference build sheets, a hardening checklist for open-weight model hosting and a data-sovereignty map for running AI on your own hardware.

Get the free guide
Need help implementing these strategies? Our cybersecurity experts can assess your environment and build a tailored plan. Prefer to write? Send us a message.
Call Penny 919-348-4912

About the Author

Craig Petronella, CEO and Founder of Petronella Technology Group
CEO, Founder & AI Architect, Petronella Technology Group

Craig Petronella founded Petronella Technology Group in 2002 and has spent 30+ years professionally at the intersection of cybersecurity, AI, compliance, and digital forensics. He holds the CMMC Registered Practitioner credential issued by the Cyber AB and leads Petronella as a Cyber AB Registered Provider Organization (RPO #1449). Craig is an NC Licensed Digital Forensics Examiner (License #604180-DFE) and completed MIT Professional Education programs in AI, Blockchain, and Cybersecurity. He also holds CompTIA Security+, CCNA, and Hyperledger certifications.

He is an Amazon #1 Best-Selling Author of 15+ books on cybersecurity and compliance, host of the Encrypted Ambition podcast (95+ episodes on Apple Podcasts, Spotify, and Amazon), and a cybersecurity keynote speaker with 200+ engagements at conferences, law firms, and corporate boardrooms. Craig serves as Contributing Editor for Cybersecurity at NC Triangle Attorney at Law Magazine and is a guest lecturer at NCCU School of Law. He serves as a digital forensics expert witness for law firms on matters involving cybercrime, cryptocurrency fraud, SIM-swap attacks, and data breaches.

Under his leadership, Petronella Technology Group has served hundreds of regulated SMB clients across NC and the southeast since 2002, earned a BBB A+ rating every year since 2003, and been featured as a cybersecurity authority on CBS, ABC, NBC, FOX, and WRAL. The company leverages SOC 2 Type II certified platforms and specializes in AI implementation, managed cybersecurity, CMMC/HIPAA/SOC 2 compliance, and digital forensics for businesses across the United States.

CMMC-RP NC Licensed DFE MIT Certified CompTIA Security+ Expert Witness 15+ Books
Related Service
Need Cybersecurity or Compliance Help?

Talk with our cybersecurity experts about your security and compliance needs.

Call Penny 919-348-4912

or send us a message

Previous All Posts Next
Questions about this topic? Talk to our team. Call Penny 919-348-4912 Message us