Gemma CMMC compliance starts with a practical question: how does a defense contractor run a capable AI model on CUI without moving that data outside the boundary its security plan already covers? The answer is to run Google DeepMind's Gemma 4 on your own hardware under the Apache 2.0 license, inside the boundary, with access and audit controls mapped to the same NIST SP 800-171 requirements that govern everything else. This post covers the model, the license, the hardware, the CMMC Level 2 fit, and what Section 1532 of the FY2026 NDAA does and does not restrict.
Key takeaways
- Gemma 4 is Google's current open-weight model family, released March 31, 2026 in E2B, E4B, 26B A4B and 31B Dense sizes, with a 12B Unified size added June 3, 2026, under the Apache 2.0 license.
- Running Gemma 4 locally keeps prompts, documents, and outputs on hardware the contractor controls, which keeps CUI inside the System Security Plan boundary instead of opening a new data flow to an external service. Our on-premise AI service is built on that principle.
- The unquantized bfloat16 weights fit a single 80GB NVIDIA H100 GPU, and quantized versions run natively on consumer GPUs, from phones and Raspberry Pi to workstation cards.
- Petronella Technology Group, Inc. deploys open-weight models with a seven-stage method: define the data boundary; run a paid scoping engagement that builds a prototype on the client's own data; size the GPU hardware from what the prototype measured; isolate the cluster; deploy the model; layer access and audit controls; and validate before production.
- Section 1532 of the FY2026 NDAA (Public Law 119-60) defines covered artificial intelligence solely as AI developed by DeepSeek, or by High Flyer or entities High Flyer owns, funds, supports, or holds at least a 20 percent stake in. Gemma appears nowhere in that definition.
Gemma CMMC compliance starts with the model and the license
Gemma 4 is the current generation of Google's open model family, built from the same research and technology as Gemini 3. Google DeepMind, Google's AI research organization under Alphabet, released it on March 31, 2026 in four sizes, E2B, E4B, 26B A4B and 31B Dense, and calls the family "our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows." A 12B Unified size followed on June 3, 2026, bringing the family to five sizes. Google puts Gemma downloads over 400 million, with more than 100,000 variants in the ecosystem.
The license is Apache 2.0
Gemma 4 is released under Apache 2.0, which Google's launch post describes as "a commercially permissive Apache 2.0 license," and the Hugging Face and Kaggle model cards carry the same line. Gemma 3, the prior generation, shipped under the Gemma Terms of Use instead, so teams that evaluated Gemma 3 will find the review materially simpler this time.
What Apache 2.0 simplifies for a legal review is concrete. Instead of a vendor-specific terms-of-use document with its own prohibited-use policy, your legal team reviews a standard open-source license, the same class the rest of your software stack already uses. There is no click-through gate on the weights and no separate policy document to track as a procurement condition. We compared this generation on our own hardware in Mistral 3.2 and Gemma 4 benchmarked on four GPUs, and our Gemma 2 overview covers the prior lineage.
After the license review, the weights download and run on infrastructure you control. Google names Hugging Face, Kaggle, and Ollama as download sources, with day-one support across Transformers, vLLM, llama.cpp, MLX, Ollama, and LM Studio. Running on your own stack means no prompt or document ever leaves your network to generate a response.
Sizes, context window, and what the model handles
Gemma 4 ships in five sizes: E2B, E4B, 12B Unified, 26B A4B, and 31B Dense. The "E" stands for effective parameters: those models use Per-Layer Embeddings, presenting an effective 2 billion and 4 billion parameter footprint during inference. The 26B A4B model is a Mixture-of-Experts model that activates only 3.8 billion of its 25.2 billion total parameters, so it runs almost as fast as a 4B model while carrying a 26B model's knowledge. The 31B Dense model maximizes raw quality and is the foundation Google recommends for fine-tuning.
Context and modality come from the model card. Context length is 128K tokens for E2B and E4B and 256K tokens for 12B Unified, 26B A4B, and 31B Dense. All five sizes take text and image input; E2B, E4B, and 12B Unified also take audio input natively, up to 30 seconds per clip. For CUI work, a 31B deployment reads a long technical document in one request without splitting it, and a small E2B deployment runs entirely on-device: the E2B and E4B models run completely offline with near-zero latency on phones, Raspberry Pi, and NVIDIA Jetson Orin Nano.
Why local AI keeps CUI inside the SSP boundary
CMMC Level 2 covers Controlled Unclassified Information: all 110 practices of NIST SP 800-171 Rev 2, organized into 14 domains, with a C3PAO assessment for most programs. The moment a contractor pastes CUI into a hosted AI service, the data crosses into someone else's system, which then needs the controls, the agreements, and in many cases the FedRAMP standing the clause requires. Local inference removes that crossing: prompts, documents, and outputs stay inside the boundary the System Security Plan already describes, and the only new thing to assess is the AI stack itself.
The same logic appears in DoD's own guidance. DFARS 252.204-7012 requires adequate security, defined as the 110 controls of NIST SP 800-171, on every covered contractor information system, and requires a cloud service provider handling covered defense information to meet security requirements equivalent to the FedRAMP Moderate baseline. A public AI endpoint that aggregates prompts in cloud logs sits outside that structure; an on-premises inference server sits inside it, under the contractor's own policies, segmentation, and logging. Enclaving CUI into a narrowly scoped boundary is standard practice, and an AI host inside that enclave inherits the same treatment as the rest of it: defined access, monitored traffic, and audit logging.
The boundary grows when you add a model, so describe it
When an organization adds an AI model, the boundary the SSP describes grows to include the model's runtime, the checkpoint storage, and any inference caches. Treat the inference server as part of the scoped system so the required controls are applied deliberately instead of letting the stack blend into unrelated business applications. If a requirement is not yet met, it belongs on a POA&M within the limits 32 CFR 170.21 allows. A model file cannot be compliant by itself; CMMC scopes systems, not weights. The deployment is what gets assessed, and the deployment is what you control when the model runs locally.
What the seven-stage method looks like for a Gemma 4 deployment
Petronella Technology Group, Inc. deploys private AI with a seven-stage method, and the order is the point: define the data boundary; run a paid scoping engagement that builds an MVP or prototype on the client's own data, often starting with a data ingestion project; size GPU hardware from what the prototype measured; isolate the cluster; deploy open-weight models on an inference stack you control; layer access and audit controls; and validate. Hardware is never sized before the prototype has run, because a spec table cannot measure your document set. Applied to Gemma 4:
- Define the boundary: name the data the model will touch and the frameworks that apply, usually CMMC Levels 1, 2, or 3, HIPAA, and DFARS 252.204-7012.
- Run the paid scoping engagement: build the prototype on the client's own data, often as a data ingestion project, and record context lengths, concurrency, and the accuracy the task demands.
- Size the GPU hardware from what the prototype measured: a 31B Dense deployment serving a 256K context under concurrency makes a different hardware case than an E4B field utility, and a client ingesting 40 TB of data may need different hardware than a smaller company.
- Isolate the cluster: segmented VLAN or full air gap, so the inference host talks only to what the boundary allows.
- Deploy the open-weight model on an inference stack you control, not on a vendor's endpoint.
- Layer role-based access control, encryption at rest and in transit, and audit logging mapped to framework controls: role-based access control serves 3.1 Access Control, audit logging serves 3.3 Audit and Accountability, and VLAN isolation serves 3.13 System and Communications Protection.
- Validate against those controls before production, and document the result in the SSP.
The framework mapping matters because it shows the AI stack was not bolted on. NIST SP 800-171 requires FIPS-validated cryptography to protect the confidentiality of CUI; encryption that is strong but not validated does not meet the requirement as written, and that applies to the volumes holding the model and the caches as much as to any other component.
Hardware sizing: what Google publishes, and what our benchmarks measured
Google's launch post gives the headline fact: the unquantized bfloat16 weights of the larger models fit efficiently on a single 80GB NVIDIA H100 GPU, and quantized versions run natively on consumer GPUs. At the small end, the E2B and E4B models run completely offline with near-zero latency on phones, Raspberry Pi, and Jetson Orin Nano.
Google's QAT post, published June 5, 2026, adds the mobile-format checkpoint: "we've reduced the memory footprint of Gemma 4 E2B to 1GB." The QAT weights ship in Q4_0 GGUF formats for llama.cpp and compressed tensors for vLLM. One caveat a sizing conversation needs: quantized figures cover loading the model weights, and running the model also requires VRAM for the KV cache, which grows with context length. Plan the card for weights plus cache, not weights alone.
Numbers on our own hardware are in our May 2026 benchmark post, which measured Gemma 4 31B Dense across vLLM and Ollama with a 2,000-token cybersecurity briefing prompt and 2,000 output tokens, single user. The best result was 63.5 tokens per second on the RTX PRO 6000 Blackwell workstation with 192 GB of DDR5 running the Ollama Q4_K_M build, and 37.7 tokens per second under vLLM on the same host. The GB10 Grace Superchip kit delivered 6.6 tokens per second under vLLM and 10.0 under Ollama, and a 32-way concurrency stress test reached 2,173 tokens per second aggregate. Our LLM benchmarks page carries the full method and per-model tables.
For larger footprints, the reference private AI cluster runs on GB10 Grace Blackwell nodes with 128GB unified memory each, clustered over a QSFP112 400G interconnect to pool 256GB for larger models, and single-box workstations are sized around RTX 5090, RTX 6000, or H200 class GPUs. The 31B Dense model at BF16 on a single 80GB H100 fits when a contractor already operates data-center cards.
Why Gemma 4 instead of GPT-OSS 20B or 120B
Both are good American open-weight choices, and both are Apache 2.0, so this comparison is not about licenses or legal exposure. Our post on GPT-OSS for regulated industries covers the OpenAI models in depth. The question is which shape fits which deployment.
OpenAI released gpt-oss-120b and gpt-oss-20b on August 5, 2025 under Apache 2.0. gpt-oss-20b runs within 16 GB of memory on a single workstation GPU; gpt-oss-120b runs on a single 80 GB GPU such as an NVIDIA H100 using its native MXFP4 quantization. Both are mixture-of-experts models: gpt-oss-120b carries 117B total parameters and activates 5.1B per token, and gpt-oss-20b carries 21B total parameters and activates 3.6B per token. Both handle text only.
| Attribute | Gemma 4 | GPT-OSS 20B / 120B |
|---|---|---|
| Released | March 31, 2026 (12B Unified June 3, 2026) | August 5, 2025 |
| License | Apache 2.0 | Apache 2.0 |
| Developer | Google DeepMind | OpenAI |
| Sizes | E2B, E4B, 12B Unified, 26B A4B (MoE), 31B Dense | 20B (MoE), 120B (MoE) |
| Smallest footprint | E2B with the QAT mobile format at a 1GB memory footprint; offline on phones, Raspberry Pi, Jetson Orin Nano | gpt-oss-20b within 16 GB of memory |
| Largest single-GPU fit | 31B Dense BF16 on one 80GB H100 | gpt-oss-120b with MXFP4 on one 80 GB GPU |
| Context window | 128K tokens (E2B, E4B); 256K tokens (12B Unified, 26B A4B, 31B) | 128k tokens (gpt-oss-120b) |
| Input modalities | Text and image on all sizes; native audio on E2B, E4B, 12B Unified | Text only |
Four differences decide most defense deployments. Size range: Gemma 4 spans from a 1GB-footprint E2B on a phone to the 31B Dense model on a single H100, so one family covers the field laptop through the enclave server under one license review; with GPT-OSS the small end is a 16 GB workstation model and there is nothing below it. Modality: every Gemma 4 size takes images and E2B, E4B, and 12B Unified take audio natively, which covers document scans, redacted screenshots, and voice notes, while GPT-OSS is text-only. Context: the larger Gemma 4 sizes carry 256K tokens against 128k on gpt-oss-120b, which matters when a request must hold a full contract set. And there is no license trade-off: the choice is made on deployment shape, not legal terms.
GPT-OSS is the better pick in specific cases, and we say so. If your team has standardized on OpenAI tooling and workflows, adopting the models OpenAI publishes directly reduces integration work. If the workload is pure text reasoning, such as code analysis or structured drafting with no document images and no audio, gpt-oss-120b's single-GPU deployment and documented agentic tooling are a strong fit. Both run inside the same compliance boundary, and the boundary work in the rest of this post applies to either family.
Gemma CMMC compliance and Section 1532 of the FY2026 NDAA
Section 1532 of the FY2026 NDAA (Public Law 119-60, enacted December 18, 2025; the prohibition effective January 17, 2026) provides that no contractor may, during the period of performance of a contract with the Department of Defense, use covered artificial intelligence with respect to the performance of that contract. Covered artificial intelligence is defined as AI developed by DeepSeek, or by High Flyer or entities High Flyer owns, funds, supports, or holds at least a 20 percent stake in. The statute restricts the model's developer, not its deployment mode: the text never mentions cloud versus on premises, so a covered model falls under the prohibition the same way either way. Gemma is not named in the enacted law, so it is not subject to the Section 1532 restriction.
Two scope points keep this accurate. The prohibition is scoped to DoD contract performance, not to everything a contractor does. And no other Chinese-origin AI model is named in the enacted law: Qwen, GLM, Kimi, MiniMax, Alibaba, Baidu and Tencent appear nowhere in it. Section 6604 of the same act separately requires removal of the DeepSeek application from intelligence community systems. CMMC itself contains no AI-model rule: the CMMC program rule, 32 CFR Part 170, maps assessment levels to NIST SP 800-171, and NIST SP 800-171 contains no AI-model-origin prohibition.
For a contractor choosing a local model, the consequence is straightforward. A model the law does not name keeps you clear of the statutory prohibition, and holding the weights in-house keeps the whole inference path inside your own boundary. Gemma is developed by Google DeepMind, a subsidiary of Alphabet, and is not named in any enacted US restriction on Chinese-developed AI.
Why we do not recommend larger Chinese models like GLM-5.3 for defense work
Start with what the law actually says, because the distinction matters. No enacted law prohibits a defense contractor from using GLM. Section 1532 names only DeepSeek and High Flyer, and GLM's developer, Zhipu AI, appears nowhere in the enacted statute. We do not say GLM is banned, because it is not, and we do not say it cannot be used, because contractors make their own decisions. What follows are the citable reasons we steer DoD contractors away from it as a matter of risk management.
The first is export control. Zhipu AI, which does business as Z.AI, has been on the Commerce Department's BIS Entity List since January 16, 2025, when ten Zhipu-family entities were added under Federal Register citation 90 FR 4619 with a presumption of denial. The listing requires a license to export, reexport, or transfer items subject to the Export Administration Regulations to a listed entity as purchaser, end-user, or consignee, and extends to entities 50 percent or more owned by listed parties. Downloading publicly released model weights and running them on your own hardware involves no export to Zhipu, so self-hosting is not prohibited by the listing; the practical exposure is the hosted-API channel, where sending controlled technology into Zhipu's hosted service is an export a license application would presumptively not survive.
The second is that the legal picture is moving toward Zhipu, not away from it. The pending FY2027 Senate bill, S. 4784, would amend Section 1532's definition of covered artificial intelligence to add Zhipu AI by name; it has been reported out of committee and is not law. A model the next NDAA may name is a model you may have to rip out mid-contract, and that operational risk lands on the program, not on the model vendor.
The third is the federal security record. A joint CISA, NSA, and FBI advisory, AA26-251A, published September 8, 2026, names Z.AI among the China-based AI companies that extracted billions of tokens from U.S. frontier AI models in industrial-scale distillation campaigns. NIST's CAISI assessment of GLM-5.2, published July 17, 2026, found that GLM-5.2's safeguards allow assistance with agentic cyber exploit development, and noted that safeguards on any open-weight model can be circumvented when self-hosted. Neither document prohibits anything by itself, but together they are the record a chief information security officer reads before approving a model and an assessor or a prime's flow-down question will ask about after an incident.
The fourth is fit for the deployment a defense contractor actually runs. GLM-5.3 is a 744B total parameter model with 40B active, and its FP8 checkpoint is roughly 744 GB, which puts local serving on a multi-GPU H200-class node. Its license is a custom GLM-5.3 license, not MIT or Apache 2.0. Compare the shape: Gemma 4 31B runs on a single 80GB H100 at BF16, in a size class a mid-size contractor can buy, power, and cool inside a CUI enclave.
None of these reasons is a legal prohibition, and the honest framing is risk management: the flow-down clauses primes add, the questions assessors ask, and the cost of unwinding a dependency when the law or the advisory record moves. Gemma 4 and GPT-OSS give you the open-weight capability without that exposure. For the enacted-law mechanics, see our post on DeepSeek and the FY2026 NDAA.
Operating the AI stack as part of the CMMC program
A local model is one component. The program work around it is what an assessor sees:
- Access control: role-based permissions decide who can submit prompts, who can view outputs, and who can change the model configuration, which is the 3.1 Access Control family applied to the AI stack.
- Audit logging: every inference request is logged with a record of who asked what, when, and with which model version, serving the 3.3 Audit and Accountability family. Our page on creating and retaining system audit logs covers the requirement.
- Communications protection: the inference host sits on a segmented VLAN or an air-gapped segment, serving the 3.13 System and Communications Protection family.
- Encryption: FIPS-validated cryptography protects CUI at rest and in transit, as NIST SP 800-171 requires.
- Incident reporting: if a reported incident occurs, DFARS 252.204-7012 requires preserving images of affected systems and monitoring data for at least 90 days, so AI request logs belong in that preservation window.
- Change discipline: a new model checkpoint is a change to the environment. Record the model family, the developer, the license, and the exact checkpoint you downloaded, and check the developer against Section 1532's definition when you record it. The Gemma 4 weights are distributed by Google DeepMind through Hugging Face, Kaggle, and Ollama, so the provenance record is short.
Petronella Technology Group, Inc. runs its own private AI stack in production. The company's private AI cluster and 24/7 AI-plus-human hybrid threat analysis stack underpin managed detection and response for defense industrial base and healthcare clients that cannot send CUI or PHI to a public-cloud SOC. In that hybrid SOC the AI never closes a ticket on its own, never touches production systems without human authorization, and every action is logged for CMMC and HIPAA audit. The wider service for primes and subcontractors is on our AI for defense contractors page, and compliance documentation runs on ComplianceArmor®, the company's platform for SSP authoring, POA&M tracking, and evidence repository organization.
What a deployment engagement looks like
A typical CMMC Level 2 readiness engagement runs 12 to 14 weeks for a mid-size contractor: Discovery in week 1, Gap Analysis in weeks 2-4, Remediation Sprint in weeks 5-12, and C3PAO Readiness Handoff in weeks 13-14. A Gemma 4 deployment slots into that program: the boundary is defined first, then the paid scoping engagement builds a prototype on the contractor's own data, often a data ingestion project, so the workload is measured before anyone quotes a GPU. Hardware sizing follows the prototype, because a 31B checkpoint serving a 256K context under concurrency is a different build than an E4B field utility. During remediation the inference host is built, the chosen Gemma 4 checkpoint is installed, and the audit logging is pointed at the collection the SSP already describes.
Two roles matter and they are not the same. Only a C3PAO can issue a CMMC Level 2 certificate; the C3PAO submits results to eMASS and the certificate is stored in SPRS. An RPO prepares a contractor but cannot conduct the assessment or issue a certificate. Petronella Technology Group, Inc. is a Cyber AB Registered Provider Organization, RPO #1449, offering CMMC consulting, and does not perform Level 2 assessments for clients it has prepared, because Cyber AB independence rules prohibit it. Every Petronella engineer assigned to a defense client holds the CMMC Registered Practitioner (CMMC-RP) credential: Blake Rea, Justin Summers, Jonathan Wood and Craig Petronella.
Keeping the model local does not answer the assessor's questions for you, but it keeps every one of them inside the environment you document, instead of adding a third-party service whose controls you do not own.
How this post fits the series
This post is part of a series on open-weight AI under the new rules. The second post, GPT-OSS for regulated industries, takes up the OpenAI open-weight models, and the third, DeepSeek and the FY2026 NDAA: what DoD contractors need to know, walks through what the Section 1532 prohibition means for contractors using DeepSeek. For the broader context, the AI Solutions hub and the private AI solutions page cover the program side, the private AI deployment page and the private AI blueprint cover the stack, and the CMMC compliance page and the CUI page cover the boundary the AI stack must protect. Managed detection and response is covered on the Managed Detection and Response page.
Call to action
If you are evaluating Gemma 4 inside a CMMC program, start with the boundary, not the model. Call Penny at 919-348-4912, or use the contact form. The first call is a free 30-minute consultation run by a CMMC Registered Practitioner, and engagements are delivered remote-first across all 50 states.
Author: Craig Petronella, CMMC-RP - Founder, Petronella Technology Group, Inc. Craig Petronella holds CMMC-RP, CCNA (Cisco Certified Network Associate), CWNE (Certified Wireless Network Expert), NC Licensed Digital Forensic Examiner license #604180, and an MIT AI certificate; he has 30+ years of experience and founded the company in 2002.
Petronella Technology Group, Inc., 5540 Centerview Dr Suite 200, Raleigh NC 27606, 919-348-4912.
Related reading
- Self-Hosted LLM Benchmarks for Private AI
- OpenCode, Antigravity, and Gemma 4: Safe AI Coding Tools
- Why a Private AI Appliance Secures Your Firm's Data
- Mistral 3.2 and Gemma 4 Benchmarked on 4 GPUs
- Jan AI: Free Local LLM App Review
FAQ
What makes Gemma 4 a defensible choice for CUI workloads?
It is not the model file; it is the deployment. Gemma 4 is Google's current open-weight model family, released March 31, 2026 under the Apache 2.0 license. Running it on infrastructure the contractor controls keeps prompts, documents, and outputs inside the boundary the System Security Plan already describes, so the AI stack is assessed under the same 110 NIST SP 800-171 practices as the rest of the environment. A model file cannot be compliant by itself; CMMC scopes systems, not weights.
What hardware does Gemma 4 need?
Google's launch post states that the unquantized bfloat16 weights fit a single 80GB NVIDIA H100 GPU, and that quantized versions run natively on consumer GPUs. The QAT checkpoints released June 5, 2026 include a mobile format that reduces the E2B footprint to 1GB. Weights are not the whole card: the KV cache needs additional VRAM that depends on context length, which is why hardware gets sized from a measured prototype rather than from a spec table.
Does Section 1532 of the FY2026 NDAA restrict Gemma?
No. Section 1532 of the FY2026 NDAA (Public Law 119-60, enacted December 18, 2025; prohibition effective January 17, 2026) prohibits any contractor from using covered artificial intelligence with respect to the performance of a contract with the Department of Defense, and covered artificial intelligence is defined as AI developed by DeepSeek, or by High Flyer or entities High Flyer owns, funds, supports, or holds at least a 20 percent stake in. Gemma is not named in the enacted law, and the statute is silent on deployment mode, so the deployment question does not change the answer.
How does Gemma 4 compare with GPT-OSS or GLM-5.3 for defense work?
Gemma 4 and the OpenAI GPT-OSS models are both Apache 2.0 open-weight families, so neither raises a license or origin barrier; the choice between them is about size range, input modalities, and context length. GLM-5.3 is different: its developer Zhipu AI has been on the Commerce Entity List since January 16, 2025, a federal advisory named Z.AI in distillation campaigns, and pending Senate legislation would add Zhipu AI to the statutory covered list. Enacted law still names only DeepSeek and High Flyer, so GLM is not legally banned today, but we treat it as an avoidable risk for DoD contract work.
What does Petronella Technology Group, Inc. actually do for a Gemma engagement?
The company designs, builds, and operates private AI clusters for regulated businesses end to end. The engagement defines the data boundary first, runs a paid scoping engagement that builds a prototype on the client's own data, sizes the GPU hardware from what the prototype measured, isolates the cluster, deploys the model, layers the access and audit controls, and validates against the framework before production. ComplianceArmor®, the company's compliance documentation platform, automates SSP authoring, POA&M tracking, and evidence repository organization, and ongoing monitoring runs through managed detection and response. The first call is a free 30-minute consultation.
Free, practical, and specific to regulated environments. We will email it to you.
No spam. Unsubscribe anytime.