If your hard requirement is on-prem and no data leaving your network, I’d shortlist open-weight models you can self-host, rather than any API-hosted model.
My default recommendation
Mistral Small 4 is a strong starting point if you have serious GPU capacity. It’s an Apache 2.0 open-weight model with 119B total / 6.5B active parameters, and supports reasoning, coding, tool use, and multimodal workloads.
For a smaller deployment, Qwen3-30B-A3B is particularly attractive: it is an open-weight MoE model with 30B total / 3B active parameters, 128K context, and can be served with vLLM or SGLang.
Qwen3-30B-A3B or a larger model depending on workload
Multiple 80 GB GPUs
Mistral Small 4
Very large GPU cluster
Mistral Large 3 / Qwen3 235B-A22B
For production serving, vLLM is a sensible default; Mistral and Qwen both document self-hosted deployment with it.
One important distinction: "on-prem" doesn't automatically mean "no data leaves the network." Keep the inference server, model weights, embeddings/RAG store, telemetry, MCP tools, and any connectors inside your network. Mistral's own offline documentation, for example, explicitly recommends disabling telemetry/background network calls when operating fully offline.
If you tell me (1) your GPU(s)/RAM, (2) number of concurrent users, and (3) what you're doing with the model—RAG, coding, agents, document processing, general chat, etc., I can narrow this to a specific model + quantization + serving stack.
For an on-premises, fully offline (air-gapped) deployment where zero data leaves your network, you want to look at open-weights models . The right choice depends heavily on your available hardware (VRAM) and primary use case.
The top enterprise-grade options are categorized by capability and hardware footprint:
1. General Purpose & High Performance (Best Overall)
Llama 3.3 (70B) or Llama 3.1 (70B) via Meta Llama
Best for: General text, complex summarization, natural dialogue, and broad ecosystem tool integration.
Hardware required: Needs multi-GPU or high-end workstation setups (e.g., 2 × or 4 × RTX 3090/4090 or enterprise accelerators like A100/H100 depending on quantization).
Qwen (72B or 32B/14B variants) via Alibaba Qwen
Best for: Exceptional multilingual performance, heavy logic, and instruction-following. Many Qwen variants also use very permissive licensing (Apache 2.0).
Hardware required: Scales well from 72B down to efficient mid-sized options.
2. Heavy Reasoning & Logic
DeepSeek-R1 (Distilled versions, e.g., 32B or 70B) via DeepSeek AI
Best for: Advanced problem-solving, math, and transparent step-by-step reasoning (showing chain-of-thought in <think> tags).
Hardware required: The 32B distilled version fits comfortably on a single high-end consumer/workstation GPU (like a 24GB VRAM card when quantized).
3. Coding & Technical Tasks
Qwen-Coder (e.g., 32B) or Codestral via Mistral AI
Best for: Local code generation, auto-completion, and developer tool integration.
Hardware required: Mid-to-high tier single GPU setups (16GB–24GB VRAM with moderate quantization).
4. Lightweight / Edge Deployment (Low Resource)
Llama 3.2 (8B / 3B) or Google Gemma (9B/4B) via Google Gemma
Best for: Running locally on standard developer laptops or single-GPU office workstations without melting the cooling fans.
Hardware required: Standard workstation or consumer GPU with 8GB–16GB VRAM.
Recommended Serving Stack for On-Prem
To actually run these models securely behind your firewall, use local inference engines packaged for containerized or bare-metal deployment:
Ollama or vLLM : Excellent for spinning up API-compatible local endpoints inside an air-gapped network.
Quantization (GGUF or EXL2) : Use quantized versions (like 4-bit or 8-bit) to drastically reduce the VRAM footprint with minimal loss in intelligence.
To help narrow this down to a specific model and hardware spec , let me know:
What GPU hardware or server budget do you have available on-prem?
Is your primary use case RAG (searching internal documents), coding , or general chat/automation?
If “no data leaving our network” is a hard requirement, I’d use an open-weight model deployed entirely inside your infrastructure, rather than an API-hosted model. Current on-prem options include Qwen, Llama, Mistral, DeepSeek, and Gemma.
My default recommendation
Qwen 3.5 would be my first model to evaluate for a new enterprise on-prem deployment, particularly if you want a good combination of general reasoning, coding, multilingual support, and permissive licensing. Current Qwen 3.5 variants span relatively small models through very large MoE models, so you can match the deployment to your GPU budget.
A practical shortlist:
Situation
Model family I'd evaluate
General enterprise assistant / RAG
Qwen 3.5
Strong general-purpose alternative
Llama
Coding / technical workloads
Qwen / DeepSeek
Maximum capability with a multi-GPU cluster
DeepSeek / large Qwen MoE
Smaller hardware footprint
Qwen 3.5 27B-class or similar mid-size model
The important caveat is that model choice should follow your hardware and workload, not just benchmark rankings. For example, a 30–70B-class model can be a sensible enterprise sweet spot, while very large MoE models can require multi-GPU or multi-node infrastructure.
For a genuinely isolated deployment
I'd architect it as:
Your applications → internal API → vLLM/TensorRT-LLM → model weights → internal RAG/vector DB
with:
no external API calls
outbound network blocked at the infrastructure level
model weights downloaded/imported through your controlled software-supply-chain process
embeddings and reranking models also hosted internally
internal audit logging
pinned model/container versions and checksums
RAG access controls enforced before documents reach the model
That's important because “self-hosted model” doesn't automatically mean “no data leaves the network.” The entire inference and retrieval stack needs to be inside the boundary.
If you tell me your GPU setup (e.g. 2×H100, 4×A100, 8×H200, RTX 6000, etc.), primary workload, expected concurrent users, and context length, I can narrow this to a specific model/quantization and give you the approximate hardware and throughput.
For an on-premise deployment where zero data leaves your network , you should run top-tier open-weights/open-source models hosted on your own local infrastructure.
The ideal model choice depends heavily on your available hardware (specifically VRAM/GPUs) and your primary use case.
Recommended Open-Weights Models by Use Case
Llama Series (Meta): Llama 3.3 / Llama 4 variants
Best For: General-purpose enterprise tasks, chatting, text summarization, and broad out-of-the-box instruction following.
Why: Excellent balance of performance and efficiency, with strong community backing and easy enterprise integration.
Qwen Series (Alibaba): Qwen2.5 / Qwen3 series
Best For: Multilingual support, heavy coding tasks, math, and structured data extraction.
Why: Qwen models punch well above their weight class on benchmarks and handle diverse language and code structures exceptionally well.
DeepSeek Series: DeepSeek-V3 / V4
Best For: Advanced reasoning, complex logic, and rigorous coding workflows.
Mistral / Codestral: Mistral Nemo or Mixtral variants
Best For: Lightweight, high-throughput, and low-latency deployments on tighter hardware budgets.
Why: Highly efficient architectures (including Mixture of Experts) that deliver great performance with fewer resource constraints.
Sizing and Hardware Rule of Thumb
Small Models (7B – 14B parameters): Can run on a single high-end workstation or consumer GPU (e.g., RTX 4090 / Mac Studio). Great for basic tasks and fast responses.
Medium Models (32B – 70B parameters): Require enterprise-grade accelerators (e.g., NVIDIA A100/H100 or multi-GPU setups). Necessary for nuanced reasoning, coding, and reliable instruction-following.
Recommended Local Serving Stacks
To keep everything strictly on-prem without telemetry or external calls, pair your chosen model with a local runtime:
Ollama: The easiest tool to spin up and manage local open models with a clean CLI and local API.
vLLM or TGI (Text Generation Inference): Best for production-grade, high-concurrency on-premise serving with optimized memory management (PagedAttention).
To help narrow down the exact model size and architecture for your setup, could you share:
What hardware/GPUs you have available on-prem?
Your primary use case (e.g., internal chatbot, automated coding, document search/RAG)?
For a true air-gapped, on-premise deployment where zero data leaves your network , you need to use top-tier open-weight/open-source models running on your own local infrastructure.
The ideal model depends heavily on your available hardware (specifically VRAM/RAM) and your primary use case. The industry standards for local, private deployment fall into these key categories:
Qwen Series (e.g., Qwen3 or Qwen2.5) — Best overall for general tasks, multilingual support, and coding. Alibaba’s Qwen family (ranging from efficient 7B/8B sizes up to massive MoE variants) punches well above its weight class, handles over 100 languages exceptionally well, and features permissive commercial/open licenses.
Llama Series (e.g., Llama 3.1 / 3.3) — Best for general-purpose enterprise baseline. Meta’s Llama models are heavily optimized, widely supported across almost every local inference engine, and offer great general conversational, summarization, and extraction performance.
DeepSeek-R1 (Distilled variants or full models) — Best for heavy reasoning, logic, and complex coding. If your local tasks require math, structured logic, or multi-step problem solving, running a distilled local version (like the Qwen-distilled R1 variants) provides frontier-level reasoning entirely offline.
Recommended Local Stacks to Run Them
To actually serve and interface with these models on-premise without phoning home, use these battle-tested tools:
Inference Engine: Use Ollama for the simplest setup, or vLLM / llama.cpp for high-performance production serving on enterprise GPUs.
User Interface: Deploy Open WebUI locally—it gives you a ChatGPT-style interface on your internal network with role-based access control, completely air-gapped.
To narrow this down to the exact model size and hardware you'll need, could you tell me:
What GPU hardware (or total VRAM/RAM) do you have available on-premise?
What is your primary use case (e.g., general chatbot, private document search/RAG, automated coding, or complex reasoning)?
If your hard requirement is on-prem, with no application data leaving your network, I’d start with an open-weight model you can download and run entirely inside your infrastructure.
My default recommendation
Qwen3.5 is a strong starting family for an enterprise deployment. It has a broad range of model sizes, including versions that can run on a single GPU, and the smaller dense models are Apache 2.0 licensed. Current self-hosting guides also identify Qwen's smaller models as particularly practical for private deployments.
A rough hardware-based choice:
1× 24 GB GPU: Qwen3.5 ~27B-class or gpt-oss-20b. The latter is Apache 2.0 and designed to be practical on local hardware.
1× 80 GB GPU: Consider gpt-oss-120b, Qwen's larger models, or Llama 4 Scout depending on workload. Llama Scout is 109B total / 17B active and has a very large context window, but its license is more restrictive than Apache 2.0.
Multiple GPUs / serious inference server: Look at Qwen, Mistral Large 3, DeepSeek, or larger reasoning models. Mistral Large 3 is currently listed as Apache 2.0 and open-weight.
For a corporate/private-data use case, I'd pay particular attention to the license as well as model quality. "Open weights" does not necessarily mean the same thing as "unrestricted commercial use"; for example, Meta's Llama 4 uses its own Community License.
Also, "no data leaving our network" is really a deployment architecture requirement, not just a model choice. You'd want inference, embeddings, vector DB, logging, telemetry, model downloads, and any external tool/API calls contained within your network.
If you tell me (1) your GPU/server specs, (2) number of concurrent users, and (3) whether this is mainly chat, RAG over documents, coding, or agents, I can narrow it down to a specific model + quantization + serving stack (e.g. vLLM/llama.cpp) for your environment.
For an air-gapped, on-premises deployment where zero data can leave your network, the best open-weight model to run depends heavily on your available hardware (GPU VRAM) and your primary use case.
The top open-weight model families for local infrastructure are categorized below by capability and hardware footprint:
Qwen (Qwen3 / Qwen2.5) Family
Best For: Overall general assistant tasks, multilingual apps, complex reasoning, and coding.
Why it shines: Qwen models offer incredible performance relative to their size, frequently rivaling much larger proprietary models. They range from lightweight sizes (1.7B/7B/8B) up to massive enterprise tiers (70B+).
Hardware fit: An 8B or 14B version runs smoothly on a single consumer/workstation GPU (like an RTX 3090/4090 with 24GB VRAM), while larger variants need multi-GPU setups.
DeepSeek-R1 (and its distilled variants)
Best For: Advanced math, deep logic, and complex step-by-step reasoning (chain-of-thought workflows).
Why it shines: It delivers top-tier reasoning capabilities completely locally, meaning you get advanced analytical output without any telemetry or per-token cloud costs.
Hardware fit: Full R1 requires massive clusters, but the distilled versions (built on Llama and Qwen architectures in smaller 7B, 14B, or 32B sizes) fit nicely on standard local enterprise servers.
Llama (Llama 3 / 3.1 / 3.2) Family
Best For: General enterprise chat, RAG (Retrieval-Augmented Generation) pipelines, and standard text generation.
Why it shines: Widely adopted with the most mature ecosystem support for tool integration, fine-tuning, and guardrails.
Hardware fit: Llama 8B fits on a single modest GPU; 70B requires a heavy server node (e.g., 2 × or 4 × A100/H100 or equivalent enterprise accelerators).
Mistral / Mixtral Family
Best For: High-throughput, efficient reasoning, and strong coding performance with a smaller footprint.
Why it shines: Mixture-of-Experts (MoE) architectures like Mixtral give you high intelligence while only activating a fraction of the total parameters per token, maximizing generation speed on local hardware.
Recommended Deployment Stacks
To serve these models securely on your own hardware without phoning home, pair them with local inference engines:
Ollama: Easiest for quick setup, testing, and single-node CLI/API hosting.
vLLM or llama.cpp: Best for production-grade throughput, continuous batching, and handling concurrent team requests locally.
To narrow this down to a specific recommendation , let me know:
What hardware/GPUs do you have available on-prem (e.g., single 24GB VRAM card vs. multi-node enterprise server)?
What is the primary use case (e.g., internal coding assistant, document Q&A / RAG, automated data extraction)?
gpt-oss-120b: Best choice when you have serious GPU infrastructure. It has 117B total parameters / 5.1B active parameters and is designed to fit on a single 80 GB GPU in its provided MXFP4 quantization. OpenAI positions it for production, general-purpose, high-reasoning workloads.
gpt-oss-20b: Choose this if latency, hardware cost, or footprint matters more. It can run with as little as 16 GB memory and is aimed at local/specialized use cases.
Most importantly for your requirement, these are open-weight models that you run on infrastructure you control. OpenAI says data sent to a self-hosted gpt-oss deployment is not received or processed by OpenAI unless you explicitly share it or use a managed hosting partner.
For an enterprise deployment, I'd architect it roughly as:
Users/apps → internal API gateway → vLLM → gpt-oss-120b → internal data/RAG systems
with outbound network access disabled at the infrastructure level.
If you tell me your GPU(s), RAM, expected concurrent users, and whether this is primarily chat, coding, RAG, or agents, I can recommend the exact model/quantization and server configuration.
If your requirement is true on-prem deployment with no data leaving your network, I’d recommend OpenAI’s gpt-oss-120b as the default choice.
gpt-oss-120b — best for production workloads where quality/reasoning matters. It has 117B total parameters but only 5.1B active per token, and its MXFP4 weights fit in one 80 GB GPU.
gpt-oss-20b — choose this if latency, cost, or hardware constraints dominate. It can run with as little as 16 GB memory and is substantially lighter.
Both are open-weight, Apache 2.0 licensed, and explicitly designed to run on infrastructure you control. OpenAI states that data sent to self-hosted models is not received or processed by OpenAI.
My recommendation
Requirement
Model
Highest quality / enterprise production
gpt-oss-120b
Limited GPU budget
gpt-oss-20b
Single H100 server
gpt-oss-120b
Edge / workstation deployment
gpt-oss-20b
Sensitive data that must stay entirely on-prem
You can serve them with vLLM, Ollama, llama.cpp, or other supported local inference stacks.
If you tell me your GPU/server specs, expected concurrent users, context length, and whether this is RAG, coding, agents, or general chat, I can recommend the exact model + hardware configuration.
For an on-premises deployment where zero data can leave your network, you need to look at high-performing open-weight or open-source models that can be self-hosted via local inference runtimes like Ollama, vLLM , or llama.cpp.
The ideal model depends heavily on your hardware capacity (GPU VRAM) and your primary use case (general chat, deep reasoning, or coding). The top enterprise-ready open models fit into distinct categories:
Qwen3 (Alibaba): Best Overall / Multilingual. The Qwen family dominates enterprise open-source adoption. It offers exceptional reasoning, multilingual support, and heavy instruction-following capabilities that rival closed frontier models while running locally.
Gemma 4 (Google DeepMind): Best for Single-GPU / Efficiency. Models like the Gemma 4 31B dense or 26B MoE (Mixture-of-Experts) provide phenomenal performance on modest enterprise hardware (comfortably fitting on standard enterprise server GPUs with strong VRAM) while supporting robust function calling and multi-step agent workflows.
DeepSeek-R1 / V4: Best for Advanced Reasoning & Logic. If your on-prem workload requires deep math, complex analysis, or advanced coding capabilities, running a localized distillation or the core DeepSeek open weights delivers top-tier logic without external data egress.
Mistral Small 3.1 / Devstral: Best Apache 2.0 / Permissive Commercial Use. If legal compliance and strict open licensing (Apache 2.0) are mandatory for your corporate governance, Mistral and specialized code-assist variants offer predictable, high-speed performance for enterprise integrations.
To help narrow down the exact model and sizing (e.g., 8B, 30B, or 70B+ parameters), tell me:
What hardware/GPUs do you have available on-premises?
What is the primary use case (e.g., internal chatbot, automated coding, RAG over internal documents)?