Insights · Hardware

What a private AI server actually needs

When people hear "run the AI on your own hardware," they often picture a server room and a six-figure infrastructure project. For the work most businesses actually do, the reality is a lot smaller than that, and a lot more boring, which is exactly what you want.

A single modern GPU can serve a whole team for private chat and document work. You scale the box to the job, not the other way around.

The model sets the hardware

The main thing that determines what you need is the size of the model you run, and most business tasks do not need a giant one. A compact open model handles private chat, drafting, summarizing, and document Q&A comfortably, and it fits on a single GPU. Step up to a mid-sized model for heavier reasoning across a company, and you are looking at a small multi-GPU box, not a rack of them. Only the largest, frontier-scale open models call for serious iron, and most businesses never need to go there.

What sits on the box

The hardware is only half of it. On top runs an open, well-understood stack: an inference server that serves the model on your GPUs, a gateway that gives your apps one clean endpoint, a vector database for retrieval over your documents, a chat interface your team actually uses, and monitoring so you can see usage and cost. None of it is exotic. All of it runs locally, with nothing phoning home.

Right-sizing, in plain terms

A small practice or team starts with a single-GPU box for chat and document work. A growing company moves to a few GPUs to serve a larger model to everyone. A larger or air-gapped operation runs a bigger build, still measured in a box or two, not a building. You buy the hardware close to cost and own it outright, which means the spend is an asset on your books rather than a subscription that never ends.

Builds by tier, in real hardware

Current-generation hardware as of mid-2026, sized to the model you actually need. You buy it close to cost and own it outright — an asset on your books, not a meter.

TierThe buildMemoryWhat it serves
StarterCompact GPU boxup to 96 GBCompact open models up to ~32B and gpt-oss — a pilot or very small team’s chat, drafting and document Q&A.
Core1× NVIDIA RTX PRO 6000 Blackwell96 GB7–32B models plus gpt-oss-120b — one team, served privately on a single box.
Enterprise4× RTX PRO 6000384 GB70B-class models at full FP8, served company-wide with headroom for concurrency.
Sovereign8× NVIDIA H200 (HGX node)1,128 GB · 1.1 TBFrontier open weights — GLM-5.2 (~750B), Mistral Large 3 — in one air-gapped node at native FP8.

Need more than a single node? We scale to multi-node B200 / GB200 clusters. Prefer a single box on a desk? Newer AI-specific machines — NVIDIA DGX Spark, Apple Mac Studio Ultra, AMD Strix Halo mini-PCs — run sizable models locally too. We size every build to your workload, compliance needs, and budget.

You do not have to operate it

The part that stops most teams is not buying a GPU, it is running the stack: sizing it, hardening it, patching it, keeping it fast. That is the part a managed service takes off your plate. You own the box; someone else keeps it healthy. See the full stack and process, compare the tiers, or run the numbers against what you spend on cloud AI today.

Sign up to learn more and we will size it to your use case.

More from Insights
Use cases

Private AI that actually knows your business

Read the article
Cost

On-prem AI vs. ChatGPT: the real cost over three years

Read the article
Healthcare

Can AI be HIPAA compliant? A straight answer for healthcare

Read the article
The hardware

The box that runs your AI.

STAVRYN · ON-PREM CLUSTER H200141 GB HBM3eH200141 GB HBM3eH200141 GB HBM3eH200141 GB HBM3eH200141 GB HBM3eH200141 GB HBM3eH200141 GB HBM3eH200141 GB HBM3e 8 × NVIDIA H200 1,128 GB VRAM · ≈ 1.1 TB CLUSTER UTILISATION 62% air-gappedyou own itflat cost running in your building · nothing leaves · you own it
Sovereign tier · 8 × NVIDIA H200 · 1.1 TB VRAM · air-gapped · you own it.
Common questions

Private AI hardware, answered.

How much hardware do I actually need?
Most business tasks run comfortably on a single GPU. A compact box handles chat, drafting, and document search for a team; a 70B-class model for a whole company fits on four workstation GPUs; only frontier-scale open models call for an eight-GPU H200 node. We size the build to your workload, not the other way around.
What GPUs do you use?
Current-generation NVIDIA: RTX PRO 6000 Blackwell with 96 GB for single and small multi-GPU builds, and H200 with 141 GB in eight-GPU nodes for the largest models. We can also spec AMD Instinct or newer local-AI machines where they fit a workload better.
Do I own the hardware?
Yes. Hardware is sourced and passed through close to cost; you buy it directly and own it from day one. It is an asset on your books rather than a lease, with no lock-in and no per-token meter.
Can it run fully air-gapped?
Yes. The whole stack runs locally with nothing phoning home, so it can be air-gapped inside your network, which is what regulated and defense work usually requires.