Sizing GPU servers for self-hosted AI: models, users, retrieval and headroom

GPU memory, more than raw compute, decides what a self-hosted AI server can do. How to size for the models you want, the users you actually have, document retrieval and room to grow.

The first question most organisations ask about self-hosted AI is which GPU to buy. The more useful one is what the hardware will be asked to do: which models, for how many people at once, with how much document context, and how much of the work needs to stay on-premises at all. Answer those and the hardware mostly follows. Skip them and it is easy to buy a server that loads the model perfectly well, then struggles the week a second department comes on board.

Memory decides what fits

Running a large language model is dominated by memory. The model’s weights have to sit in GPU memory, and for a conventional dense model, generating each token means reading essentially all of them. Memory capacity therefore decides which models will run at all, and memory bandwidth largely decides how quickly a single user sees words appear.

The arithmetic for weights is simple: parameters multiplied by bytes per parameter. At 16-bit precision each parameter takes two bytes, so an 8-billion-parameter model needs about 16 GB and a 70-billion-parameter model about 140 GB, before anything else is loaded. Quantisation reduces this. At 8 bits the requirement roughly halves, and at 4 bits the same 70-billion-parameter model fits in around 40 GB.

Quantisation is not free. Eight-bit versions of good models usually perform very close to the original; four-bit results vary more by model and task, so test on your own documents before settling on one.

Mixture-of-experts models use only a fraction of their parameters for each token, which makes them fast for their size, but every expert still has to be held in memory. Size for the total parameter count, not the active one.

The part everyone forgets: context and concurrency

Weights are the fixed cost. The variable cost is the key-value cache, the working memory a model keeps for every token in every conversation it is currently handling. It grows in proportion to context length and to the number of conversations in flight.

The numbers are larger than people expect. For a typical 70-billion-parameter model that uses grouped-query attention, the cache at 16-bit precision works out to roughly a third of a megabyte per token. A single 32,000-token conversation, say a long contract plus the discussion about it, needs around 10 GB. Ten of those at the same time need around 100 GB, on top of the weights.

Modern serving engines manage this memory well, allocating the cache in small pages, batching concurrent requests and, in some cases, storing the cache at reduced precision. What they cannot do is create memory that is not there. When the cache is full, new requests wait.

Document-native platforms push this further: a user who drops a 60-page PDF into a chat has just asked for a long context. Size for the documents people actually work with, not short test questions.

Users: count concurrency, not headcount

Named users and concurrent requests are very different numbers. Most people with access are not using the tool at any given moment, and use comes in bursts. If 300 staff have access and, at the busiest minute of the week, 20 of them have a request in progress, then 20 concurrent sequences is the planning figure, and the length of those sequences decides the cache.

The only reliable way to get these numbers is to run a pilot and measure peak concurrent requests, typical input length including attached documents, and typical output length.

Each request also has two phases. Prefill processes the prompt and any documents and is compute-heavy, so a long document means a noticeable wait before the first word. Decode generates the answer token by token and is limited by memory bandwidth; batching users together raises total throughput at some cost to each user’s speed. The practical targets are a short time to first token and a generation speed comfortably faster than people read.

Retrieval has its own appetite

Retrieval-augmented generation, where the platform searches a document library and hands relevant passages to the model, brings its own workloads. An embedding model converts documents into vectors when they are ingested and converts each question at query time, and a reranker may score candidate passages. These are far smaller than the main model and can often share a GPU or run on CPUs, but bulk ingestion of a large library is a real burst of work: schedule it, or give it its own capacity, so it does not compete with interactive users.

Parsing and text recognition on scanned files take compute too, and the vector index wants memory and fast storage. And because retrieved passages are added to the prompt, good retrieval makes contexts longer, which brings us back to the cache.

Not everything has to run locally

A multi-model platform changes the sizing question. If sensitive work runs on open-weight models on your own hardware, and tasks whose data classification allows it go to frontier models through a governed gateway, the local GPUs only carry the sensitive workload, not the whole organisation’s demand. We covered that routing model in self-hosted enterprise AI.

The same thinking applies to model choice. A well-chosen mid-sized model on hardware with room for long contexts will often serve staff better than the largest available model squeezed onto too little memory.

The server around the GPUs

  • Splitting models across GPUs. A model too large for one GPU can be split across several, which works best with fast GPU-to-GPU links. Where a model fits on one GPU, running separate copies on each is usually simpler and scales better.
  • Power and cooling. A server carrying several data-centre GPUs can draw several kilowatts under sustained load. Many comms rooms and older racks were never built for that density, so check the power feed, UPS capacity and cooling before the purchase order, not after delivery.
  • Data-centre or consumer cards. Data-centre GPUs bring error-correcting memory, cooling designed for rack chassis and enterprise support. Consumer cards are cheaper per gigabyte, but check the vendor’s licence terms for your intended use, and how they will be cooled in a server.
  • Everything else. Enough system memory, and fast NVMe so that loading a model does not take minutes.

Headroom and growth

Do not plan to run GPU memory at its limit. Serving engines reserve memory for their own use, and contexts and user numbers grow once people find the tool useful. Open-weight models also change quickly, and next year’s preferred model may be larger or support longer contexts. Servers with spare GPU slots, or a design that adds a second node, age better than a box that is full on day one.

Decide early how much an outage matters. A single GPU server is a single point of failure, and once the platform is part of how work gets done it needs a second server or a fallback route to another model.

A practical starting point

  • Pick candidate models and work out their weight memory at the precision you intend to run.
  • Run a pilot, and measure peak concurrent requests and typical context length.
  • Add cache memory for that concurrency and context, then retrieval, then meaningful headroom.
  • Check the power and cooling in the room the server will actually live in.
  • Decide which workloads must stay local and which can be routed elsewhere.

Where Bizix fits

Bizix designs and operates self-hosted AI platforms that put frontier and open-weight models behind one governed interface, with per-user routing and cost visibility by user and by model, on your infrastructure, in our Australian cloud or as a hybrid. The same engineering team designs the infrastructure underneath, so sizing questions like these are answered alongside the platform rather than handed to a separate vendor.