Skip to content

On-premise AI

Running AI on your own hardware

Open-weight language models can now run on a single workstation or server in your office. For most small businesses a cloud API is still cheaper and simpler. For some, local hardware wins on privacy, cost at high volume or speed, and the deciding number is how much memory the model needs.

Start with the API

A hosted model from a large provider is the default for good reasons. It needs no hardware, it is usually the most capable model you can get, you pay only for what you use, and someone else handles updates, uptime and security patches. Business and enterprise plans from the major providers generally commit not to train on your data. Read those terms before assuming the privacy case for local hardware.

When local makes sense

  • Data that cannot leave the building. Client files under a strict confidentiality agreement, regulated health or financial records, or material a customer contract says stays on your systems. Local models remove the third party entirely.
  • Steady, high volume. API bills scale with every token. A workload that runs all day, every day (classifying documents, summarizing tickets, extracting fields) can pay back a GPU within its life. A workload used a few times a day almost never does.
  • Latency and offline work. A model on the local network answers without a round trip to a data center and keeps working when the internet link does not.
  • Control. The model does not change underneath you. A hosted model can be updated or retired on the provider's schedule, and prompts tuned for one version may behave differently on the next.

The trade-off is capability. The largest frontier models are not available to run locally, and the open models a small business can host are smaller. For narrow, repetitive tasks the gap is often small. For open-ended reasoning and writing it can be large. Test your real task on both before buying anything.

The memory math

The model has to fit in memory: GPU memory (VRAM) on a graphics card, or unified memory on machines where the CPU and GPU share it. The weights take parameters x bytes per parameter:

  • FP16 or BF16: 2 bytes per parameter. Full precision as most models are published.
  • FP8 or INT8: 1 byte. Usually a small quality loss.
  • 4-bit quantized: about half a byte. The common choice for local use, with some quality loss that varies by model and task.
Model sizeFP168-bit4-bit
7 to 8 billion16 GB8 GB4 GB
13 to 14 billion28 GB14 GB7 GB
30 to 32 billion64 GB32 GB16 GB
70 billion140 GB70 GB35 GB

Those figures are the weights alone. On top sits the KV cache, the working memory the model keeps for the conversation so far. It grows with the length of the prompt and with the number of users served at once. Add the runtime's own overhead and a margin of about 20% over the weights is a reasonable starting estimate for one user with moderate prompts. Long documents or several simultaneous users need more. The calculator lets you change both.

Hardware tiers

  • A workstation with one high-memory GPU. Consumer cards with 24 or 32 GB run 8 to 14 billion parameter models comfortably and around 30 billion at 4-bit. Professional cards with 48 GB or more reach further, and add ECC memory and certified drivers.
  • Unified-memory machines. Some desktops and compact workstations share a large pool of memory between CPU and GPU, making big models fit on a quiet desktop. They are usually slower per user than a dedicated GPU of the same memory size.
  • Servers with several GPUs. For 70 billion parameters and up at higher precision, or many users at once. These need rack space, cooling and serious power. See the server register and our server guide.

Browse specific machines and cards in the AI hardware directory. A GPU server also needs a properly sized UPS; a single high-end card can draw several hundred watts under load.

Costs people miss

  • Electricity and heat. A GPU running all day adds to the power bill and to the cooling load of the room it sits in.
  • Someone to run it. Updates, model upgrades, access control and monitoring are ongoing work.
  • Model licenses. Open-weight does not always mean unrestricted. Read each model's license for commercial use terms.
  • Security. A local model server is another system on your network that holds sensitive data. Treat it that way.

Before you buy

  1. Run your real task through a hosted API and a local model of the size you can afford. Compare the output.
  2. Estimate monthly API spend at your true volume. Local hardware pays back only at steady, high use.
  3. Size memory with the calculator, then add headroom for long prompts and more users.
  4. Check the model license allows your commercial use.
  5. Plan power, cooling, backup power and who maintains it.
  6. Write the system into your AI use policy: who can use it, for what data.

High-memory GPUs for local AI

Affiliate links

Graphics cards with 24 GB or more of memory, the practical floor for running useful language models locally. Check power supply capacity and case clearance first.

Prices retrieved from Amazon on 2026-09-29 and may since have changed; the price on Amazon at the time of purchase is the one that applies. As an Amazon Associate we earn from qualifying purchases, at no extra cost to you. No manufacturer pays for a place in this list.