“When ‘just call the API’ isn’t an option, every convenience you took for granted becomes a thing you have to build.”
There is a category of organization for which the entire modern AI industry is simply unavailable. Not too expensive, not too risky – unavailable. The data cannot leave the building. Not “shouldn’ t,” not “we’d prefer it didn’ t.” Can‘t-because a law, a contract, or a policy says so, and the penalty for being wrong is not a bad quarter but a shut-down program.
For that organization, “just call the API” is a non-starter, and so is almost everything built on top of it. The hosted model is off the table. So is the cloud vector database, the embeddings endpoint, the managed pipeline, and the helpful little SDK that quietly phones home with telemetry. Every convenience the GenAI stack is built on turns out, on inspection, to be a network call to someone else’ s datacenter – and you’ ve just lost the right to make any of them.
What’ s left is on-prem: the model, and everything around it, running on hardware inside the boundary, with nothing crossing the line. And the first thing you learn building it is that on-prem GenAI is not cloud GenAI with a different endpoint. It’ s a different discipline. Every convenience you used to consume as a service is now a component you own – one you have to select, size, secure, and keep alive yourself.
This is a field account of what that actually involves.
It’ s worth being precise about which on-prem you mean, because the word covers a spectrum. Some teams go on-prem for cost at scale, some for latency, some because a security team prefers it. Those are the soft versions, where the cloud is still a fallback if things get hard.
This piece is about the hard version: data sovereignty as a hard requirement. The information is sensitive enough – personal, regulated, or legally confined to a jurisdiction – that sending it to a third party isn’ t a tradeoff to weigh but a door that’ s locked. When that’ s the constraint, the cloud isn’ t a fallback. There is no fallback. Whatever you can’ t build inside the boundary, the organization simply doesn’ t get.
That reframes the whole engineering posture. In the cloud, a model is a service you consume – someone else runs it, scales it, patches it, and hands you a clean interface. On-prem, the model is a system you operate. And you don’ t own one system; you own four, each of which the cloud used to hide from you completely: the model, the retrieval stack, the hardware, and the lifecycle. Take them in turn.
The first substitution is the obvious one. The hosted API becomes a local serving runtime – an inference server running open-weight models you’ re actually permitted to run offline, on your own metal, with no license phoning out and no request leaving the box.
The runtime itself is the easy part; a lightweight local server will have a model answering prompts within an afternoon. What’ s not easy is everything you inherit the moment the model becomes yours.
Model selection is now your problem. The frontier hosted model you’d have reached for doesn’ t come in a form you can run in a locked room. So you choose
among open weights, and you live inside the capability gap between what you can host and what you wish you could. That gap is real, and pretending otherwise is how on-prem projects lose trust early. The honest move is to pick the smallest model that clears the task’ s actual bar, not the biggest one that fits.
Quantization is the central lever. A model’ s weights at full precision may not fit the VRAM you have, and even if they do, they leave no room for anything else.
Quantization – serving the model at reduced numerical precision – is how you make it fit, and it’ s the single most consequential tuning decision you’ ll make.
Push too far and quality degrades in ways that are subtle until they’ re embarrassing; not far enough and you can’ t fit the context length or the concurrency you need. Most of the practical craft of on-prem inference lives on this dial.
Quantization is the central lever. A model’ s weights at full precision may not fit the VRAM you have, and even if they do, they leave no room for anything else.
Quantization – serving the model at reduced numerical precision – is how you make it fit, and it’ s the single most consequential tuning decision you’ ll make.
Push too far and quality degrades in ways that are subtle until they’ re embarrassing; not far enough and you can’ t fit the context length or the concurrency you need. Most of the practical craft of on-prem inference lives on this dial.
Throughput is finite, and rationing it is your job. In the cloud, concurrency felt infinite because someone else absorbed the spikes. On one box, it isn’ t. Every simultaneous request shares the same GPU, the same memory, the same compute budget. You are now the capacity planner for a resource the cloud spent years making feel unlimited – and it isn’ t unlimited anymore, it’ s whatever you bought.
A self-hosted model on its own is a closed book – it knows only what was baked into its weights, which for a domain-specific task is rarely enough. Retrieval-augmented generation is how you fix that: you ground the model’ s answers in your own documents. And on-prem, RAG matters morethan it does in the cloud,
because it’ s the mechanism by which a smaller local model punches above its weight. What the model doesn’ t know, retrieval supplies.
But every piece of the retrieval stack that the cloud offered as a service, you now build inside the boundary.
The embedding model is a second model you host. Turning documents and queries into vectors requires an embedding model, and it can’ t be an API call any more than the LLM can. So you run it locally too — a second model to select, size, and keep consistent. That consistency is load-bearing: change the embedding model and every vector you’ ve ever stored is now incompatible, which means re-indexing your entire corpus. It’ s a decision you want to make once.
The vector store is infrastructure now. A self-hosted vector database – the on-prem equivalent of the managed service you can’ t use – comes with everything a managed service quietly handled for you: persistence, backups, index rebuilds, schema, upgrades. It’ s a database, and you’ re now its DBA.
Ingestion is the unglamorous eighty percent. Parsing documents, choosing a chunking strategy, attaching metadata, and keeping the index fresh as source material changes – none of it is glamorous and all of it determines whether retrieval actually works. There’ s no managed pipeline doing it in the background.
If the index goes stale, it goes stale because you didn’ t build the thing that keeps it fresh.
Here’ s the question every on-prem discussion dances around, and the one you can finally answer concretely: how much box do you need to serve how many people?
The trap is to answer “does the model fit?” That’ s the wrong question. Fitting is the floor, not the goal. A model that loads and sits at rest is not a model serving a room full of concurrent users at acceptable latency – and the gap between those two states is where sizing actually happens.
What determines it is VRAM, but not in the way people first assume. The GPU’ s memory is a single pool, and the model’ s weights are only the first tenant. Sharing that pool: the KV cache, which grows with every concurrent request and every token of context; the embedding model, resident alongside the LLM; and the headroom you need so the whole thing doesn’ t fall over under load. Spec the box for concurrency and contextlength, not for whether the weights fit when nothing’ s happening. The moment you serve real traffic, the KV cache is what runs you out of memory, not the model.
There’ s a further wrinkle if your on-prem platform is one of the newer unified-memory, ARM-based machines rather than a commodity x86 box with a discrete GPU. Unified memory changes the mental model – CPU and GPU sharing one physical pool has real advantages for fitting large models, and real quirks in how that memory gets allocated and contended. And the aarch64 software ecosystem still has gaps a seasoned x86 engineer won’ t expect: a wheel that doesn’ t exist for your architecture, a driver stack with sharper edges, a dependency that assumes an instruction set you don’ t have. None of it is fatal. All of it is time you have to budget for, because the internet’ s collective on-prem knowledge is overwhelmingly x86 and won’ t rescue you.
This is the part the cloud hides so thoroughly that most engineers have never had to think about it. In a managed world, deployment is a git push and a green checkmark. On-prem, deployment is a physical act you have to make reproducible– on hardware you may not control, and often for an operations team that will run the system long after you’ ve gone. That demands a few things the cloud never asked of you.
Setup and teardown have to be automation, not a runbook. You will stand this stack up more than once — on a second box, a replacement box, a box in a different site – and you cannot depend on a human remembering step fourteen. The whole thing, from bare OS to answering prompts, has to come up from scripts, and it has to come downcleanly too, without leaving orphaned state behind. A deployment you can’ t reproduce on demand isn’ t a deployment; it’ s a demo that happens to still be running.
Everything has to be offline–installable. Inside a strict boundary, you cannot assume a package index is reachable, a model registry will answer, or a container pull will succeed. Every artifact – models, dependencies, images – has to be staged and installable without the open internet. The first time you discover a hidden download-at-runtime in your own stack is the day you’ re standing in front of an air-gapped machine that will never see that URL.
Updates happen without a pipeline. Pushing a new model version or a security patch to a box behind the wall is its own small discipline. There’ s no rolling deploy from a CI system reaching in. You need a deliberate, legible way to update in place, verify it worked, and roll back if it didn’ t.
And you have to hand over the keys. On-prem usually means someone else’ s team operates the system once you leave. So it has to be legible and operable without you – documented, scripted, and boring in the best way. The golden rule underneath all of this:if standingit back up depends on you remembering,itisn‘t deployed.
Credibility on this subject lives in the honest ledger, so here it is.
What you gain is real and, for the right organization, decisive. The data never leaves – the requirement that started everything is satisfied absolutely, not mostly. There’ s no per-token bill that scales with success. There’ s no dependency on a vendor’ s uptime, no exposure to their pricing changes, no model deprecated out from under you on someone else’ s schedule. You own the stack, top to bottom.
What you pay is equally real. There’ s a capability gap between the open model you can host and the frontier model you can’ t, and while it narrows every quarter, on any given day it’ s there. You are now the SRE, the capacity planner, and the security boundary, all roles the cloud filled invisibly. Scaling means buying metal and racking it, not dragging a slider. And every convenience you used to consume is now a thing you maintain – forever, or until you hand it to someone who will.
The reframe that makes sense of the trade: on-prem isn’ t cheaper and it isn’ t easier. It’ s sovereign. You’ re exchanging convenience for control, and for the organization that started this whole exercise by saying the data can’ t leave, control was never the optional part.
The cloud’ s real magic was making generative AI feel like a faucet – you turn it on, water comes out, and you never think about the pipes. Building it on-prem is what reveals how much plumbing was always behind that wall. The model server, the embedding model, the vector store, the ingestion pipeline, the capacity
planning, the offline staging, the teardown scripts – none of it was ever absent.
You just weren’ t the one maintaining it.
The encouraging part is the direction of travel. Open models get better and denser every few months, the hardware to run them keeps gaining memory and losing cost, and the capability gap that makes on-prem sting is closing on a schedule you can watch. “Keep the data in the building” gets cheaper to say yes to every quarter. For the organizations that were never allowed to say anything else, that’ s not a minor trend. It’ s the difference between having modern AI and going without.