Where a Self-Hosted AI Stack Breaks Past One User (2026)

selfhostedollamamultiusercostai

TL;DR: A default Ollama install serves exactly one request at a time — the second user doesn’t crash it, they just wait in a queue. Two of the four scaling walls are fixed with free config changes; the other two cost $399–$1,300 in 2026 hardware. Know which wall you’re hitting before you buy anything.

Config fixes onlyRAM/GPU upgradeMove to vLLM
Best for2–4 occasional users3–8 regular users8+ concurrent, long contexts
Cost (Sep 2026)$0$399–$1,300$0 software + 24GB-class GPU
The catchKV cache eats your VRAM headroomDDR5 costs 4× what it did in 2025Ollama’s simplicity is gone

Honest take: Try OLLAMA_NUM_PARALLEL=2 and a keep-alive pin first — it’s free and it fixes the two most common complaints. Buy a used RTX 3090 only when ollama ps proves you’re out of VRAM, not because a forum told you to.

Your Ollama + Open WebUI box worked flawlessly for months — then you gave your partner, a roommate, or two teammates a login, and suddenly requests hang for a minute, models reload constantly, and the whole thing feels broken. Nothing is broken. A single-user stack makes four specific assumptions that stop holding at user number two, and each one has a known failure signature, a known fix, and a knowable price. All four below, cheapest first.

Why does a second user make your Ollama server feel frozen?

Because Ollama processes one request per model at a time by default. The official FAQ (checked September 22, 2026) states that OLLAMA_NUM_PARALLEL — “the maximum number of parallel requests each model will process at the same time” — defaults to 1, and that up to 512 further requests (OLLAMA_MAX_QUEUE) simply wait in line. So when user A asks for a 900-token summary and user B sends a prompt two seconds later, user B’s request sits in the queue until A’s generation finishes. On a 30B-class model producing ~25 tokens/second, that’s a 30–40 second wait before B sees a single token. B concludes the server is down. It isn’t — it’s serializing.

The fix costs nothing. Raise the parallel slot count where your Ollama service starts — on a systemd Linux box:

$ sudo systemctl edit ollama
[Service]
Environment="OLLAMA_NUM_PARALLEL=2"
Environment="OLLAMA_KEEP_ALIVE=24h"

$ sudo systemctl restart ollama
$ ollama ps
NAME         ID            SIZE     PROCESSOR    UNTIL
qwen3:30b    4a5c3e8f2b1d  21 GB    100% GPU     24 hours from now

Two things matter in that output. 100% GPU confirms the model still fits in VRAM after the change — if a CPU percentage appears there, part of the model spilled to system RAM and every user’s tokens got 5–10× slower, which is worse than the queue you started with. And the SIZE column will have grown compared to OLLAMA_NUM_PARALLEL=1, which brings us to the real constraint.

How much VRAM does each concurrent user actually cost?

Each parallel slot multiplies the context memory, not the model weights. Ollama’s FAQ is explicit: “a 2K context with 4 parallel requests will result in an 8K context and additional memory allocation” — total context allocation is OLLAMA_NUM_PARALLEL × OLLAMA_CONTEXT_LENGTH. The weights load once; the KV cache (the per-conversation working memory) is what scales per user.

That multiplication is why the naive move — “10 users, set parallel to 10” — falls off a cliff. A 30B-class model at Q4_K_M occupies ~19–21GB of a 24GB card. The few GB left over hold the KV cache for all slots combined. What fits, roughly, on a 24GB card running that model:

OLLAMA_NUM_PARALLELContext per userTotal context allocatedVerdict on 24GB
1 (default)8K8KFits, single-file queue
28K16KUsually fits — the sweet spot
48K32KTight; watch ollama ps for CPU spill
432K128KDoes not fit; model offloads to CPU

Exact gigabytes per slot depend on the model’s attention layout, so trust the empirical check over any table, this one included: run ollama ps after each change and back off the moment PROCESSOR shows anything but 100% GPU. The escape hatches, in ascending price: lower the per-user context, drop the model one quantization level (trade-offs mapped in our quantization guide), or buy VRAM — priced two sections down.

A related wall arrives if your users touch different models — one person chats with a 30B model while another’s editor hits a 7B coder model. A 24GB card can’t hold both plus KV, so Ollama evicts and reloads, and everyone pays a multi-second swap penalty several times an hour. Our five-person team cost breakdown hit exactly this thrashing in week two; the config-level fix (pin models with OLLAMA_KEEP_ALIVE, cap OLLAMA_MAX_LOADED_MODELS, shrink one model) is documented there.

What breaks in Open WebUI when the logins multiply?

The inference engine is only half the stack — the UI layer accumulates state, and single-user habits around that state turn into multi-user problems:

  • Accounts and roles exist; use them. Open WebUI (0.11 series as of September 2026) ships multi-user support with an admin role and per-user chat history, so the actual account layer is not the problem. The problem is that many single-user installs run with auth disabled (WEBUI_AUTH=False) or share one login — fine for you, a privacy failure the day a second person’s conversations land in the same history.
  • The default database is SQLite. According to the project docs, Open WebUI uses a single SQLite file unless you point DATABASE_URL at Postgres. A handful of users is fine on SQLite; what changes is consequence — that one file now holds everyone’s chats and RAG documents, so it needs real backups, and heavy simultaneous RAG ingestion is where users report it straining first.
  • Per-user documents multiply disk and embedding load. Every user uploading PDFs into their own knowledge collections means embedding jobs competing with chat inference for the same GPU, and a document store that grows N× faster than before.
  • Remote users force a network decision. The moment one user isn’t on your LAN, you either expose the UI through a reverse proxy with TLS and rate limiting, or you put everyone on a VPN (WireGuard/Tailscale). The clean split — UI on a small VPS a few dollars a month (Vultr entry instances do it), GPU box tunneled behind your firewall — is laid out in the team stack article. Never expose port 11434 itself: Ollama’s API has no authentication, and thousands of open instances are already indexed by scanners (details and hardening).

None of this costs hardware money. It costs admin hours — and those grow ~50% moving from a solo stack to a shared one, per our maintenance hours log. Price your hours honestly before the GPU.

When do you outgrow Ollama entirely and need vLLM?

When genuinely simultaneous requests are the norm rather than the collision — in practice somewhere past 5–8 truly concurrent users, or earlier if they run long contexts. The architectural difference: Ollama allocates fixed parallel slots and queues everything beyond them, while vLLM’s continuous batching admits new requests into the running batch the moment old ones finish, and its PagedAttention allocates KV cache in small blocks instead of reserving each user’s full context up front. Same GPU, materially higher aggregate throughput under concurrent load — the mechanics and benchmarks are in our Ollama vs vLLM comparison.

The serve command is still one line:

$ vllm serve Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 --max-model-len 16384

Two honest caveats before you switch. vLLM does not solve a VRAM shortage — it schedules the memory you have more efficiently, but a model that didn’t fit under Ollama still doesn’t fit (multi-GPU is its own project: vLLM multi-GPU setup). And you give up Ollama’s conveniences: one-command model pulls, automatic model swapping, and the lowest-friction upgrade path. For a household or small team that occasionally collides, OLLAMA_NUM_PARALLEL=2 is the better trade; vLLM earns its complexity when the GPU would otherwise sit half-idle between queued requests.

What does the multi-user hardware fix cost in September 2026?

More than the old advice assumes, because 2026 repriced both RAM and GPUs. DDR5 has roughly quadrupled since summer 2025 — 32GB kits run $399–$479 and 64GB $680–$1,070 (Tom’s Hardware RAM price index, September 11, 2026) — and street prices on current GPUs sit far above MSRP. The upgrade ladder, cheapest first:

Wall you hitThe fixPrice (verified Sep 2026)
Queue delays, VRAM fineOLLAMA_NUM_PARALLEL=2 + keep-alive pin$0
Embedding/RAG jobs starve system RAM32GB DDR5 kit$399–$479
KV cache spills at 2+ slots on a 16GB cardUsed RTX 3090 24GB~$1,252 (was ~$1,010 in March — used prices are rising)
Starting from iGPU/8GB, budget-cappedRTX 5060 Ti 16GB$679–$805 street (MSRP $429 — nobody pays MSRP)
Two models resident + headroom, no swap everUsed RTX 4090 24GB$2,150–$2,350
Big-MoE models for several users, one quiet boxGMKtec EVO-X2 128GB unified~$3,649

Three notes on that table. Don’t reflex-buy 64GB of RAM “while you’re in there” — at $680+ it costs more than stepping up a GPU tier used to, and a multi-user inference box rarely needs it; buy for the measured workload. The 16GB-card row is a compromise: it runs 13B-class models for 2–3 users well, but it cannot hold a 30B model and multi-slot KV, so treat it as the budget entry, not the destination. And the unified-memory route (Strix Halo boxes like the EVO-X2) trades raw speed for capacity — the hardware analysis lives on our sister site: Ryzen AI Max 395 “Strix Halo” review.

Also budget the invisible line item: a box that used to sleep now runs 24/7 for other people. At US average rates that’s real but tolerable; at 35¢+/kWh it can exceed $400/year — regional math in always-on server electricity costs.

When should you NOT scale your own stack past one user?

  • The other users are occasional. If user two asks five questions a week, the collision you’re engineering around almost never happens. OLLAMA_MAX_QUEUE already handles it — the default queue of 512 exists precisely so rare overlaps wait a few seconds instead of failing. Spend nothing.
  • Total usage is light across the board. A metered API key covering everyone at $10–30/month beats a $1,252 GPU for years. Measure 60 days of real usage first; the broader framework is in when NOT to self-host AI.
  • Nobody wants to be the admin. Multi-user means backups, accounts, updates, and uptime expectations — 45–75 hours in year one for a small team. A resented stack gets abandoned by month four, after the hardware is paid for.
  • Your second user needs frontier-model quality. A 30B open-weight model is excellent for chat, summarization, and code completion; it is not a frontier reasoning model. Hybrid (local for private/routine, one API seat for hard problems) often beats forcing everything local.
  • You’d be buying at the top. Used 3090s rose ~24% since March and DDR5 is at crisis pricing. If your current setup limps along at NUM_PARALLEL=2, waiting is a legitimate strategy for RAM in particular — renting covers the gap.

What to actually buy

Prices as of September 2026, all taken from the comparison above:

Your situationThe fixPriceWhere
2–4 users, requests rarely collideConfig changes only$0This article, section one
16GB card, KV cache spilling at 2 slotsUsed RTX 3090 24GB~$1,252Check price
No dedicated GPU yet, hard $800 ceilingRTX 5060 Ti 16GB$679–$805Check price
Multiple models resident, zero swap toleranceUsed RTX 4090 24GB$2,150–$2,350Check price
Remote users — UI layer off the GPU boxSmall VPS + WireGuard tunnela few $/moVultr
Unsure the extra users will stick aroundRented RTX 3090, test a monthfrom ~$0.07/hrVast.ai

The rental row deserves the last word: a rented 3090 at market rates costs less per month than the interest on regret from a $1,252 card your second user stopped touching in week three. Point your existing Open WebUI at the rented endpoint, watch actual concurrency for 30 days, then buy the row that matches what you measured.

FAQ

Does adding users slow down my own single-user requests? Only when requests actually overlap. With OLLAMA_NUM_PARALLEL=2 and enough VRAM, two simultaneous generations share GPU compute, so each streams somewhat slower than it would alone — but a request arriving to an idle server runs at full speed regardless of how many accounts exist. Accounts are free; concurrency is what costs.

Can I just run two Ollama instances on one GPU instead? You can, but you gain nothing over OLLAMA_NUM_PARALLEL=2 and lose the scheduler’s coordination — two instances will fight over VRAM with no shared eviction logic. Separate instances only make sense on separate GPUs, and at that point compare against a single vLLM deployment with tensor parallelism.

Is Open WebUI’s SQLite database a real problem at 3–5 users? Usually not for chat alone — the strain shows up with heavy simultaneous RAG ingestion, and the bigger issue is operational: one file now holds everyone’s data, so it needs scheduled backups. The project docs support Postgres via DATABASE_URL when you want the sturdier path; treat that as an at-leisure migration, not an emergency.

Sources

Was this article helpful?

What self-hosting actually costs

Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.