Where a Self-Hosted AI Stack Breaks Past One User (2026)
TL;DR: A default Ollama install serves exactly one request at a time — the second user doesn’t crash it, they just wait in a queue. Two of the four scaling walls are fixed with free config changes; the other two cost $399–$1,300 in 2026 hardware. Know which wall you’re hitting before you buy anything.
| Config fixes only | RAM/GPU upgrade | Move to vLLM | |
|---|---|---|---|
| Best for | 2–4 occasional users | 3–8 regular users | 8+ concurrent, long contexts |
| Cost (Sep 2026) | $0 | $399–$1,300 | $0 software + 24GB-class GPU |
| The catch | KV cache eats your VRAM headroom | DDR5 costs 4× what it did in 2025 | Ollama’s simplicity is gone |
Honest take: Try
OLLAMA_NUM_PARALLEL=2and a keep-alive pin first — it’s free and it fixes the two most common complaints. Buy a used RTX 3090 only whenollama psproves you’re out of VRAM, not because a forum told you to.
Your Ollama + Open WebUI box worked flawlessly for months — then you gave your partner, a roommate, or two teammates a login, and suddenly requests hang for a minute, models reload constantly, and the whole thing feels broken. Nothing is broken. A single-user stack makes four specific assumptions that stop holding at user number two, and each one has a known failure signature, a known fix, and a knowable price. All four below, cheapest first.
Why does a second user make your Ollama server feel frozen?
Because Ollama processes one request per model at a time by default. The official FAQ (checked September 22, 2026) states that OLLAMA_NUM_PARALLEL — “the maximum number of parallel requests each model will process at the same time” — defaults to 1, and that up to 512 further requests (OLLAMA_MAX_QUEUE) simply wait in line. So when user A asks for a 900-token summary and user B sends a prompt two seconds later, user B’s request sits in the queue until A’s generation finishes. On a 30B-class model producing ~25 tokens/second, that’s a 30–40 second wait before B sees a single token. B concludes the server is down. It isn’t — it’s serializing.
The fix costs nothing. Raise the parallel slot count where your Ollama service starts — on a systemd Linux box:
$ sudo systemctl edit ollama
[Service]
Environment="OLLAMA_NUM_PARALLEL=2"
Environment="OLLAMA_KEEP_ALIVE=24h"
$ sudo systemctl restart ollama
$ ollama ps
NAME ID SIZE PROCESSOR UNTIL
qwen3:30b 4a5c3e8f2b1d 21 GB 100% GPU 24 hours from now
Two things matter in that output. 100% GPU confirms the model still fits in VRAM after the change — if a CPU percentage appears there, part of the model spilled to system RAM and every user’s tokens got 5–10× slower, which is worse than the queue you started with. And the SIZE column will have grown compared to OLLAMA_NUM_PARALLEL=1, which brings us to the real constraint.
How much VRAM does each concurrent user actually cost?
Each parallel slot multiplies the context memory, not the model weights. Ollama’s FAQ is explicit: “a 2K context with 4 parallel requests will result in an 8K context and additional memory allocation” — total context allocation is OLLAMA_NUM_PARALLEL × OLLAMA_CONTEXT_LENGTH. The weights load once; the KV cache (the per-conversation working memory) is what scales per user.
That multiplication is why the naive move — “10 users, set parallel to 10” — falls off a cliff. A 30B-class model at Q4_K_M occupies ~19–21GB of a 24GB card. The few GB left over hold the KV cache for all slots combined. What fits, roughly, on a 24GB card running that model:
OLLAMA_NUM_PARALLEL | Context per user | Total context allocated | Verdict on 24GB |
|---|---|---|---|
| 1 (default) | 8K | 8K | Fits, single-file queue |
| 2 | 8K | 16K | Usually fits — the sweet spot |
| 4 | 8K | 32K | Tight; watch ollama ps for CPU spill |
| 4 | 32K | 128K | Does not fit; model offloads to CPU |
Exact gigabytes per slot depend on the model’s attention layout, so trust the empirical check over any table, this one included: run ollama ps after each change and back off the moment PROCESSOR shows anything but 100% GPU. The escape hatches, in ascending price: lower the per-user context, drop the model one quantization level (trade-offs mapped in our quantization guide), or buy VRAM — priced two sections down.
A related wall arrives if your users touch different models — one person chats with a 30B model while another’s editor hits a 7B coder model. A 24GB card can’t hold both plus KV, so Ollama evicts and reloads, and everyone pays a multi-second swap penalty several times an hour. Our five-person team cost breakdown hit exactly this thrashing in week two; the config-level fix (pin models with OLLAMA_KEEP_ALIVE, cap OLLAMA_MAX_LOADED_MODELS, shrink one model) is documented there.
What breaks in Open WebUI when the logins multiply?
The inference engine is only half the stack — the UI layer accumulates state, and single-user habits around that state turn into multi-user problems:
- Accounts and roles exist; use them. Open WebUI (0.11 series as of September 2026) ships multi-user support with an admin role and per-user chat history, so the actual account layer is not the problem. The problem is that many single-user installs run with auth disabled (
WEBUI_AUTH=False) or share one login — fine for you, a privacy failure the day a second person’s conversations land in the same history. - The default database is SQLite. According to the project docs, Open WebUI uses a single SQLite file unless you point
DATABASE_URLat Postgres. A handful of users is fine on SQLite; what changes is consequence — that one file now holds everyone’s chats and RAG documents, so it needs real backups, and heavy simultaneous RAG ingestion is where users report it straining first. - Per-user documents multiply disk and embedding load. Every user uploading PDFs into their own knowledge collections means embedding jobs competing with chat inference for the same GPU, and a document store that grows N× faster than before.
- Remote users force a network decision. The moment one user isn’t on your LAN, you either expose the UI through a reverse proxy with TLS and rate limiting, or you put everyone on a VPN (WireGuard/Tailscale). The clean split — UI on a small VPS a few dollars a month (Vultr entry instances do it), GPU box tunneled behind your firewall — is laid out in the team stack article. Never expose port 11434 itself: Ollama’s API has no authentication, and thousands of open instances are already indexed by scanners (details and hardening).
None of this costs hardware money. It costs admin hours — and those grow ~50% moving from a solo stack to a shared one, per our maintenance hours log. Price your hours honestly before the GPU.
When do you outgrow Ollama entirely and need vLLM?
When genuinely simultaneous requests are the norm rather than the collision — in practice somewhere past 5–8 truly concurrent users, or earlier if they run long contexts. The architectural difference: Ollama allocates fixed parallel slots and queues everything beyond them, while vLLM’s continuous batching admits new requests into the running batch the moment old ones finish, and its PagedAttention allocates KV cache in small blocks instead of reserving each user’s full context up front. Same GPU, materially higher aggregate throughput under concurrent load — the mechanics and benchmarks are in our Ollama vs vLLM comparison.
The serve command is still one line:
$ vllm serve Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 --max-model-len 16384
Two honest caveats before you switch. vLLM does not solve a VRAM shortage — it schedules the memory you have more efficiently, but a model that didn’t fit under Ollama still doesn’t fit (multi-GPU is its own project: vLLM multi-GPU setup). And you give up Ollama’s conveniences: one-command model pulls, automatic model swapping, and the lowest-friction upgrade path. For a household or small team that occasionally collides, OLLAMA_NUM_PARALLEL=2 is the better trade; vLLM earns its complexity when the GPU would otherwise sit half-idle between queued requests.
What does the multi-user hardware fix cost in September 2026?
More than the old advice assumes, because 2026 repriced both RAM and GPUs. DDR5 has roughly quadrupled since summer 2025 — 32GB kits run $399–$479 and 64GB $680–$1,070 (Tom’s Hardware RAM price index, September 11, 2026) — and street prices on current GPUs sit far above MSRP. The upgrade ladder, cheapest first:
| Wall you hit | The fix | Price (verified Sep 2026) |
|---|---|---|
| Queue delays, VRAM fine | OLLAMA_NUM_PARALLEL=2 + keep-alive pin | $0 |
| Embedding/RAG jobs starve system RAM | 32GB DDR5 kit | $399–$479 |
| KV cache spills at 2+ slots on a 16GB card | Used RTX 3090 24GB | ~$1,252 (was ~$1,010 in March — used prices are rising) |
| Starting from iGPU/8GB, budget-capped | RTX 5060 Ti 16GB | $679–$805 street (MSRP $429 — nobody pays MSRP) |
| Two models resident + headroom, no swap ever | Used RTX 4090 24GB | $2,150–$2,350 |
| Big-MoE models for several users, one quiet box | GMKtec EVO-X2 128GB unified | ~$3,649 |
Three notes on that table. Don’t reflex-buy 64GB of RAM “while you’re in there” — at $680+ it costs more than stepping up a GPU tier used to, and a multi-user inference box rarely needs it; buy for the measured workload. The 16GB-card row is a compromise: it runs 13B-class models for 2–3 users well, but it cannot hold a 30B model and multi-slot KV, so treat it as the budget entry, not the destination. And the unified-memory route (Strix Halo boxes like the EVO-X2) trades raw speed for capacity — the hardware analysis lives on our sister site: Ryzen AI Max 395 “Strix Halo” review.
Also budget the invisible line item: a box that used to sleep now runs 24/7 for other people. At US average rates that’s real but tolerable; at 35¢+/kWh it can exceed $400/year — regional math in always-on server electricity costs.
When should you NOT scale your own stack past one user?
- The other users are occasional. If user two asks five questions a week, the collision you’re engineering around almost never happens.
OLLAMA_MAX_QUEUEalready handles it — the default queue of 512 exists precisely so rare overlaps wait a few seconds instead of failing. Spend nothing. - Total usage is light across the board. A metered API key covering everyone at $10–30/month beats a $1,252 GPU for years. Measure 60 days of real usage first; the broader framework is in when NOT to self-host AI.
- Nobody wants to be the admin. Multi-user means backups, accounts, updates, and uptime expectations — 45–75 hours in year one for a small team. A resented stack gets abandoned by month four, after the hardware is paid for.
- Your second user needs frontier-model quality. A 30B open-weight model is excellent for chat, summarization, and code completion; it is not a frontier reasoning model. Hybrid (local for private/routine, one API seat for hard problems) often beats forcing everything local.
- You’d be buying at the top. Used 3090s rose ~24% since March and DDR5 is at crisis pricing. If your current setup limps along at
NUM_PARALLEL=2, waiting is a legitimate strategy for RAM in particular — renting covers the gap.
What to actually buy
Prices as of September 2026, all taken from the comparison above:
| Your situation | The fix | Price | Where |
|---|---|---|---|
| 2–4 users, requests rarely collide | Config changes only | $0 | This article, section one |
| 16GB card, KV cache spilling at 2 slots | Used RTX 3090 24GB | ~$1,252 | Check price |
| No dedicated GPU yet, hard $800 ceiling | RTX 5060 Ti 16GB | $679–$805 | Check price |
| Multiple models resident, zero swap tolerance | Used RTX 4090 24GB | $2,150–$2,350 | Check price |
| Remote users — UI layer off the GPU box | Small VPS + WireGuard tunnel | a few $/mo | Vultr |
| Unsure the extra users will stick around | Rented RTX 3090, test a month | from ~$0.07/hr | Vast.ai |
The rental row deserves the last word: a rented 3090 at market rates costs less per month than the interest on regret from a $1,252 card your second user stopped touching in week three. Point your existing Open WebUI at the rented endpoint, watch actual concurrency for 30 days, then buy the row that matches what you measured.
FAQ
Does adding users slow down my own single-user requests?
Only when requests actually overlap. With OLLAMA_NUM_PARALLEL=2 and enough VRAM, two simultaneous generations share GPU compute, so each streams somewhat slower than it would alone — but a request arriving to an idle server runs at full speed regardless of how many accounts exist. Accounts are free; concurrency is what costs.
Can I just run two Ollama instances on one GPU instead?
You can, but you gain nothing over OLLAMA_NUM_PARALLEL=2 and lose the scheduler’s coordination — two instances will fight over VRAM with no shared eviction logic. Separate instances only make sense on separate GPUs, and at that point compare against a single vLLM deployment with tensor parallelism.
Is Open WebUI’s SQLite database a real problem at 3–5 users?
Usually not for chat alone — the strain shows up with heavy simultaneous RAG ingestion, and the bigger issue is operational: one file now holds everyone’s data, so it needs scheduled backups. The project docs support Postgres via DATABASE_URL when you want the sturdier path; treat that as an at-leisure migration, not an emergency.
Sources
- Ollama official FAQ — concurrency,
OLLAMA_NUM_PARALLEL, queue and memory behavior (checked September 22, 2026) - Tom’s Hardware RAM Price Index (September 11, 2026 — DDR5 kit pricing)
- Open WebUI — GitHub repository and documentation (multi-user, auth, and database configuration)
- vLLM documentation — serving and continuous batching (PagedAttention, OpenAI-compatible server)
- Vast.ai pricing (market rental rates for consumer GPUs, September 2026)
Recommended Gear
- Used RTX 3090 24GB — the multi-user VRAM fix, ~$1,252 used (Sep 2026)
- RTX 5060 Ti 16GB — budget entry for 2–3 users on 13B-class models, $679–$805 street
- Used RTX 4090 24GB — two resident models plus KV headroom, $2,150–$2,350
- GMKtec EVO-X2 128GB — unified-memory capacity play, ~$3,649
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →What self-hosting actually costs
Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.