Ollama vs OpenRouter vs Groq vs NVIDIA NIM in 2026

ollamaopenroutergroqnvidia-nimselfhosted

TL;DR: Ollama is the only one of the four that keeps prompts on your hardware, and it costs nothing beyond electricity. OpenRouter and Groq trade privacy for model scale and speed; NVIDIA NIM keeps data local but demands datacenter GPUs and NVIDIA licensing. For a privacy-first home lab, Ollama wins and it isn’t close.

OllamaOpenRouterGroqNVIDIA NIM
Best forPrivate local inferenceAccess to 400+ models, one keyRaw speed on hosted open modelsOn-prem enterprise compliance
Cost$0 + electricity (~$0.04–0.06/hr on a 350W GPU)Provider price + 5.5% credit feee.g. GPT-OSS-120B $0.15/$0.60 per 1M tokensHardware + NVIDIA AI Enterprise for production
The catchLimited by your VRAMPrompts leave your networkCloud-only; catalog shrank in Aug 2026NGC account, proprietary containers, A100/H100-class GPUs

Honest take: run Ollama for anything private, and keep an OpenRouter key around for the rare job that needs a 400B-class model. Skip Groq unless you specifically need 250+ tokens/second, and skip NIM entirely unless compliance forces you on-prem at enterprise scale.

The “which inference backend?” question keeps resurfacing on r/LocalLLaMA because these four tools get lumped together despite doing fundamentally different jobs. Two of them (Ollama, NIM) run models on hardware you control. Two of them (OpenRouter, Groq) are cloud APIs that happen to serve open-weight models. If you care about self-hosting, that split matters more than any benchmark.

What’s the actual difference between these four?

Ollama and NVIDIA NIM run inference on your own GPU; OpenRouter and Groq run it on someone else’s. Everything else follows from that.

TypeWhere inference runsLicense / termsOpenAI-compatible API
Ollama (v0.34.3, Sep 19 2026)Local model runnerYour machineMIT (open source)Yes, localhost:11434
OpenRouterCloud API routerThird-party providersProprietary serviceYes
GroqCloud inference on custom LPU chipsGroq’s datacentersProprietary service (serves open models)Yes
NVIDIA NIMContainerized model microservicesYour NVIDIA GPU (or cloud)Proprietary containers; free for dev via NVIDIA Developer ProgramYes

Ollama (MIT license, v0.34.3 as of September 19, 2026) pulls quantized GGUF models and serves them through a local REST API. OpenRouter is a routing layer: one API key, and as of September 2026 its catalog lists over 400 models — your request gets forwarded to whichever upstream provider hosts the model. Groq runs a much smaller catalog of open-weight models on its custom LPU hardware, selling pure speed. NIM packages models (Llama, Nemotron, and others) into Docker containers with TensorRT-LLM optimization — the weights may be open, but the container and runtime are proprietary NVIDIA software.

Which one keeps your prompts private?

Only Ollama and NIM keep prompt data on your network, and Ollama is the only one that’s fully open source.

Ollama runs inference entirely locally. The CLI reaches out to ollama.com to pull models and check for updates, but generation itself needs no network at all — you can pull a model, block outbound traffic on the box, and keep serving. That’s the test that matters: if the firewall goes up and the tool keeps working, your prompts aren’t going anywhere. (One caveat from our own testing: Ollama’s newer cloud-proxy models are the exception — see Ollama cloud models and privacy for which model tags route to Ollama’s servers.)

OpenRouter forwards your prompts to third-party providers — OpenAI, Anthropic, DeepInfra, whoever hosts the model you picked. OpenRouter itself doesn’t log prompt contents by default (it offers a 1% usage discount if you opt in to logging), and it added a Zero Data Retention setting that restricts routing to providers who commit to not storing data. But the honest framing is: your prompt is processed on someone else’s server under that provider’s retention policy. ZDR narrows the exposure; it doesn’t eliminate it.

Groq is cloud-only. Your prompts go to Groq’s datacenters, full stop. Its data-handling terms live at groq.com/legal and are reasonable for a commercial API, but there is no local deployment option for individuals.

NIM runs on your own GPU, so inference data stays local — that’s its genuine strength. You do need an NGC account and API key to pull containers, and the container phones NGC at startup to fetch model artifacts. For a hard air-gap, NVIDIA documents an offline deployment path, but it’s enterprise-grade friction.

If self-hosting is about not shipping your data to a third party, two of these four options fail the test by design.

How hard is each one to set up?

Ollama takes one command; NIM takes an account, a registry login, and datacenter-class hardware. The gap is enormous.

Ollama — install, pull, run:

$ curl -fsSL https://ollama.com/install.sh | sh
$ ollama run qwen3:32b
>>> Why is the sky blue?
The sky appears blue because of Rayleigh scattering...

First token in under a minute on a warm model. The API is live at http://localhost:11434/v1 and works with any OpenAI SDK.

OpenRouter — sign up, create a key, change two lines:

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
)

Groq — same pattern: sign up at console.groq.com, export GROQ_API_KEY, point the SDK at Groq’s endpoint. Five minutes.

NVIDIA NIM — you need an NVIDIA GPU with the Container Toolkit installed, an NGC account, and an API key with “NGC Catalog” scope:

$ echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
$ docker run --gpus all -e NGC_API_KEY nvcr.io/nim/meta/llama-3.3-70b-instruct:latest

A problem we hit doing exactly this: docker login nvcr.io returning 401 Unauthorized even with a valid key. The cause is an NGC key generated without the “NGC Catalog” service scope — the key authenticates but can’t pull from the catalog. Fix: regenerate the key at org.ngc.nvidia.com/setup/api-keys with NGC Catalog checked. Nothing in the error message tells you this.

Setup-time verdict: Ollama < OpenRouter ≈ Groq << NIM. And most LLM NIM containers target A100 80GB or H100-class GPUs — per NVIDIA’s own support matrix, only the smaller embedding and vision NIMs run comfortably on RTX-class cards. NIM is not a home-lab tool.

What does each one cost in September 2026?

Ollama’s marginal cost is electricity — around $0.04–$0.06 per GPU-hour. The cloud options bill per token, and the numbers below were checked September 2026.

Pricing modelVerified numbers (Sep 2026)
OllamaHardware you own + power350W GPU at $0.10–0.17/kWh ≈ $0.035–$0.06/hr
OpenRouterProvider list price, no per-token markup5.5% fee on credit purchases ($0.80 min)
GroqPer tokenGPT-OSS-20B: $0.075 in / $0.30 out per 1M; GPT-OSS-120B: $0.15 / $0.60
NIMHardware + licenseFree for development (Developer Program); production requires NVIDIA AI Enterprise

The Ollama math assumes you already own the GPU. As of September 2026 a used RTX 3090 runs about $1,252 — up roughly 24% since March, because the DRAM crisis has repriced the whole used market. A used RTX 4090 sits around $2,150–$2,350. That sunk cost is real, which is why “rent first” is the right move if you don’t already own hardware: on Vast.ai, RTX 3090s rent from $0.07/hr and 4090s from $0.14/hr (marketplace prices float — those are September 2026 floors).

One pricing event worth knowing about if you build on Groq: on August 26, 2026, Groq moved Llama 3.1 8B and Llama 3.3 70B to enterprise-only “contact sales” pricing. If your pipeline had a Groq Llama model pinned, it broke — that’s the failure mode of building on a hosted catalog you don’t control. The fix that worked for us: switch the model string to openai/gpt-oss-120b (still self-serve on Groq), or route through OpenRouter with a fallback list so a delisted model degrades instead of erroring. A local Ollama model can’t be delisted out from under you.

When is Groq’s speed actually worth it?

When you need interactive latency on a model too big for your VRAM — that’s the one case where Groq is the clear pick.

Independent benchmarking by Artificial Analysis measured Groq serving Llama 3.3 70B at around 276 tokens/second (Groq’s own figures claim higher still). A 24GB consumer card running a 27–32B model at Q4 typically generates 30–50 tokens/second, and a 70B doesn’t fit in 24GB at all without heavy offloading that drops you to single digits. So for a latency-sensitive app on a 70B-class model — a voice assistant, a real-time agent loop — Groq is roughly an order of magnitude faster than anything a single consumer GPU will do.

For batch work (overnight summarization, embedding pipelines, evals), speed per request matters much less than cost and privacy, and local inference wins back the advantage.

When NOT to use each one

  • Don’t use Ollama when the model genuinely doesn’t fit your hardware. A 24GB card tops out around 32B dense models at Q4 with usable context. If your workload needs a 400B-class frontier model, no amount of self-hosting enthusiasm changes the VRAM math — rent a bigger GPU or use an API for that one job.
  • Don’t use OpenRouter for anything covered by confidentiality obligations — client code, medical notes, legal drafts — unless ZDR mode plus the specific provider’s terms actually satisfy your requirements. Read the provider-routing page, not just the homepage.
  • Don’t use Groq as your only backend. The August 2026 catalog change showed that self-serve models can vanish behind a sales form with two weeks’ notice.
  • Don’t use NIM in a home lab. Between the NGC account, proprietary containers, A100/H100-class targets, and AI Enterprise licensing for production, it solves an enterprise compliance problem you probably don’t have. LocalAI or Ollama gives you the same OpenAI-compatible local endpoint with none of that — see LocalAI vs Ollama.

Which should you pick?

Ollama for anything private or recurring; OpenRouter for occasional big-model calls; Groq for speed-critical hosted inference; NIM only under enterprise constraints.

Costs as of September 2026, all verified above:

Your situationPickCostWhere
Privacy matters and the model fits in 24GBOllama on a used RTX 3090~$1,252 once + powerCheck price
Need occasional 400B-class model accessOpenRouter key alongside local Ollamalist price + 5.5% credit feeopenrouter.ai
Latency-critical app on a 70B-class modelGroq (GPT-OSS-120B self-serve)$0.15/$0.60 per 1M tokensconsole.groq.com
Want to test workloads before buying a GPURented RTX 3090/4090from $0.07/hrVast.ai
On-prem mandate, enterprise budgetNVIDIA NIMhardware + AI Enterprise licensebuild.nvidia.com

The pattern most home labs land on isn’t one platform — it’s Ollama as the default plus one cloud key for overflow. That combination covers 95% of jobs privately at zero marginal cost and rents the last 5% instead of buying an 8-GPU server for it. If you’re choosing between local runners rather than local-vs-cloud, llamafile vs Ollama vs LM Studio covers that decision; for the GPU hardware side, our sister site runaihome.com tracks what current street prices do to the build math, and aicoderscope.com covers wiring any of these endpoints into Cursor or Cline as a BYOK backend.

FAQ

Can Ollama serve multiple users like OpenRouter or Groq? Yes, with limits. Ollama v0.34.x handles parallel requests via OLLAMA_NUM_PARALLEL, but each concurrent request consumes KV-cache VRAM, so a 24GB card realistically serves 2–4 simultaneous users on a mid-size model. For more, you’re into multi-GPU territory or vLLM.

Is NVIDIA NIM open source? No. NIM containers wrap open-weight models in proprietary NVIDIA runtime software (TensorRT-LLM optimized). Development use is free through the NVIDIA Developer Program; production deployment is licensed through NVIDIA AI Enterprise. The models can be open; the serving stack isn’t.

Does OpenRouter train on my prompts? OpenRouter says it doesn’t log prompt contents by default and offers a Zero Data Retention routing mode. But requests are processed by the upstream provider serving your chosen model, under that provider’s retention policy — check the per-provider logging table in OpenRouter’s docs before sending anything sensitive.

Sources

  • Used RTX 3090 24GB — ~$1,252 as of September 2026; still the cheapest 24GB CUDA card for running Ollama at home.
  • Used RTX 4090 24GB — ~$2,150–$2,350 as of September 2026; roughly 1.7× the 3090’s price for about 2× the generation speed.

Was this article helpful?

What self-hosting actually costs

Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.