Qwen3.8-27B Self-Hosting Guide 2026: Ollama, GGUF, vLLM

qwenollamavllmggufselfhosted

TL;DR: Qwen3.8-27B is a dense 27B multimodal model under plain Apache 2.0, with open weights on Hugging Face since August 14, 2026. A Q4_K_M GGUF is 17.1–17.8GB, which makes a 24GB card the comfortable floor. ollama pull qwen3.8:27b works out of the box on Ollama v0.32.12 or newer; vLLM needs 0.17+. As of September 2026 this is the default model to run at the 24GB tier.

The short answer
LicenseApache 2.0, no MAU caps, no attribution clause (verified Sep 2026)
ArchitectureDense 27B + vision tower + MTP draft head — not an MoE
Context262,144 tokens native (Ollama lists it as 256K — same number)
Weights~56GB BF16; Q4_K_M GGUF ~17GB
Hardware floor24GB VRAM for Q4_K_M with usable context
Ollamaollama pull qwen3.8:27b (18GB download), requires v0.32.12+
vLLMSupported since 0.17; official recipe uses the NVFP4 quant

Honest take: If you already run Qwen3.6-35B-A3B and care most about tokens per second, this is a sidegrade, not an upgrade — dense 27B decodes slower than a 3B-active MoE. Switch for the vision input, the long context, or the quality ceiling, not for speed.

Is Qwen3.8-27B actually Apache 2.0?

Yes — plain Apache 2.0, verified against the license file on the Hugging Face repo (Qwen/Qwen3.8-27B). No monthly-active-user cap, no attribution requirement, no field-of-use restriction, and none of the Tongyi Qianwen license terms that some expected when the 2.4T flagship shipped API-only in early August 2026.

That matters because Qwen’s own history made it worth checking: Alibaba has used the restrictive Qwen License for its 72B-class models while keeping ≤32B models Apache 2.0. Qwen3.8-27B follows the clean pattern. Commercial self-hosted inference, fine-tuning, and redistribution are all unrestricted. If you track licenses across the ecosystem, the open-source LLM license shootout has the full map.

Is Qwen3.8-27B dense or a MoE?

Dense. The model card lists a single 27B-parameter transformer (Qwen3_5ForConditionalGeneration) with a vision config and a built-in multi-token-prediction (MTP) draft head — it is a smaller sibling of the 2.4T Qwen3.8-Max flagship, not a shrunken MoE.

Three practical consequences:

  1. All 27B parameters are active every token. Decode speed will be noticeably slower than Qwen3.6-35B-A3B, whose MoE design only moves ~3B parameters per token. If throughput is your constraint, the Qwen3.6-35B-A3B setup guide covers the faster option.
  2. VRAM math is simple. Unlike MoE models, there is no “active parameters fit but total parameters don’t” trap. If the quant file fits in VRAM with KV-cache headroom, it runs.
  3. The MTP draft head gives you speculative decoding for free in engines that support it, which claws back some of the dense-model speed penalty.

It is also natively multimodal — image and video input, not a bolted-on adapter — which neither Qwen3.6-35B-A3B nor most 24GB-class competitors offer.

How much VRAM does Qwen3.8-27B need?

Around 17GB for a Q4_K_M GGUF, plus KV-cache. Measured file sizes from the two main community quant sources, September 2026:

QuantSizeFits onNotes
BF16 (original)~56GB3× 24GB or A100 80GBFull precision, vLLM territory
Q4_K_M (Unsloth)17.11GB24GB cardUnsloth Dynamic v3.0 quants
Q4_K_M (Bartowski)17.77GB24GB cardConventional quant
UD-Q4_K_XL (Unsloth)17.9GB24GB cardBest quality-per-GB at 4-bit
Ollama default tag18GB download24GB cardQ4_K_M under the hood

On a 24GB card (RTX 3090/4090), a ~17GB quant leaves roughly 6GB for KV-cache — enough for tens of thousands of tokens of context, nowhere near the full 262K. On 16GB cards the Q4 tiers do not fit with usable context; you would need Q3-class quants and the quality cost of going below 4-bit is rarely worth it on a 27B. The general size-vs-quality tradeoff is covered in the quantization guide.

GGUF sources: unsloth/Qwen3.8-27B-GGUF and Bartowski’s repo on Hugging Face. Unsloth also publishes an NVFP4 build (unsloth/Qwen3.8-27B-NVFP4) aimed at vLLM on Blackwell hardware.

How do you run Qwen3.8-27B with Ollama?

One command, if your Ollama is current:

ollama pull qwen3.8:27b   # 18GB download
ollama run qwen3.8:27b
>>> Describe what a dense multimodal model is in one sentence.
A dense multimodal model is a neural network in which every parameter
participates in processing each input token, and which accepts more
than one input type — such as text and images — in the same context.

The bare qwen3.8 tag resolves to the same 27B build, and qwen3.8:27b-mlx exists for Apple Silicon. The library listing shows text and image input enabled and a 256K context window.

Ollama shipped day-one support in v0.32.12 (August 14, 2026). That version floor is the single most common failure point — see the fix below.

For a front end, the pairing that works well is Open WebUI, which picked up image-input handling in its 0.11 overhaul — setup in the Open WebUI 0.11 guide.

Fix: “unknown model architecture” or a failed pull on older Ollama

The problem you will actually hit: any Ollama older than v0.32.12 cannot run this model. Depending on how stale the install is, the failure shows up either as a manifest error during ollama pull or as an unknown model architecture error at load time, because the Qwen3_5ForConditionalGeneration architecture support simply isn’t in the binary.

The fix is a version check and an upgrade, not a config change:

ollama --version        # needs 0.32.12 or newer
curl -fsSL https://ollama.com/install.sh | sh   # Linux upgrade-in-place
sudo systemctl restart ollama

Docker users: pull ollama/ollama:latest and recreate the container — the model store in the volume survives.

The second common stumble is context. Ollama defaults to a small context window, and naively setting 262K will OOM the KV-cache on a 24GB card. Set num_ctx to what your VRAM actually affords (32K is a sane starting point at Q4 on 24GB) and raise it only while ollama ps still shows the model 100% on GPU.

Does vLLM support Qwen3.8-27B?

Yes, since vLLM 0.17. The official vLLM recipe for this model serves the NVFP4 quant across two GPUs:

vllm serve unsloth/Qwen3.8-27B-NVFP4 \
  --tensor-parallel-size 2 \
  --max-model-len 262144 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml

Notes from the recipe worth keeping: the qwen3 reasoning parser and qwen3_xml tool-call parser are what make agent frameworks and OpenAI-compatible clients behave; --kv-cache-dtype fp8 is what makes the 262K window plausible at all. Full-precision BF16 serving needs ~56GB of weights before KV-cache, so single-card home setups should stay on the Ollama/llama.cpp GGUF path and leave vLLM for multi-GPU or datacenter-class cards. The tradeoff between the two engines hasn’t changed since we wrote Ollama vs vLLM: vLLM for concurrent throughput, Ollama for one-user simplicity.

What does the hardware actually cost in September 2026?

The entry ticket is a 24GB card, and 2026 GPU prices are not kind. Verified street prices as of September 2026: a used RTX 3090 runs $1,150–$1,350, a used RTX 4090 $2,150–$2,350, and the 16GB RTX 5060 Ti ($679–$805, against a $429 MSRP) doesn’t clear the VRAM bar for this model anyway. If you’re buying, the used RTX 3090 remains the cheapest 24GB entry; runaihome.com’s GPU guides cover the full buy decision.

Before spending four figures, rent the exact workload first: a 3090 on Vast.ai starts from about $0.07/hr (marketplace pricing, September 2026), so an evening of testing Qwen3.8-27B against your real prompts costs less than a coffee.

When NOT to self-host Qwen3.8-27B

  • You have 16GB of VRAM or less. Sub-4-bit quants of a dense 27B lose real quality. Run Gemma 4 26B-A4B QAT (~15GB) or a 12B-class model instead — the August 2026 leaderboard maps every VRAM tier.
  • Throughput is the whole point. Qwen3.6-35B-A3B decodes several times faster at the same VRAM budget thanks to its 3B-active MoE design.
  • Your workload is agent loops. Meta’s Muse Glimmer 30B (also Apache 2.0, also 24GB-class) was tuned specifically for tool-calling; test both before committing.
  • You need the full 262K context. At home you won’t have the VRAM for it. Long-context serving of this model is rented-GPU territory, not a 3090 job.
  • You’d use it via a coding tool anyway. If the only consumer is Cursor or Cline, weigh a hosted endpoint first — aicoderscope.com covers the BYOK backend angle.

FAQ

What is the exact Ollama command for Qwen3.8-27B? ollama pull qwen3.8:27b, then ollama run qwen3.8:27b. The download is 18GB (Q4_K_M). Requires Ollama v0.32.12 or newer; older versions fail with a manifest or unknown-architecture error.

Can Qwen3.8-27B run on a 16GB GPU? Not well. The Q4_K_M GGUF alone is ~17GB. You would need a Q3-class quant with almost no KV-cache headroom, and quality drops noticeably below 4-bit on dense 27B models. A 24GB card is the practical floor.

Is Qwen3.8-27B free for commercial use? Yes. The weights ship under plain Apache 2.0 (verified September 2026) — no user caps, no revenue clauses, no attribution requirement. Commercial inference, fine-tuning, and redistribution are all permitted.

Sources

Was this article helpful?

What self-hosting actually costs

Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.