Qwen3.8-27B Self-Hosting Guide 2026: Ollama, GGUF, vLLM
TL;DR: Qwen3.8-27B is a dense 27B multimodal model under plain Apache 2.0, with open weights on Hugging Face since August 14, 2026. A Q4_K_M GGUF is 17.1–17.8GB, which makes a 24GB card the comfortable floor. ollama pull qwen3.8:27b works out of the box on Ollama v0.32.12 or newer; vLLM needs 0.17+. As of September 2026 this is the default model to run at the 24GB tier.
| The short answer | |
|---|---|
| License | Apache 2.0, no MAU caps, no attribution clause (verified Sep 2026) |
| Architecture | Dense 27B + vision tower + MTP draft head — not an MoE |
| Context | 262,144 tokens native (Ollama lists it as 256K — same number) |
| Weights | ~56GB BF16; Q4_K_M GGUF ~17GB |
| Hardware floor | 24GB VRAM for Q4_K_M with usable context |
| Ollama | ollama pull qwen3.8:27b (18GB download), requires v0.32.12+ |
| vLLM | Supported since 0.17; official recipe uses the NVFP4 quant |
Honest take: If you already run Qwen3.6-35B-A3B and care most about tokens per second, this is a sidegrade, not an upgrade — dense 27B decodes slower than a 3B-active MoE. Switch for the vision input, the long context, or the quality ceiling, not for speed.
Is Qwen3.8-27B actually Apache 2.0?
Yes — plain Apache 2.0, verified against the license file on the Hugging Face repo (Qwen/Qwen3.8-27B). No monthly-active-user cap, no attribution requirement, no field-of-use restriction, and none of the Tongyi Qianwen license terms that some expected when the 2.4T flagship shipped API-only in early August 2026.
That matters because Qwen’s own history made it worth checking: Alibaba has used the restrictive Qwen License for its 72B-class models while keeping ≤32B models Apache 2.0. Qwen3.8-27B follows the clean pattern. Commercial self-hosted inference, fine-tuning, and redistribution are all unrestricted. If you track licenses across the ecosystem, the open-source LLM license shootout has the full map.
Is Qwen3.8-27B dense or a MoE?
Dense. The model card lists a single 27B-parameter transformer (Qwen3_5ForConditionalGeneration) with a vision config and a built-in multi-token-prediction (MTP) draft head — it is a smaller sibling of the 2.4T Qwen3.8-Max flagship, not a shrunken MoE.
Three practical consequences:
- All 27B parameters are active every token. Decode speed will be noticeably slower than Qwen3.6-35B-A3B, whose MoE design only moves ~3B parameters per token. If throughput is your constraint, the Qwen3.6-35B-A3B setup guide covers the faster option.
- VRAM math is simple. Unlike MoE models, there is no “active parameters fit but total parameters don’t” trap. If the quant file fits in VRAM with KV-cache headroom, it runs.
- The MTP draft head gives you speculative decoding for free in engines that support it, which claws back some of the dense-model speed penalty.
It is also natively multimodal — image and video input, not a bolted-on adapter — which neither Qwen3.6-35B-A3B nor most 24GB-class competitors offer.
How much VRAM does Qwen3.8-27B need?
Around 17GB for a Q4_K_M GGUF, plus KV-cache. Measured file sizes from the two main community quant sources, September 2026:
| Quant | Size | Fits on | Notes |
|---|---|---|---|
| BF16 (original) | ~56GB | 3× 24GB or A100 80GB | Full precision, vLLM territory |
| Q4_K_M (Unsloth) | 17.11GB | 24GB card | Unsloth Dynamic v3.0 quants |
| Q4_K_M (Bartowski) | 17.77GB | 24GB card | Conventional quant |
| UD-Q4_K_XL (Unsloth) | 17.9GB | 24GB card | Best quality-per-GB at 4-bit |
| Ollama default tag | 18GB download | 24GB card | Q4_K_M under the hood |
On a 24GB card (RTX 3090/4090), a ~17GB quant leaves roughly 6GB for KV-cache — enough for tens of thousands of tokens of context, nowhere near the full 262K. On 16GB cards the Q4 tiers do not fit with usable context; you would need Q3-class quants and the quality cost of going below 4-bit is rarely worth it on a 27B. The general size-vs-quality tradeoff is covered in the quantization guide.
GGUF sources: unsloth/Qwen3.8-27B-GGUF and Bartowski’s repo on Hugging Face. Unsloth also publishes an NVFP4 build (unsloth/Qwen3.8-27B-NVFP4) aimed at vLLM on Blackwell hardware.
How do you run Qwen3.8-27B with Ollama?
One command, if your Ollama is current:
ollama pull qwen3.8:27b # 18GB download
ollama run qwen3.8:27b
>>> Describe what a dense multimodal model is in one sentence.
A dense multimodal model is a neural network in which every parameter
participates in processing each input token, and which accepts more
than one input type — such as text and images — in the same context.
The bare qwen3.8 tag resolves to the same 27B build, and qwen3.8:27b-mlx exists for Apple Silicon. The library listing shows text and image input enabled and a 256K context window.
Ollama shipped day-one support in v0.32.12 (August 14, 2026). That version floor is the single most common failure point — see the fix below.
For a front end, the pairing that works well is Open WebUI, which picked up image-input handling in its 0.11 overhaul — setup in the Open WebUI 0.11 guide.
Fix: “unknown model architecture” or a failed pull on older Ollama
The problem you will actually hit: any Ollama older than v0.32.12 cannot run this model. Depending on how stale the install is, the failure shows up either as a manifest error during ollama pull or as an unknown model architecture error at load time, because the Qwen3_5ForConditionalGeneration architecture support simply isn’t in the binary.
The fix is a version check and an upgrade, not a config change:
ollama --version # needs 0.32.12 or newer
curl -fsSL https://ollama.com/install.sh | sh # Linux upgrade-in-place
sudo systemctl restart ollama
Docker users: pull ollama/ollama:latest and recreate the container — the model store in the volume survives.
The second common stumble is context. Ollama defaults to a small context window, and naively setting 262K will OOM the KV-cache on a 24GB card. Set num_ctx to what your VRAM actually affords (32K is a sane starting point at Q4 on 24GB) and raise it only while ollama ps still shows the model 100% on GPU.
Does vLLM support Qwen3.8-27B?
Yes, since vLLM 0.17. The official vLLM recipe for this model serves the NVFP4 quant across two GPUs:
vllm serve unsloth/Qwen3.8-27B-NVFP4 \
--tensor-parallel-size 2 \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml
Notes from the recipe worth keeping: the qwen3 reasoning parser and qwen3_xml tool-call parser are what make agent frameworks and OpenAI-compatible clients behave; --kv-cache-dtype fp8 is what makes the 262K window plausible at all. Full-precision BF16 serving needs ~56GB of weights before KV-cache, so single-card home setups should stay on the Ollama/llama.cpp GGUF path and leave vLLM for multi-GPU or datacenter-class cards. The tradeoff between the two engines hasn’t changed since we wrote Ollama vs vLLM: vLLM for concurrent throughput, Ollama for one-user simplicity.
What does the hardware actually cost in September 2026?
The entry ticket is a 24GB card, and 2026 GPU prices are not kind. Verified street prices as of September 2026: a used RTX 3090 runs $1,150–$1,350, a used RTX 4090 $2,150–$2,350, and the 16GB RTX 5060 Ti ($679–$805, against a $429 MSRP) doesn’t clear the VRAM bar for this model anyway. If you’re buying, the used RTX 3090 remains the cheapest 24GB entry; runaihome.com’s GPU guides cover the full buy decision.
Before spending four figures, rent the exact workload first: a 3090 on Vast.ai starts from about $0.07/hr (marketplace pricing, September 2026), so an evening of testing Qwen3.8-27B against your real prompts costs less than a coffee.
When NOT to self-host Qwen3.8-27B
- You have 16GB of VRAM or less. Sub-4-bit quants of a dense 27B lose real quality. Run Gemma 4 26B-A4B QAT (~15GB) or a 12B-class model instead — the August 2026 leaderboard maps every VRAM tier.
- Throughput is the whole point. Qwen3.6-35B-A3B decodes several times faster at the same VRAM budget thanks to its 3B-active MoE design.
- Your workload is agent loops. Meta’s Muse Glimmer 30B (also Apache 2.0, also 24GB-class) was tuned specifically for tool-calling; test both before committing.
- You need the full 262K context. At home you won’t have the VRAM for it. Long-context serving of this model is rented-GPU territory, not a 3090 job.
- You’d use it via a coding tool anyway. If the only consumer is Cursor or Cline, weigh a hosted endpoint first — aicoderscope.com covers the BYOK backend angle.
FAQ
What is the exact Ollama command for Qwen3.8-27B?
ollama pull qwen3.8:27b, then ollama run qwen3.8:27b. The download is 18GB (Q4_K_M). Requires Ollama v0.32.12 or newer; older versions fail with a manifest or unknown-architecture error.
Can Qwen3.8-27B run on a 16GB GPU? Not well. The Q4_K_M GGUF alone is ~17GB. You would need a Q3-class quant with almost no KV-cache headroom, and quality drops noticeably below 4-bit on dense 27B models. A 24GB card is the practical floor.
Is Qwen3.8-27B free for commercial use? Yes. The weights ship under plain Apache 2.0 (verified September 2026) — no user caps, no revenue clauses, no attribution requirement. Commercial inference, fine-tuning, and redistribution are all permitted.
Sources
- Qwen/Qwen3.8-27B on Hugging Face — weights, license, model card
- Ollama library: qwen3.8 tags — pull tags and download sizes
- unsloth/Qwen3.8-27B-GGUF — Dynamic v3.0 GGUF quants
- vLLM recipe: Qwen3.8-27B — official serve command and flags
- Yotta Labs — Qwen 3.8 27B specs and hardware requirements
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →What self-hosting actually costs
Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.