
Strata
Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
The Lens
By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.
Updated Sep 2026
Strata runs Qwen3.8-Flash-Next, a very large open-weight model, on a gaming PC. The model has 125B parameters but only 6B do work on each token, so Strata keeps the busiest parts on your graphics card and the rest in system RAM, where the CPU handles them. It serves an OpenAI- and Anthropic-compatible API on localhost. The engine is MIT and free; the model weights carry Qwen's own community license.
RAM is the real requirement, not the GPU. Plan on an NVIDIA RTX 30, 40 or 50 card with 12 GB or more (8 GB runs, slowly), 48 to 64 GB of system RAM (32 GB for the coding-only variant), and 70 to 80 GB of free disk. On the author's RTX 5070 with 64 GB of DDR5, it writes about 40 to 90 tokens per second, depending on model size and context length. It answers one request at a time.
Solo developers with the hardware: a way to run a frontier-sized open model locally at $0 per token. For general local models, llama.cpp (ggml-org/llama.cpp) and Ollama (ollama/ollama) are the mature picks.
The catch: the project is days old with one maintainer, and the Windows installer downloads a hand-built strata.exe and runs it without checking a hash. Cautious users should build from source (--build). Also read the setup prompts: the "experimental speed projection" it offers is, by its own docs, a vector that strips the model's refusals, not a speed boost.
Free vs Self-Hosted vs Paid
fully freeFree: The Strata engine is MIT licensed with no paid tier, account or cloud. The model weights are not MIT: Qwen3.8-Flash-Next is under the Qwen Community License 1.0, which requires a separate license from Qwen for commercial Model-as-a-Service or AI work assistant businesses. Strata downloads third-party quantized versions (ISTA-DASLab GGUFs) from Hugging Face.
Self-hosted (your own PC): The cost is hardware.
- NVIDIA RTX 30/40/50 with 12 GB+ VRAM (8 GB runs, slowly), driver 580+
- 64 GB RAM fits every model size; 48 GB fits the two smallest; the Coder variant fits 32 GB
- 70 to 80 GB free disk
- Windows 10/11 or Linux; AMD RX 7900 XT/XTX is experimental on Linux
Paid: None. The alternative is paying per token for Qwen's hosted Qwen3.8-Flash.
Free engine, $0 per token, but only if you already own a 12 GB GPU and 48 to 64 GB of RAM.
What to do by team size
- Solo
- free if you have the RAM and GPU; build from source if you don't trust the prebuilt exe
- Small team
- free, but one request at a time makes it a personal tool
- Medium team
- use vLLM or a hosted API for concurrent users
- Large team
- not built for serving; use vLLM on server GPUs or a hosted API
Get tools like this every Wednesday
One featured tool, three on the radar. No fluff.
A low score is not a verdict on quality. Young and niche tools start low by design. How we calculate scores
License: MIT License
Use freely, including commercial. Just keep the license.
Commercial use: ✓ Yes
About
- Owner
- Niko (User)
- Stars
- 1,234
- Forks
- 147
Explore Further
More tools in the directory
openclaw
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
390.8k ★everything-claude-code
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
269.2k ★hermes-agent
The agent that grows with you
249.9k ★