Tools/Niko1221/Strata

Strata

Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

1.2k ★emergingC++MIT License →new this week

The Lens

By Erik Loyd, SaaS CEO and former COO/CFO of an AWS Premier Partner.

Updated Sep 2026

Strata runs Qwen3.8-Flash-Next, a very large open-weight model, on a gaming PC. The model has 125B parameters but only 6B do work on each token, so Strata keeps the busiest parts on your graphics card and the rest in system RAM, where the CPU handles them. It serves an OpenAI- and Anthropic-compatible API on localhost. The engine is MIT and free; the model weights carry Qwen's own community license.

RAM is the real requirement, not the GPU. Plan on an NVIDIA RTX 30, 40 or 50 card with 12 GB or more (8 GB runs, slowly), 48 to 64 GB of system RAM (32 GB for the coding-only variant), and 70 to 80 GB of free disk. On the author's RTX 5070 with 64 GB of DDR5, it writes about 40 to 90 tokens per second, depending on model size and context length. It answers one request at a time.

Solo developers with the hardware: a way to run a frontier-sized open model locally at $0 per token. For general local models, llama.cpp (ggml-org/llama.cpp) and Ollama (ollama/ollama) are the mature picks.

The catch: the project is days old with one maintainer, and the Windows installer downloads a hand-built strata.exe and runs it without checking a hash. Cautious users should build from source (--build). Also read the setup prompts: the "experimental speed projection" it offers is, by its own docs, a vector that strips the model's refusals, not a speed boost.

Free vs Self-Hosted vs Paid

fully free

Free: The Strata engine is MIT licensed with no paid tier, account or cloud. The model weights are not MIT: Qwen3.8-Flash-Next is under the Qwen Community License 1.0, which requires a separate license from Qwen for commercial Model-as-a-Service or AI work assistant businesses. Strata downloads third-party quantized versions (ISTA-DASLab GGUFs) from Hugging Face.

Self-hosted (your own PC): The cost is hardware.

  • NVIDIA RTX 30/40/50 with 12 GB+ VRAM (8 GB runs, slowly), driver 580+
  • 64 GB RAM fits every model size; 48 GB fits the two smallest; the Coder variant fits 32 GB
  • 70 to 80 GB free disk
  • Windows 10/11 or Linux; AMD RX 7900 XT/XTX is experimental on Linux

Paid: None. The alternative is paying per token for Qwen's hosted Qwen3.8-Flash.

Free engine, $0 per token, but only if you already own a 12 GB GPU and 48 to 64 GB of RAM.

What to do by team size

Solo
free if you have the RAM and GPU; build from source if you don't trust the prebuilt exe
Small team
free, but one request at a time makes it a personal tool
Medium team
use vLLM or a hosted API for concurrent users
Large team
not built for serving; use vLLM on server GPUs or a hosted API
Self-hosting ops:moderate

Get tools like this every Wednesday

One featured tool, three on the radar. No fluff.

Score
68/100 · B
Adoption13/30
Maintenance25/25
Community5/20
License15/15
Analysis10/10

A low score is not a verdict on quality. Young and niche tools start low by design. How we calculate scores

License: MIT License

Use freely, including commercial. Just keep the license.

Commercial use: ✓ Yes

About

Owner
Niko (User)
Stars
1,234
Forks
147

Explore Further

More tools in the directory