A 125B Qwen model on a gaming PC: what Strata does, and the catch
A free project called Strata runs Qwen3.8-Flash-Next, a 125-billion-parameter open model, on an ordinary gaming PC. It was the top story on Hacker News on October 4. Before you install it, two details from the comments matter.
What happened
Strata is a free, open-source installer that runs Qwen3.8-Flash-Next on a home computer. Its README asks for an NVIDIA RTX 20-series or newer card, or a recent AMD Radeon, with at least 12 GB of video memory, 32 GB of RAM or more, about 80 GB of disk, and Windows or Linux. The project says nothing leaves your PC.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
The model itself is not new. Qwen published it on Hugging Face in August and calls it an experimental preview of the architecture that will underpin Qwen4. According to its model card, it has 125 billion parameters, of which only 6 billion are active for each token, reads text and images, and handles 262,144 tokens of context natively. It is released under Qwen's own community license, not a standard open-source one. What is new this week is that someone made it run on hardware people already own.
The speed, and where it comes from
Strata's README reports 94 tokens per second when writing answers on an RTX 5070 with 12 GB, using its smallest version (Q2_0), and 53 tokens per second with the larger IQ3_S version. One Hacker News commenter reported 124 tokens per second on an RTX 4090 with 128 GB of RAM. For scale, the README notes that 60 tokens per second is already faster than you can read.
The catch is in those version names. They are 2-bit and 3-bit quantizations: copies of the model compressed far below its original precision so that it fits. Several replies on Hacker News were skeptical. One developer said they would not go below 4-bit for real work and run a 4-bit copy on a rented RTX Pro 6000 for about $1 an hour instead. Strata's speed figures are real, but they are for a smaller, rougher copy of the model than the one Qwen benchmarked.
The install step to read first
The recommended install is unusual. Instead of commands, Strata's README gives you one line to paste into your AI coding assistant (Claude Code, Cursor, Codex or Copilot), which then follows a setup document from the repository on your machine. One commenter compared it to piping a download straight into your shell. Nothing suggests Strata's document is unsafe, but the habit is worth thinking about: an agent following someone else's instructions can install drivers and run scripts with your permissions. Read the document first, or ask a model to read it for you without running anything.
What it means for you
If you have a gaming PC with 12 GB of video memory or more and want a private model for drafting or coding, this is the easiest way so far to try a model of this size at home. Treat it as a test, not a replacement for ChatGPT or Claude: give it the same three tasks you gave your usual model last week and compare the answers before you rely on it. If quality matters more than privacy, a 4-bit copy on a rented GPU or an API is the safer bet.
Try it
Read this setup document and every script it tells you to download, but do not run anything: [paste the link or the document]. List every command it would run, everything it downloads and from where, and everything it changes outside the project folder. Flag anything that needs administrator rights, pipes a download into a shell, or sends data off this machine. Quote the exact line for each. End with a verdict: safe for an agent to run, run by hand step by step, or do not run.
The full version is in the library: Setup Doc Safety Check.
Sources
- Strata, README on GitHub: requirements, speed tables, install
- Qwen, Qwen3.8-Flash-Next model card: parameters, context length, license
- Hacker News, discussion thread, October 4, 2026