Thanks, trying it now. On my 3090 it seems to run at around 60tps, the IQ3_S variant. I am testing it now to see if it is better than Qwen 3.8 27b
a11r
I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I'm running 4-bit quants on an RTX Pro 6000 rented for approximately $1/hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: https://www.reddit.com/r/BlackwellPerformance/s/FrKwk3GoDK
show comments
Jackson__
I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.
To put that into perspective, here are some more numbers from other models via llama.cpp:
Median/Average
Qwen 3.5 9B BF16: 46.5 / 193.3
Qwen 3.6 35B Q4 K XL: 38.4 / 76.4
Qwen 3.5 122B Q3 K M: 32.9 / 68.6
The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.
I have done no further testing, as these results line up perfectly with my expectations.
show comments
AntiRush
I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.
Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:
I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.
LLM threads the world over are spammed with Strata links, it remains to be seen how much of the breathless hype remains standing once the honeymoon period is over. I've tried it but so far I have not seen anything that overly impressed me in terms of accuracy, though the speed is definitely there. I'm sure there are applications for LLMs where the quality of the answers is less important but I don't have any of those. YMMV.
cuvinny
Haven't had time to do much quality testing but the IQ2 model is running at 65 tps on a 9070xt/5900x. That is wild.
kamranjon
Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.
Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?
show comments
mmaunder
More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
show comments
SuperV1234
We're getting closer and closer to the day we can have an Opus-like model running locally. The dream!
show comments
zkmon
I don't get it. It's file size is about 6 times larger than 27B model for the same quant, but the performance improvement is hardly 10% across all benchmarks, according the metrics on it's hf page. Why should one devote so much more hardware for so little benefit?
Tepix
Q2 quantization. Not interested.
show comments
Tepix
All headlines about LLM performance MUST have the quantization also mentioned in the headline.
You know, so you're not wasting your time like in this post.
show comments
Luker88
Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.
Surprisingly useful as long as you can leave it running a couple of hours at the very least.
While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.
show comments
prettyblocks
I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).
nialv7
There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
show comments
hecturchi
- Tiny context size or hours to load it
- Hard to benefit from thinking and preserve thinking given token cost.
- Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.
- K/V quants probably quantized too make things less accurate.
Useful would be combinations with:
- Full context size so it can code and think a bit.
- Draft MTP <= 2 so it doesn't trip
- Q4 quants or better so its accurate
- q8 cache or better so it stays accurate.
- 20 token/s so it finishes while reviewing previous step.
- And enough left RAM for 50+ context checkpoints so that it can progress quuckly.
Closest you have is Qwen3.6-35B-A3B-MTP.
Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.
Source: I have low specs and tried them all for agentic use + coding.
show comments
b212
I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.
I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?
show comments
1-6
I hope we're coming to a plateau with the HBM/GDDR7/on-chip RAM hype and get back to normalcy with system RAM alternatives for the rest of us.
pilooch
My goto private setup, runs ~50t/sex on a dgx spark with sglang, nvfp4. Excellent model.
show comments
fsiefken
I wonder if a higher Qwen3.8-27b quant could beat or match these lower < 16/24/48/64G Qwen3.8-Flash Next quants given similar quality.
What speed are you willing the sacrifice to debug/program for more complex jobs faster?
Anyone know how it compares to GLM 5.3 for real world use?
show comments
mark_l_watson
Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.
Progress on running local models has been amazing.
show comments
sweetboy
I think it would be great if you could try models with lesser parameters that could fit on 6GB VRAM-ish, which could work for "gaming laptops" as well.
show comments
paulez
Pretty impressive so far, but needs more testing.
It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.
Local LLM is getting more exciting every day!
show comments
hemedanmert
I need a version of this that runs 3.8 27B on 8 gigs of VRAM
amazing project, congrats on the launch
ai_ja_nai
I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
show comments
hypfer
Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?
The Readme doesn't say, but it's all AI generated, so..
show comments
jameslholcombe
I might try combining this a FreeToken
esafak
Has anyone calculated the effective intelligence of these quantized models?
I think publishing benchmarks with quantized models should become standard practice.
show comments
Neywiny
I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
show comments
Jeeetendra
getting it to fit is impressive, but i'd want to compare the smaller quants on a real coding task before picking one. how much quality do you lose going from IQ3_S to Q2_0?
show comments
bt1a
80 t/s w/ 3090s and 3.05bpw exllamav3
lousken
lm studio bionic, unsloth, now this... it would be nice if it worked at least in one of those without installing another component
show comments
boredatoms
It hard to take below-8bit quants seriously
cesarvarela
It's funny that most of the AI industry is built around the assumption (which is most likely true) that it is not possible to run SOTA models on current consumer hardware.
Imagine if someone managed to run an Astra- or Fable-level model on a 5090 at reasonable speeds.
show comments
quietFalcon
Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
show comments
0xbadcafebee
Lol, sure, if you quant it to hell (Q2) it'll go real fast...
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
show comments
panny
I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
show comments
lsb
There’s other slop projects to run of Qwen, like ds4, would be interesting to see a comparison
show comments
api
Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.
Why would I want to use a model that has no knowledge of Tiananmen Square or Winnie The Pooh?
show comments
tracerbulletx
The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.
Thanks, trying it now. On my 3090 it seems to run at around 60tps, the IQ3_S variant. I am testing it now to see if it is better than Qwen 3.8 27b
I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I'm running 4-bit quants on an RTX Pro 6000 rented for approximately $1/hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: https://www.reddit.com/r/BlackwellPerformance/s/FrKwk3GoDK
I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.
To put that into perspective, here are some more numbers from other models via llama.cpp:
Median/Average
Qwen 3.5 9B BF16: 46.5 / 193.3
Qwen 3.6 35B Q4 K XL: 38.4 / 76.4
Qwen 3.5 122B Q3 K M: 32.9 / 68.6
The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.
I have done no further testing, as these results line up perfectly with my expectations.
I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.
Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:
Most important for me, I can run 4 concurrent streams at 400+ tok/s.https://github.com/fairfieldt/ds4
I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.
https://huggingface.co/Qwen/Qwen3.8-Flash-Next
LLM threads the world over are spammed with Strata links, it remains to be seen how much of the breathless hype remains standing once the honeymoon period is over. I've tried it but so far I have not seen anything that overly impressed me in terms of accuracy, though the speed is definitely there. I'm sure there are applications for LLMs where the quality of the answers is less important but I don't have any of those. YMMV.
Haven't had time to do much quality testing but the IQ2 model is running at 65 tps on a 9070xt/5900x. That is wild.
Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.
https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...
Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?
More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
We're getting closer and closer to the day we can have an Opus-like model running locally. The dream!
I don't get it. It's file size is about 6 times larger than 27B model for the same quant, but the performance improvement is hardly 10% across all benchmarks, according the metrics on it's hf page. Why should one devote so much more hardware for so little benefit?
Q2 quantization. Not interested.
All headlines about LLM performance MUST have the quantization also mentioned in the headline.
You know, so you're not wasting your time like in this post.
Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.
Surprisingly useful as long as you can leave it running a couple of hours at the very least.
While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.
I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).
There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
- Tiny context size or hours to load it
- Hard to benefit from thinking and preserve thinking given token cost.
- Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.
- K/V quants probably quantized too make things less accurate.
Useful would be combinations with:
- Full context size so it can code and think a bit.
- Draft MTP <= 2 so it doesn't trip
- Q4 quants or better so its accurate
- q8 cache or better so it stays accurate.
- 20 token/s so it finishes while reviewing previous step.
- 1000 tokens/s context load so compactions don't waste 10+ minutes.
- And enough left RAM for 50+ context checkpoints so that it can progress quuckly.
Closest you have is Qwen3.6-35B-A3B-MTP.
Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.
Source: I have low specs and tried them all for agentic use + coding.
I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.
I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?
I hope we're coming to a plateau with the HBM/GDDR7/on-chip RAM hype and get back to normalcy with system RAM alternatives for the rest of us.
My goto private setup, runs ~50t/sex on a dgx spark with sglang, nvfp4. Excellent model.
I wonder if a higher Qwen3.8-27b quant could beat or match these lower < 16/24/48/64G Qwen3.8-Flash Next quants given similar quality.
What speed are you willing the sacrifice to debug/program for more complex jobs faster?
Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
>> The model is a team of 24,576 small specialists ("experts")
That's a neat number (576 is the square of 24). Ofcourse it must have come from 24 * 2^10.
I had this working with the FreeToken inference engine a month ago when they launched.
https://github.com/FlashML-org/FreeToken
Anyone know how it compares to GLM 5.3 for real world use?
Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.
Progress on running local models has been amazing.
I think it would be great if you could try models with lesser parameters that could fit on 6GB VRAM-ish, which could work for "gaming laptops" as well.
Pretty impressive so far, but needs more testing.
It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.
Local LLM is getting more exciting every day!
I need a version of this that runs 3.8 27B on 8 gigs of VRAM
amazing project, congrats on the launch
I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?
The Readme doesn't say, but it's all AI generated, so..
I might try combining this a FreeToken
Has anyone calculated the effective intelligence of these quantized models?
I think publishing benchmarks with quantized models should become standard practice.
I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
getting it to fit is impressive, but i'd want to compare the smaller quants on a real coding task before picking one. how much quality do you lose going from IQ3_S to Q2_0?
80 t/s w/ 3090s and 3.05bpw exllamav3
lm studio bionic, unsloth, now this... it would be nice if it worked at least in one of those without installing another component
It hard to take below-8bit quants seriously
It's funny that most of the AI industry is built around the assumption (which is most likely true) that it is not possible to run SOTA models on current consumer hardware.
Imagine if someone managed to run an Astra- or Fable-level model on a 5090 at reasonable speeds.
Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
Lol, sure, if you quant it to hell (Q2) it'll go real fast...
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
There’s other slop projects to run of Qwen, like ds4, would be interesting to see a comparison
Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.
> Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
And I thought piping to bash was bad
Why would I want to use a model that has no knowledge of Tiananmen Square or Winnie The Pooh?
The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.