262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).
Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?
lxe
On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.
Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.
Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.
show comments
kmike84
This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)
I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.
3 main failure modes I observed in the engines:
* Not using best available spec decoding
* Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)
* Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline
show comments
herf
I have two NVIDIA GPUs (16GB+16GB) here, and it detects them each twice (says I have 4 GPUs). But then, it says most models are too big (anything >8GB?) and seems to run only on one GPU (5070ti).
Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:
set CUDA_VISIBLE_DEVICES=0
build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0
show comments
mncharity
Fwiw, top of my own pain-point list (I suppose given the first item, that's a pun) includes:
External/policy-based throttling for temperature control. Unthrottled, my laptop bottom goes skin-burn hot. But fixed compute caps can have non-linearly dreadful performance impacts in particular cases. Plan is a runtime knob, to replace manual limits-kludgery.
I'll use models which barely fit in VRAM+RAM, and are order-1 tok/s slow. So tool call step overhead can be painful - a world where `ls` costs tens of seconds. Plan is blending harness plugins with inference loop, for "no, don't stop - I already have the call result for you - just keep going" (and also some logit games).
show comments
c7b
Cool idea! Do you happen to have benchmarks for Strix Halo (AMD Ryzen AI Max+ 395)? I take it that Qwen3.8-Flash-Next is not supported?
And a more general question: does your engine detect and optimize for custom setups like multiple (possibly different) GPUs, eGPUs,...? Because if all you have is a stock major system like a Mac or DGX Spark, that's all you're going to care about, and there are a lot of highly optimized single-hardware engines out there that will be hard to beat in the long run. Something that automatically adapts to custom systems that don't have their own subreddits could really fill a gap.
show comments
nateb2022
Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
show comments
msdz
Congratulations on the launch, it looks like an impressive product and tool!
Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is that model- or at least architecture-specific code is required for most, if not each new open-weight model coming out.
Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?
if fully custom compiler would find best settings for given setup, upload the setup to mothership and allow new peers to download it as good starting point.
Looks interesting! Is there any way to skip or speed up the 'Assessing Models' step? I'm unable to download anything because it's been taking forever. I'm sure you could apply some quick heuristics or do a lookup or something to filter models. Or trust the user a bit more -- I already know which models fit on my machine. As it stands I'm stuck at that step and can't use the app.
Maybe have it run silently in the background and assess on demand when a user selects / attempts to download a model. It's not quite clear why all need to be assessed before I can download the first model to try.
show comments
hypercube33
From your description looks like this isn't for AMD or Strix Halo at all? Also one of the things I'm not sure of but definitely plays a huge factor is the variant of the model you download - how does this help select the fastest version for your specific hardware / context size?
show comments
kenzic
How long does tuning take (on an M3 MacBook Pro for example)?
show comments
amirhesham
Oh this is so cool. Curious about the business model, too.
show comments
paulgerhardt
Trying to run this but keep hitting bugs. Can you open up issues reporting on your repo?
show comments
sgtwompwomp
This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
show comments
p-e-w
What is the business model?
show comments
yolandac
does it allow us to run larger models that weren't possible before?
How accurate are speed estimates in the UI? I'm asking because for Qwen 3.8 (Q8) the speed numbers cited in the UI look quite poor:
262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?
On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.
Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.
Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.
This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)
I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.
3 main failure modes I observed in the engines:
* Not using best available spec decoding
* Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)
* Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline
I have two NVIDIA GPUs (16GB+16GB) here, and it detects them each twice (says I have 4 GPUs). But then, it says most models are too big (anything >8GB?) and seems to run only on one GPU (5070ti).
Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:
set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0
Fwiw, top of my own pain-point list (I suppose given the first item, that's a pun) includes:
External/policy-based throttling for temperature control. Unthrottled, my laptop bottom goes skin-burn hot. But fixed compute caps can have non-linearly dreadful performance impacts in particular cases. Plan is a runtime knob, to replace manual limits-kludgery.
I'll use models which barely fit in VRAM+RAM, and are order-1 tok/s slow. So tool call step overhead can be painful - a world where `ls` costs tens of seconds. Plan is blending harness plugins with inference loop, for "no, don't stop - I already have the call result for you - just keep going" (and also some logit games).
Cool idea! Do you happen to have benchmarks for Strix Halo (AMD Ryzen AI Max+ 395)? I take it that Qwen3.8-Flash-Next is not supported?
And a more general question: does your engine detect and optimize for custom setups like multiple (possibly different) GPUs, eGPUs,...? Because if all you have is a stock major system like a Mac or DGX Spark, that's all you're going to care about, and there are a lot of highly optimized single-hardware engines out there that will be hard to beat in the long run. Something that automatically adapts to custom systems that don't have their own subreddits could really fill a gap.
Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
Congratulations on the launch, it looks like an impressive product and tool!
Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is that model- or at least architecture-specific code is required for most, if not each new open-weight model coming out.
Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?
How does this compare to ZML's llmd https://zml.ai/llmd/ ?
if fully custom compiler would find best settings for given setup, upload the setup to mothership and allow new peers to download it as good starting point.
Can it do all the shenanigans that allows to run qwen flash on 12GB vram over 40 toks like people seems to be getting in this thread?: https://www.reddit.com/r/LocalLLaMA/comments/1wp7zyb/qwen38f...
Looks interesting! Is there any way to skip or speed up the 'Assessing Models' step? I'm unable to download anything because it's been taking forever. I'm sure you could apply some quick heuristics or do a lookup or something to filter models. Or trust the user a bit more -- I already know which models fit on my machine. As it stands I'm stuck at that step and can't use the app.
Maybe have it run silently in the background and assess on demand when a user selects / attempts to download a model. It's not quite clear why all need to be assessed before I can download the first model to try.
From your description looks like this isn't for AMD or Strix Halo at all? Also one of the things I'm not sure of but definitely plays a huge factor is the variant of the model you download - how does this help select the fastest version for your specific hardware / context size?
How long does tuning take (on an M3 MacBook Pro for example)?
Oh this is so cool. Curious about the business model, too.
Trying to run this but keep hitting bugs. Can you open up issues reporting on your repo?
This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
What is the business model?
does it allow us to run larger models that weren't possible before?