What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
show comments
TN1ck
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.
It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.
Ask Jeeves - Only took us 30 years to come full circle.
show comments
itzikkatz
Cool engineering, but 17s p90 latency kind of defeats the point of a Jev-class model, which is supposed to be fast and cheap. Losing 10 points on MMLU along the way doesn't help.
trencedamp
Jev noob here. I'm seeing all this jev talk and I understand the difference between this and normal models, but what are some actual use cases for jev?
theanonymousone
This reminds me of "on-premise cloud".
teravor
you don't need to post-train anything for this.
just get an LLM to think and then force it to output a specific json with prefill post-think.
make sure to include good conditioning text in the prompt with examples of exactly what the output should be like. you don't want dissonance in the probabilities on the prefill.
show comments
alienbaby
Just curious, where has this term 'noul' come from for yes/no ansers?
/a bit more digging and..
A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.
I hate it :)
show comments
winddude
That completely defeats the point of something to make decisions faster.
swader999
Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
zerop
Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
show comments
swingboy
Any good classifiers like this or Jev that support image input?
show comments
HarHarVeryFunny
The number of people who feel the need to try and argue that you don't need Jev can only be astroturfing by those with something to lose - Anthropic and OpenAI employees.
Like it or not, companies are going to use Jev unless you can offer something just as cheap and fast.
I wonder just how much of the business automation market, previously held by LLMs, is at risk here?
druskacik
How's the performance compared to ordinary 9B LLM with structured outputs? Both accuracy and speed?
RamblingCTO
Super dope. If it would ship as prod ready code supporting mps as well that would be even doper.
But funny that jev is getting its lunch eaten apparently in under two weeks?
show comments
woadwarrior01
This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
loclol101
How general really are these jev type models? Has anyone done any broad very cross-domain eval on them?
Naitik88
what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
AnodicElegy
I'm surprised we haven't seen a "Jehovah" yet.
show comments
mxkuzn
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
quantized_state
The diffusion drafter adaptation is nice
captainbland
See if it can beat Jev's Pokémon benchmark
show comments
raverbashing
Jeeves, that's a name I haven't heard in a long time...
show comments
singularity2001
In my experience, Jev is only faster because it's a small shitty model. Any objections?
speedping
Meh. Wake me up when it reasons in latent space and answers in less than a second
esafak
Jev-like models give calibrated decision probabilities, but at low accuracy.
So why didn't they show both??
phplovesong
So "askjeeves" has been resurrected?
SV_BubbleTime
“Somehow Jeeves returned…”
hjun1052
If the model does autoregressive reasoning before the decision, doesn't that give up much of what a Jev-style model buys you (a single forward pass, cheap calibrated probabilities)? Or is the point mainly to keep the typed output and probability interface while getting better accuracy on harder cases?
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.
It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.
[1] https://tn1ck.com/blog/jevdit
Ask Jeeves - Only took us 30 years to come full circle.
Cool engineering, but 17s p90 latency kind of defeats the point of a Jev-class model, which is supposed to be fast and cheap. Losing 10 points on MMLU along the way doesn't help.
Jev noob here. I'm seeing all this jev talk and I understand the difference between this and normal models, but what are some actual use cases for jev?
This reminds me of "on-premise cloud".
you don't need to post-train anything for this.
just get an LLM to think and then force it to output a specific json with prefill post-think.
make sure to include good conditioning text in the prompt with examples of exactly what the output should be like. you don't want dissonance in the probabilities on the prefill.
Just curious, where has this term 'noul' come from for yes/no ansers?
/a bit more digging and..
A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.
I hate it :)
That completely defeats the point of something to make decisions faster.
Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
Any good classifiers like this or Jev that support image input?
The number of people who feel the need to try and argue that you don't need Jev can only be astroturfing by those with something to lose - Anthropic and OpenAI employees.
Like it or not, companies are going to use Jev unless you can offer something just as cheap and fast.
I wonder just how much of the business automation market, previously held by LLMs, is at risk here?
How's the performance compared to ordinary 9B LLM with structured outputs? Both accuracy and speed?
Super dope. If it would ship as prod ready code supporting mps as well that would be even doper.
But funny that jev is getting its lunch eaten apparently in under two weeks?
This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
How general really are these jev type models? Has anyone done any broad very cross-domain eval on them?
what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
I'm surprised we haven't seen a "Jehovah" yet.
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
The diffusion drafter adaptation is nice
See if it can beat Jev's Pokémon benchmark
Jeeves, that's a name I haven't heard in a long time...
In my experience, Jev is only faster because it's a small shitty model. Any objections?
Meh. Wake me up when it reasons in latent space and answers in less than a second
Jev-like models give calibrated decision probabilities, but at low accuracy.
So why didn't they show both??
So "askjeeves" has been resurrected?
“Somehow Jeeves returned…”
If the model does autoregressive reasoning before the decision, doesn't that give up much of what a Jev-style model buys you (a single forward pass, cheap calibrated probabilities)? Or is the point mainly to keep the typed output and probability interface while getting better accuracy on harder cases?