Going directly for the logprobs is always icky when you use a chat model as base, because they are trained to write prose as output. So your "choice" tokens and thus their probabilities might get diluted in whatever else it wanted to say. If you have to do it in the same way as this post, at least add clear system instructions and a carefully worded beginning to the assistant output section of the prompt to lower the chances of it wandering off immediately.
I've found that using structured outputs solves this problem much better. Instead of letting a model generate only "A", "B" or "C" and looking at the probs, have it directly generate "Legitimate", "Spam" or "Phishing" or any other pre-defined option from a set of multi-token sequences. Behind the scenes it boils down to something quite similar, but you're not running into the risk that the model actually wanted to say "A phishing attempt seems likely, so answer (C) is correct.", which would lead "A" to have the highest probability in the first token. You can even use a reasoning budget this way either via inherent reasoning or a free-form part preceding the remaining output structure. You can also have it assign probabilities (either in words or numbers) using more complex output structures, but I would not rely on them much more than the token logprobs (they can still be quite good though).
show comments
antirez
Because of masked attention in LLMs, if you put the options before the body (the email to analyze), the transformer already knows what it needs to look for, and can use more tokens to create state to address that specific task (BERT has no mask in the attention, so tokens attend also to next tokens). You could also do a few examples in the system prompt to improve calibration.
Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."
show comments
iamflimflam1
The number of - “I did/invented Jev last year”, or, “here’s a version of Jev I vibed up last night” is getting a bit ridiculous.
Especially ridiculous is how the hacker news crowd seems to be taking these at face value…
There was one the other day with a compelling demo. But when you looked closely at it, it was feeding in the options with the word “best” on the option to pick and a fine tuned model designed to recognise that word…
philipbk
> "25 lines of python"
> "import Solution"
ok
show comments
visarga
I also built one, but mine uses embeddings. It classifies concepts defined by a collection of positive and negative examples. The classifier model is trained in <1 second using ridge regression. The model itself is exactly the same shape as the embedding, so it works as a concept embedding. Since I already have a dataset, I can use it to do conformal prediction in order to calibrate confidence scores. Jev, on the other hand, has a generic model, not trained on in-domain examples, so its confidence scores are uncalibrated for any non-generic task.
So you might ask: how do I obtain the training examples? Just collect samples and use a coding agent to classify them as match or no match. From time to time you can add more examples to the dataset to have your concept adapt to changes in input distribution. It's all automated, but it only uses LLMs to train concept vectors, after that it works like a regular embedding model with a calibrated classifier on top. It's also 20-30x faster than Jev, free, and runs on CPU.
Beyond the missing latency and compute comparisons that Heaney commenter mentioned, also nothing about its error rate compared to Jev (nor if it even always outputs in a format the app can parse, not sure how solved that is).
But then at the end it says it’s parody. Maybe HN title should say it’s a joke.
show comments
SylonZero
Haha - I did enjoy this read! And there is a point to the whole marketing-dresses-up-stuff that is certainly true. I think it's worth pointing out the other HN story earlier https://news.ycombinator.com/item?id=49765348 about an open-weight model called Laya.
P.S. I am evaluating that model for a production use case where I would have used Jev
0123456789ABCDE
this is such a trivial thing to do in DSPY, no one bothered to give it a name…
here's 7 lines
import os
import dspy
lm = dspy.LM("openrouter/z-ai/glm-5.3-flash", api_key=os.environ["OPENROUTER_API_KEY"])
jev = dspy.Predict('email:str -> choice:Literal["Legitimate", "Spam", "Phishing"]')
email = "Payroll asks for your password on a non-company sign-in page."
pred = jev(email=email, lm=lm)
print(pred.choice)
there are other options, obviously. you can choose to give it some tools, maybe some reasoning stage before picking a choice, and that's on top of the "reasoning" the llm model already does api side
show comments
rcarmo
Very nice as a conceptual thing, but there's a bit more to it. I've bolted Gemma 4 onto a custom pipeline for that (https://rcarmo.github.io/projects/go-system-one/) and it's OK-ish (a bit slow on my puny 3060, but I can use it to prototype a bunch of things locally until the Jev mania settles and we have better models).
yipinwong
Any analogy works at certain abstraction level, and this works with a premise that it's a classifier that makes decisions.
Still good. In practice for Jev the devils in the details.
As you all know by now, it's easy to write PoC and understand with AIs (or even manually, which is now a prestious practice).
That demo will get you 80% there
Getting to that 100% or even 99% to JEV level will be hard with all the edge cases, infra, API, communications, etc.
Still a good article.
zeroq
How to write Jev in 25 lines of Python:
1. draw a circle
2. import the rest of the owl
xigoi
This is like saying that cars are useless because you can achieve the same thing by removing the cannon from a tank.
show comments
jorisw
Highly suspect of content marketing.
Ends with referring to a product, and saying "this is a parody post", after pretending to make a serious point.
show comments
alun
The one thing I can't wrap my head around with Jev is why they're trying to create that "System One" narrative.
In real life, a human doesn't do classification tasks with the System One part of their brain, they use System Two. So by definition what Jev does isn't System One thinking.
If anything, regular programming that automatically executes based on logic, without requiring "thinking" would be "System One".
show comments
beamy
> We didn't train a model with Reinforcement Learning for Calibrated Decisions (RLCD) to calibrate the decisions and probabilities
But aren’t calibrated predictions one of the defining features?
onion2k
It's fast.
If you're comparing with something, you need to state 'fast' in relative terms. Jev is definitely fast, and if this Python takes the same time to get a decision then it's also fast. If it's 100* slower than Jev though, you shouldn't be calling it 'fast', because relatively speaking it's really, really slow.
show comments
vonStackelberg
Hey, I enjoyed the article! I think it’s helpful to break stuff down as much as possible to nail down what is happening. They did that. The important part is not importing a model or something, it’s what’s happening after. I learned something
bruhhhhhh
I am hearing about Jev for the first time here so no idea about the hype.
So their(Jev) is that the thing is faster at classification than a frontier model? Because the whole type safe aspect is already fully solvable with structured output.
But their example is classification but that would also be possible and faster with a classic BERT model.
So their pitch is a task specific smaller model or am I completely misunderstanding the whole thing?
show comments
dhsysusbsjsi
Whilst I do like reading these things for technical know how, I can sympathise with the creator of jev who now presumably has to apply an order of magnitude effort to explain why the 100 smaller things done better than this add up to a much better product.
show comments
tducret
The website is currently returning "Site not available"
First release people are getting more balanced at least. Less 'it's gonna change everything' to 'might be viable use cases'. If I have to look at one more video high lighting google flights, I might loose it though.
armcat
Looking at the logprobs on tokens works for the local models, but not on the frontier ones. It's been more or less broken since GPT-4o for example. I wrote about it two years ago: https://medium.com/data-science/9-11-or-9-9-which-one-is-hig.... Also, I've done some work in estimating confidence and on rubric evals using the same method, and you actually get better correlation to "real confidence" by just getting the LLM to say it.
show comments
cupofjoakim
I wonder if this could be a good stepping stone to write a local prompt router to optimise what model get what prompt. I.e. if the prompt is just a lookup, send it to haiku, if it's reasoning, send it to opus and if it's implementation send it to sonnet.
show comments
chpatrick
Could someone explain how Jev is different from using any old model and constraining the output to "My choice is a/b/c..."?
show comments
davidfekke
This is a system two model, and not a system one. To get the performance of Jev, and you want to run locally, use Laya. It is up on Hugging Face.
kjshsh123
Maybe someone can explain why RL is even needed for post training with Jev? We have supervised labels.
I guess it's due to the calibrated decision part (and that's what LLMs tell me).
But I figure some supervised classification post training would still improve the model.
fzysingularity
Am I missing something here:
p(y = next thinking+decision token | x = question) != p(y = next decision token | x = question)
The former is what LLMs are trained for, the latter is what Jev was likely trained on (likely used thinking alignment as an auxiliary loss, but not explicitly included in the probability calibration).
boros2me
We have Jev at home
rgbrgb
i love that people are trying to make OS jevs but what is the point of doing all this work and not ask your coding agent to do a little benchmarking. selfishly want an open weight model to beat jev here
kinda feel like the
"this is a parody blog post, see these links for better/more complete open implementations of Jev..."
is a legal cya a la "Nathan For You" 's Dumb Starbucks
xg15
I missed the hypewave so can't say a lot about Jev, but the double standards are entertaining:
About Jev:
> We didn't train a model with Reinforcement Learning for Calibrated Decisions (RLCD) to calibrate the decisions and probabilities (even though they are not always correct).
Only 99% correctness! Borderline unusable!
About their model:
> It classifies: it gets a prompt with choices and outputs probabilities.
You want numbers, it gives you numbers! What more could you want?
pjankiewicz
What I'm missing here is also type guarantees. I don't think you can do it without token level logic which forces the model to output the tokens from a predefined pool of tokens. A logic like this given some JSON schema is not that difficult to implement. If the LLM must output JSON schema compatible value then you can also add that it doesn't "hallucinate". Which is funny too because just guaranteeing the type does not mean the model does not hallucinate but this is another story.
K0IN
a hile ago (when big providers still provided logprobs) i created a VS Code highlighter that visualizes unsure tokens.
Since most chat models want to answer with a human-readable message i think their logprobs are not as meaningful. It would be interesting to see if one choice is like "correct" and if the model wants to choose it more often, cause it might not answer the question but to prose to the user.
petercooper
You can also go beyond Jev. Qwen 3.5 0.8B is fantastic at basic image classification/question answering (including OCR elements) also. Though rather than looking at logits, I get it to output a structured JSON object and it does simple object classification tasks on a Mac at under 500ms a pop (I forget how far, but I think it's like ~250ms) with good accuracy (depending on task).
What I don’t understand is, why would you not want “reasoning” in a classifier?
Speed and cost are obvious reasons, but isn’t this a tradeoff?
show comments
teravor
to be fair a Jev architecture would be better optimized for this particular workflow than an LLM.
<think>\n\n</think>
but letting an LLM think would trade latency and performance for significant reliability above that of Jev.
qurren
> Email: {email}\n\n{options}<|im_end|>
Why do I have to feed my e-mail into the model?
show comments
nlpnerd
This is basically Temu Jev
param_gupta
Pretty interesting how a simple example like this makes the idea so easy to understand.
tracyhenry
This, like the hype of Jev on Twitter, totally ignores accuracy and generality across domains.
In my experience even structured LLM output performs poorly on classifier tasks. LLMs are trained to talk and think longer. If you don't give LLM enough space to reason it would become very dumb.
I'm not saying that Jev is way better, but that people way overindexed cost and speed.
philipbk
> "25 lines of python"
> "import Solution"
ok
imranq
This guy just seems a bit salty
max979
This makes me wonder about using Jevko for configuration instead of YAML, especially with such a compact parser.
marcy_74
Reminds me of my own tiny Lisp interpreter attempts; that moment when it first evaluates a simple expression is pure magic.
shawabawa3
strong "You can build dropbox quite trivially by getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" vibes
You have built something like jev but not jev (for starters, the output of what you've built will be absolutely worthless, the whole reason Jev is getting so much hype is because the output is good enough)
show comments
qainsights
THis is not Jev.
heaney-555
Latency and compute comparison needed.
show comments
revexos
Startup coming out of 2 years of stealth to be reproduced this easily
show comments
pietz
I'm surprised something like Jev came out "so late", but the hype has been ridiculous. Yes, it's a good idea. No, it only helps when fast and cheap are important and I guarantee existing labs will have this figured out in a matter of days.
Add visual understanding, add reasoning and bring down the size to run on my computer. That's when it will be interesting.
So many people that don't understand the tech jumped on the hype train because "it cannot hallucinate" and else. It's crazy.
show comments
ricardobeat
Now, can you do it in <200ms for 45 questions at once, have 0% malformed output, and any kind of meaningful benchmark? We’ll wait!
show comments
teaonly
The principle is this.
esafak
Latency-calibration charts or it didn't happen. (LLMs are not optimized for calibration.)
zteppenwolf
I get dishonest vibes from this post? Jev claims to be cheaper/more efficient, and the post claims just to achieve the same functionality.
show comments
iLoveOncall
Nothing I hate more than bullshit articles claiming X in Y lines of code, only to use libraries abstracting hundreds of thousands of lines of code.
show comments
baobabKoodaa
I'm so sick of seeing these people who "made Jev in 25 lines of Python" or whatever the flavor of the day is. Do you people seriously think that Qwen3-0.6B-Q8_0.gguf is frontier intelligence? If you want to argue that Jev is NOT frontier intelligence, then go make that argument. Don't try to pretend that Qwen3-0.6B-Q8_0.gguf is frontier intelligence. That's retarded.
Going directly for the logprobs is always icky when you use a chat model as base, because they are trained to write prose as output. So your "choice" tokens and thus their probabilities might get diluted in whatever else it wanted to say. If you have to do it in the same way as this post, at least add clear system instructions and a carefully worded beginning to the assistant output section of the prompt to lower the chances of it wandering off immediately.
I've found that using structured outputs solves this problem much better. Instead of letting a model generate only "A", "B" or "C" and looking at the probs, have it directly generate "Legitimate", "Spam" or "Phishing" or any other pre-defined option from a set of multi-token sequences. Behind the scenes it boils down to something quite similar, but you're not running into the risk that the model actually wanted to say "A phishing attempt seems likely, so answer (C) is correct.", which would lead "A" to have the highest probability in the first token. You can even use a reasoning budget this way either via inherent reasoning or a free-form part preceding the remaining output structure. You can also have it assign probabilities (either in words or numbers) using more complex output structures, but I would not rely on them much more than the token logprobs (they can still be quite good though).
Because of masked attention in LLMs, if you put the options before the body (the email to analyze), the transformer already knows what it needs to look for, and can use more tokens to create state to address that specific task (BERT has no mask in the attention, so tokens attend also to next tokens). You could also do a few examples in the system prompt to improve calibration.
Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."
The number of - “I did/invented Jev last year”, or, “here’s a version of Jev I vibed up last night” is getting a bit ridiculous.
Especially ridiculous is how the hacker news crowd seems to be taking these at face value…
There was one the other day with a compelling demo. But when you looked closely at it, it was feeding in the options with the word “best” on the option to pick and a fine tuned model designed to recognise that word…
> "25 lines of python" > "import Solution" ok
I also built one, but mine uses embeddings. It classifies concepts defined by a collection of positive and negative examples. The classifier model is trained in <1 second using ridge regression. The model itself is exactly the same shape as the embedding, so it works as a concept embedding. Since I already have a dataset, I can use it to do conformal prediction in order to calibrate confidence scores. Jev, on the other hand, has a generic model, not trained on in-domain examples, so its confidence scores are uncalibrated for any non-generic task.
So you might ask: how do I obtain the training examples? Just collect samples and use a coding agent to classify them as match or no match. From time to time you can add more examples to the dataset to have your concept adapt to changes in input distribution. It's all automated, but it only uses LLMs to train concept vectors, after that it works like a regular embedding model with a calibrated classifier on top. It's also 20-30x faster than Jev, free, and runs on CPU.
An illustration of how it defines a concept as opposed to simple cosine similarity: https://github.com/horiacristescu/semlabel/raw/main/images/c...
Beyond the missing latency and compute comparisons that Heaney commenter mentioned, also nothing about its error rate compared to Jev (nor if it even always outputs in a format the app can parse, not sure how solved that is).
But then at the end it says it’s parody. Maybe HN title should say it’s a joke.
Haha - I did enjoy this read! And there is a point to the whole marketing-dresses-up-stuff that is certainly true. I think it's worth pointing out the other HN story earlier https://news.ycombinator.com/item?id=49765348 about an open-weight model called Laya.
P.S. I am evaluating that model for a production use case where I would have used Jev
this is such a trivial thing to do in DSPY, no one bothered to give it a name…
here's 7 lines
there are other options, obviously. you can choose to give it some tools, maybe some reasoning stage before picking a choice, and that's on top of the "reasoning" the llm model already does api sideVery nice as a conceptual thing, but there's a bit more to it. I've bolted Gemma 4 onto a custom pipeline for that (https://rcarmo.github.io/projects/go-system-one/) and it's OK-ish (a bit slow on my puny 3060, but I can use it to prototype a bunch of things locally until the Jev mania settles and we have better models).
Any analogy works at certain abstraction level, and this works with a premise that it's a classifier that makes decisions.
Still good. In practice for Jev the devils in the details. As you all know by now, it's easy to write PoC and understand with AIs (or even manually, which is now a prestious practice).
That demo will get you 80% there
Getting to that 100% or even 99% to JEV level will be hard with all the edge cases, infra, API, communications, etc.
Still a good article.
How to write Jev in 25 lines of Python:
This is like saying that cars are useless because you can achieve the same thing by removing the cannon from a tank.
Highly suspect of content marketing.
Ends with referring to a product, and saying "this is a parody post", after pretending to make a serious point.
The one thing I can't wrap my head around with Jev is why they're trying to create that "System One" narrative.
In real life, a human doesn't do classification tasks with the System One part of their brain, they use System Two. So by definition what Jev does isn't System One thinking.
If anything, regular programming that automatically executes based on logic, without requiring "thinking" would be "System One".
> We didn't train a model with Reinforcement Learning for Calibrated Decisions (RLCD) to calibrate the decisions and probabilities
But aren’t calibrated predictions one of the defining features?
It's fast.
If you're comparing with something, you need to state 'fast' in relative terms. Jev is definitely fast, and if this Python takes the same time to get a decision then it's also fast. If it's 100* slower than Jev though, you shouldn't be calling it 'fast', because relatively speaking it's really, really slow.
Hey, I enjoyed the article! I think it’s helpful to break stuff down as much as possible to nail down what is happening. They did that. The important part is not importing a model or something, it’s what’s happening after. I learned something
I am hearing about Jev for the first time here so no idea about the hype. So their(Jev) is that the thing is faster at classification than a frontier model? Because the whole type safe aspect is already fully solvable with structured output. But their example is classification but that would also be possible and faster with a classic BERT model. So their pitch is a task specific smaller model or am I completely misunderstanding the whole thing?
Whilst I do like reading these things for technical know how, I can sympathise with the creator of jev who now presumably has to apply an order of magnitude effort to explain why the 100 smaller things done better than this add up to a much better product.
The website is currently returning "Site not available"
Here is an archive: https://web.archive.org/web/20260923122959/https://www.nobod...
First release people are getting more balanced at least. Less 'it's gonna change everything' to 'might be viable use cases'. If I have to look at one more video high lighting google flights, I might loose it though.
Looking at the logprobs on tokens works for the local models, but not on the frontier ones. It's been more or less broken since GPT-4o for example. I wrote about it two years ago: https://medium.com/data-science/9-11-or-9-9-which-one-is-hig.... Also, I've done some work in estimating confidence and on rubric evals using the same method, and you actually get better correlation to "real confidence" by just getting the LLM to say it.
I wonder if this could be a good stepping stone to write a local prompt router to optimise what model get what prompt. I.e. if the prompt is just a lookup, send it to haiku, if it's reasoning, send it to opus and if it's implementation send it to sonnet.
Could someone explain how Jev is different from using any old model and constraining the output to "My choice is a/b/c..."?
This is a system two model, and not a system one. To get the performance of Jev, and you want to run locally, use Laya. It is up on Hugging Face.
Maybe someone can explain why RL is even needed for post training with Jev? We have supervised labels.
I guess it's due to the calibrated decision part (and that's what LLMs tell me).
But I figure some supervised classification post training would still improve the model.
Am I missing something here:
p(y = next thinking+decision token | x = question) != p(y = next decision token | x = question)
The former is what LLMs are trained for, the latter is what Jev was likely trained on (likely used thinking alignment as an auxiliary loss, but not explicitly included in the probability calibration).
We have Jev at home
i love that people are trying to make OS jevs but what is the point of doing all this work and not ask your coding agent to do a little benchmarking. selfishly want an open weight model to beat jev here
just found this one https://huggingface.co/spaces/multimodalart/jev-decision-ind...
kinda feel like the "this is a parody blog post, see these links for better/more complete open implementations of Jev..."
is a legal cya a la "Nathan For You" 's Dumb Starbucks
I missed the hypewave so can't say a lot about Jev, but the double standards are entertaining:
About Jev:
> We didn't train a model with Reinforcement Learning for Calibrated Decisions (RLCD) to calibrate the decisions and probabilities (even though they are not always correct).
Only 99% correctness! Borderline unusable!
About their model:
> It classifies: it gets a prompt with choices and outputs probabilities.
You want numbers, it gives you numbers! What more could you want?
What I'm missing here is also type guarantees. I don't think you can do it without token level logic which forces the model to output the tokens from a predefined pool of tokens. A logic like this given some JSON schema is not that difficult to implement. If the LLM must output JSON schema compatible value then you can also add that it doesn't "hallucinate". Which is funny too because just guaranteeing the type does not mean the model does not hallucinate but this is another story.
a hile ago (when big providers still provided logprobs) i created a VS Code highlighter that visualizes unsure tokens.
Since most chat models want to answer with a human-readable message i think their logprobs are not as meaningful. It would be interesting to see if one choice is like "correct" and if the model wants to choose it more often, cause it might not answer the question but to prose to the user.
You can also go beyond Jev. Qwen 3.5 0.8B is fantastic at basic image classification/question answering (including OCR elements) also. Though rather than looking at logits, I get it to output a structured JSON object and it does simple object classification tasks on a Mac at under 500ms a pop (I forget how far, but I think it's like ~250ms) with good accuracy (depending on task).
I've built something similar, i am hosting it.
You can test Jev like model at 26B parameter count here (built few weeks ago): https://gambler-relay-us-west1.leo-fish.ts.net/demo (might not stay up for long)
Typesafe compatible API
This is just running on old hardware.
What I don’t understand is, why would you not want “reasoning” in a classifier?
Speed and cost are obvious reasons, but isn’t this a tradeoff?
to be fair a Jev architecture would be better optimized for this particular workflow than an LLM.
but letting an LLM think would trade latency and performance for significant reliability above that of Jev.> Email: {email}\n\n{options}<|im_end|>
Why do I have to feed my e-mail into the model?
This is basically Temu Jev
Pretty interesting how a simple example like this makes the idea so easy to understand.
This, like the hype of Jev on Twitter, totally ignores accuracy and generality across domains.
In my experience even structured LLM output performs poorly on classifier tasks. LLMs are trained to talk and think longer. If you don't give LLM enough space to reason it would become very dumb.
I'm not saying that Jev is way better, but that people way overindexed cost and speed.
> "25 lines of python" > "import Solution"
ok
This guy just seems a bit salty
This makes me wonder about using Jevko for configuration instead of YAML, especially with such a compact parser.
Reminds me of my own tiny Lisp interpreter attempts; that moment when it first evaluates a simple expression is pure magic.
strong "You can build dropbox quite trivially by getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" vibes
You have built something like jev but not jev (for starters, the output of what you've built will be absolutely worthless, the whole reason Jev is getting so much hype is because the output is good enough)
THis is not Jev.
Latency and compute comparison needed.
Startup coming out of 2 years of stealth to be reproduced this easily
I'm surprised something like Jev came out "so late", but the hype has been ridiculous. Yes, it's a good idea. No, it only helps when fast and cheap are important and I guarantee existing labs will have this figured out in a matter of days.
Add visual understanding, add reasoning and bring down the size to run on my computer. That's when it will be interesting.
So many people that don't understand the tech jumped on the hype train because "it cannot hallucinate" and else. It's crazy.
Now, can you do it in <200ms for 45 questions at once, have 0% malformed output, and any kind of meaningful benchmark? We’ll wait!
The principle is this.
Latency-calibration charts or it didn't happen. (LLMs are not optimized for calibration.)
I get dishonest vibes from this post? Jev claims to be cheaper/more efficient, and the post claims just to achieve the same functionality.
Nothing I hate more than bullshit articles claiming X in Y lines of code, only to use libraries abstracting hundreds of thousands of lines of code.
I'm so sick of seeing these people who "made Jev in 25 lines of Python" or whatever the flavor of the day is. Do you people seriously think that Qwen3-0.6B-Q8_0.gguf is frontier intelligence? If you want to argue that Jev is NOT frontier intelligence, then go make that argument. Don't try to pretend that Qwen3-0.6B-Q8_0.gguf is frontier intelligence. That's retarded.