I'm a bit confused by the calibration penalty. Isn't a cross entropy loss already self-calibrating? Of course that depends on the data distribution you train on, but assuming it's representative, why do we need a second penalty?
Yeah, in theory that’s true. Minimizing NLL/Cross-entropy (in expectation) already encourages that. So an additional penalty isn’t strictly or theoretically necessary if we reach that optimum.
Guo et al. (2017) had a paper (https://proceedings.mlr.press/v70/guo17a.html) that discusses this but they observe that “neural networks can overfit to NLL without overfitting to the 0/1 loss.”
So they found test accuracy could improve while test cross-entropy worsened (due to overfitting on the training set.). So calibration can still be useful (but needs to be checked
empirically on held-out data).
Very good comment though and I should add a few paragraphs on this.
Fun footnote for the bag-of-words section. M. E. Maron was doing naive Bayes-style text classification back in 1961, sorting computer abstracts into 32 subject categories by their clue words.
The calibration discussion is a useful reminder that a higher accuracy model can still be the worse component when downstream decisions depend on confidence. I’d be curious whether you measured calibration drift separately across domains, since a single global temperature often hides class- or slice-specific failure modes.
Fair point, I only showed the averages here. Overall, calibration was beneficial, but it slightly worsened the ECE on one dataset (AGNews). So, yes, your point is right. (I will add a note about that.)
Thanks for checking. My guess is the temperature fit on the held-out split didn't match AGNews' test confidence profile, so a per-dataset or even per-class temperature might close that gap cheaply. Appreciate you adding the note.
I think there is a small typo in section 5.2 where the modified reward is defined. The line after the equation reads "where is 1 for a correct answer and 0 otherwise" instead of "where C is 1 for a correct answer and 0 otherwise".
A dedicated classifier probably still beats Jev on any single narrow job, as you note, but each one needs its own training pipeline, its own drift monitoring, its own retraining schedule. Run more than a handful of those tasks in parallel and I'd expect the maintenance surface, not the accuracy gap, to decide which approach an organization keeps running.
Jev collapses that into one thing to monitor instead of a dozen. The crossover then sits where the number of narrow classifiers a team maintains exceeds what that team can keep watched, rather than where Jev's accuracy catches up to a tuned task-specific model. Curious whether that tipping point is already showing up in how teams choose between the two approaches.
Yes, I agree with everything here, and that was the main point regarding Jev being so useful. Like you said, there is a maintenance burden for specialized models. Also, even if you develop and maintain all those, it's also going to be resource intensive to have like 500 different specialists running on your hardware (sure, one can have one main model and handle the rest with LoRA adapters, but still, it'd be lots of work)
I'm suspicious of the Tetris or other game-playing tests, because one of Jev's launch demos was playing a game, so this should be very much in-distribution where something like ModernBERT or GLiNER based models may have never trained on that type of decision.
That said I'm stoked to avoid fine-tuning for many low-volume decision tasks that were previously untenable due to time investment, and the interest plus signals like your late-breaking edit about OpenAI, and the Decisions output modality on OpenRouter show it's clearly a huge demand.
> ModernBERT or GLiNER based models may have never trained on that type of decision.
I agree. But that probably goes also for a lot of things. I do believe the reason why Jev is so good is that it's training distribution just covers such a wide range of tasks.
I am really curious if the next generation of Jev-like models could actually beat, say models that Tesla uses in production for full-self driving.
To me, it seems like we’re really narrowing down the path of training specialized models for tasks. You can really start feeling the generality of the intelligence, when you throw a task at hand and it just solves it. I see also similar things in robotics, e.g. growbot, just connect you phone to control ‘a body’ whatever it’s, it figures things out.
All of this reminds me the patato head theory of David Eagleman, that says brain is the ultimate decoder and whatever input you give it to it, it will find a way to use that input. Apparently, it is not just the brain that can do that, so we should really start accepting we’ve not only passed the turing test few years ago, we’ve almost achieved general intelligence.
Interesting point, but I remain a bit skeptical. A Jev-like model fully specialized on full-self-driving is still going to be better than a Jev-like generalist imho. Or, at least, at the same capability-level, a specialized model could potentially be smaller and faster.
Great article as always, articulating theory visually for mere mortals. One question about bag-of-words representation, which I struggle a lot to understand. You mention
"""
A bag-of-words representation results in these fixed-size inputs by assigning each word in a vocabulary its own position in a vector. We then count how often each word occurs in a document. For example, if we have a vocabulary of 50,000 unique words, it produces a fixed-size vector with 50,000 entries, regardless of whether the input consists of only ten words or 300k words
"""
How is it possible to know in advance the size of your vocabulary (unique words) considering that new documents, and hence new words, will come up in the future? Are there any predefined vocabularies provided by libraries which are subsequently used by classic classifiers? What is the standard (old) practice in the industry to address that?
Good question. It's based on the combined training set. So if there is a new word, it would simply not be counted, or you can have an "unknown" word placeholder (but there is probably no benefit in that). The CountVectorizer and TfidfVectorizer in scikit-learn, which I usually use for this, ignore those.
There is an open-source alternative to Jev, called Contrastive Language Model, that was released recently: https://contrastive-lm.notion.site/. It is not another Alpaca :) I am curious about your thoughts on it.
And yeah, examples like that are great (this one has the best docs I've seen so far), but my point stands: they are not nearly on the same level as Jev. For example, I just tried it and its IMDb Test accuracy is at just 82.90% (over 96% in Jev) and fails the Tetris test: https://sebastianraschka.com/videos/tetris-clm
Agreed. However, CLM was tested on other games (Dino, Super Mario and Wikiracing) and found to be on par with Jev. It could be that it was not trained on Tetris. As you pointed out, Typesafe probably spent a lot more time and money on building Jev than the CLM researchers. It is still encouraging to see there is a decent open-source alternative out there that allows us to fine-tune specific domains, like Tetris, but for business use cases.
A standing “insufficient evidence / none of the above” option could stop the system from being forced to manufacture certainty when every substantive choice is poorly supported.
Was Jev trained to handle such an option? And if not, would it meaningfully populate an “I don’t know” choice if one were always provided?
Yeah, sth like that would be useful. Right now, one would have to go by confidence. Like if the chosen answer has a low confidence, you can have a manual check that then escalates this to a GPT-6 Astra-level model or so.
This is a proprietary model so we can't see how they implemented it. It's possible they could just run a small model like qwen3-4B (or smaller, plus quantization) with the right system prompt and get this output, especially if you strip off the output layer and replace it with this output format.
Also, if it's predicting numbers and it's an LLM, it'd be like giving you a spreadsheet to look at and answer questions by reading it and doing math in your head. It's totally possible, but it's not optimal, and won't be accurate if dump a non-trivial amount of data in.
None of this is assertion of what they're doing since we can't know. It's just considering the easiest way they could have done it, and what the limitations would be if you expect an LLM to do things like forecasting or number crunching. They could also hook it up to a tool with a side model like tspulse or timefm if someone sent in a giant bunch of numbers (hint, model provider!).
This is entirely possible. But I do expect that they put a bit more effort behind it and curated a pretty good dataset for further training/fine-tuning. (If I were to build something like this, I would probably start with a pre-trained base model and focus on post-training mostly.)
Correct me if I am wrong but you cannot upload files to Jev in order to give specific domain knowledge plus it does not have memory so incremental learning on a specific domain via gradual feed is also not possible. What would be the best approach to creating a domain specific agent using your own information source. I have tried locally trained agents but they are so incredibly slow on a decent laptop
Yeah, I think Jev is not multimodal (yet). Regarding memory: that's true for any LLM though; memory usually comes from external context files via the harness. You can just build your own memory this way.
Thanks for the comprehensive article. One thing I wanted to mention w.r.t. Section 5.1 and 5.3 is that there's been a lot of work on calibrating neural networks in recent years, and many of these works are a substantive departure from the temperature-scaling (and related) approaches from the 2010's, in particular with regard to robustness to covariate shifts, which is important in the high-dimensional neural-network setting. (Single-parameter temperature scaling tends to be very brittle in the presence of covariate shifts.) Achieving that robustness is arguably something fundamentally new in terms of capability.
To recap, for context on calibration: The key properties for using such classification models for conditional-branching decisions in agentic stacks (and related) is that they should be well-calibrated (under the definition chosen for the task) and informative/"sharp" (e.g., always predicting the mean might be "well-calibrated" in a theoretical sense for some chosen quantities of interest, but isn't particularly useful in practice).
The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics. http://arxiv.org/abs/2509.12760
[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269. https://doi.org/10.1145/3786335.3813214
The LSTM example caught me: a model that can track word order still did worse here than the simpler baseline. Keeping the training-from-scratch detail beside it matters. It stops that result from becoming another tidy rule about which kind of model always wins.
I still have a hard time understanding how it can do any task out of the box. Is it an encoder then it needs to be tuned, if a decoder then it should be a bit slower, or i am missing something ?
I'm a bit confused by the calibration penalty. Isn't a cross entropy loss already self-calibrating? Of course that depends on the data distribution you train on, but assuming it's representative, why do we need a second penalty?
Yeah, in theory that’s true. Minimizing NLL/Cross-entropy (in expectation) already encourages that. So an additional penalty isn’t strictly or theoretically necessary if we reach that optimum.
Guo et al. (2017) had a paper (https://proceedings.mlr.press/v70/guo17a.html) that discusses this but they observe that “neural networks can overfit to NLL without overfitting to the 0/1 loss.”
So they found test accuracy could improve while test cross-entropy worsened (due to overfitting on the training set.). So calibration can still be useful (but needs to be checked
empirically on held-out data).
Very good comment though and I should add a few paragraphs on this.
Added this info in Section 5.3
thanks
Fun footnote for the bag-of-words section. M. E. Maron was doing naive Bayes-style text classification back in 1961, sorting computer abstracts into 32 subject categories by their clue words.
Nice one. I added it as a reference and linked the corresponding "Automatic Indexing: An Experimental Inquiry" paper (https://dl.acm.org/doi/10.1145/321075.321084)
Skimmed through the article, it’s a master class. I can’t wait to read it throughly. As always, thanks professor.
Thanks!!
The calibration discussion is a useful reminder that a higher accuracy model can still be the worse component when downstream decisions depend on confidence. I’d be curious whether you measured calibration drift separately across domains, since a single global temperature often hides class- or slice-specific failure modes.
Fair point, I only showed the averages here. Overall, calibration was beneficial, but it slightly worsened the ECE on one dataset (AGNews). So, yes, your point is right. (I will add a note about that.)
Thanks for checking. My guess is the temperature fit on the held-out split didn't match AGNews' test confidence profile, so a per-dataset or even per-class temperature might close that gap cheaply. Appreciate you adding the note.
I think there is a small typo in section 5.2 where the modified reward is defined. The line after the equation reads "where is 1 for a correct answer and 0 otherwise" instead of "where C is 1 for a correct answer and 0 otherwise".
Thanks for the note, I think it swallowed the $C$ here because of the latex formatting. Fixed this
A dedicated classifier probably still beats Jev on any single narrow job, as you note, but each one needs its own training pipeline, its own drift monitoring, its own retraining schedule. Run more than a handful of those tasks in parallel and I'd expect the maintenance surface, not the accuracy gap, to decide which approach an organization keeps running.
Jev collapses that into one thing to monitor instead of a dozen. The crossover then sits where the number of narrow classifiers a team maintains exceeds what that team can keep watched, rather than where Jev's accuracy catches up to a tuned task-specific model. Curious whether that tipping point is already showing up in how teams choose between the two approaches.
Yes, I agree with everything here, and that was the main point regarding Jev being so useful. Like you said, there is a maintenance burden for specialized models. Also, even if you develop and maintain all those, it's also going to be resource intensive to have like 500 different specialists running on your hardware (sure, one can have one main model and handle the rest with LoRA adapters, but still, it'd be lots of work)
I'm suspicious of the Tetris or other game-playing tests, because one of Jev's launch demos was playing a game, so this should be very much in-distribution where something like ModernBERT or GLiNER based models may have never trained on that type of decision.
That said I'm stoked to avoid fine-tuning for many low-volume decision tasks that were previously untenable due to time investment, and the interest plus signals like your late-breaking edit about OpenAI, and the Decisions output modality on OpenRouter show it's clearly a huge demand.
Regarding
> ModernBERT or GLiNER based models may have never trained on that type of decision.
I agree. But that probably goes also for a lot of things. I do believe the reason why Jev is so good is that it's training distribution just covers such a wide range of tasks.
Great article!
I am really curious if the next generation of Jev-like models could actually beat, say models that Tesla uses in production for full-self driving.
To me, it seems like we’re really narrowing down the path of training specialized models for tasks. You can really start feeling the generality of the intelligence, when you throw a task at hand and it just solves it. I see also similar things in robotics, e.g. growbot, just connect you phone to control ‘a body’ whatever it’s, it figures things out.
All of this reminds me the patato head theory of David Eagleman, that says brain is the ultimate decoder and whatever input you give it to it, it will find a way to use that input. Apparently, it is not just the brain that can do that, so we should really start accepting we’ve not only passed the turing test few years ago, we’ve almost achieved general intelligence.
Interesting point, but I remain a bit skeptical. A Jev-like model fully specialized on full-self-driving is still going to be better than a Jev-like generalist imho. Or, at least, at the same capability-level, a specialized model could potentially be smaller and faster.
Great article as always, articulating theory visually for mere mortals. One question about bag-of-words representation, which I struggle a lot to understand. You mention
"""
A bag-of-words representation results in these fixed-size inputs by assigning each word in a vocabulary its own position in a vector. We then count how often each word occurs in a document. For example, if we have a vocabulary of 50,000 unique words, it produces a fixed-size vector with 50,000 entries, regardless of whether the input consists of only ten words or 300k words
"""
How is it possible to know in advance the size of your vocabulary (unique words) considering that new documents, and hence new words, will come up in the future? Are there any predefined vocabularies provided by libraries which are subsequently used by classic classifiers? What is the standard (old) practice in the industry to address that?
Good question. It's based on the combined training set. So if there is a new word, it would simply not be counted, or you can have an "unknown" word placeholder (but there is probably no benefit in that). The CountVectorizer and TfidfVectorizer in scikit-learn, which I usually use for this, ignore those.
Thanks for another great article!
There is an open-source alternative to Jev, called Contrastive Language Model, that was released recently: https://contrastive-lm.notion.site/. It is not another Alpaca :) I am curious about your thoughts on it.
Thanks!
And yeah, examples like that are great (this one has the best docs I've seen so far), but my point stands: they are not nearly on the same level as Jev. For example, I just tried it and its IMDb Test accuracy is at just 82.90% (over 96% in Jev) and fails the Tetris test: https://sebastianraschka.com/videos/tetris-clm
Thank you for the comparison!
Agreed. However, CLM was tested on other games (Dino, Super Mario and Wikiracing) and found to be on par with Jev. It could be that it was not trained on Tetris. As you pointed out, Typesafe probably spent a lot more time and money on building Jev than the CLM researchers. It is still encouraging to see there is a decent open-source alternative out there that allows us to fine-tune specific domains, like Tetris, but for business use cases.
A standing “insufficient evidence / none of the above” option could stop the system from being forced to manufacture certainty when every substantive choice is poorly supported.
Was Jev trained to handle such an option? And if not, would it meaningfully populate an “I don’t know” choice if one were always provided?
Yeah, sth like that would be useful. Right now, one would have to go by confidence. Like if the chosen answer has a low confidence, you can have a manual check that then escalates this to a GPT-6 Astra-level model or so.
This is a proprietary model so we can't see how they implemented it. It's possible they could just run a small model like qwen3-4B (or smaller, plus quantization) with the right system prompt and get this output, especially if you strip off the output layer and replace it with this output format.
Also, if it's predicting numbers and it's an LLM, it'd be like giving you a spreadsheet to look at and answer questions by reading it and doing math in your head. It's totally possible, but it's not optimal, and won't be accurate if dump a non-trivial amount of data in.
None of this is assertion of what they're doing since we can't know. It's just considering the easiest way they could have done it, and what the limitations would be if you expect an LLM to do things like forecasting or number crunching. They could also hook it up to a tool with a side model like tspulse or timefm if someone sent in a giant bunch of numbers (hint, model provider!).
This is entirely possible. But I do expect that they put a bit more effort behind it and curated a pretty good dataset for further training/fine-tuning. (If I were to build something like this, I would probably start with a pre-trained base model and focus on post-training mostly.)
Correct me if I am wrong but you cannot upload files to Jev in order to give specific domain knowledge plus it does not have memory so incremental learning on a specific domain via gradual feed is also not possible. What would be the best approach to creating a domain specific agent using your own information source. I have tried locally trained agents but they are so incredibly slow on a decent laptop
Yeah, I think Jev is not multimodal (yet). Regarding memory: that's true for any LLM though; memory usually comes from external context files via the harness. You can just build your own memory this way.
I just wanted to thank you for writing this! It’s really clear and useful even to someone who doesn’t have a strong background in ML.
Thanks you, really glad to hear that it is useful (and still understandable)
Thanks for the comprehensive article. One thing I wanted to mention w.r.t. Section 5.1 and 5.3 is that there's been a lot of work on calibrating neural networks in recent years, and many of these works are a substantive departure from the temperature-scaling (and related) approaches from the 2010's, in particular with regard to robustness to covariate shifts, which is important in the high-dimensional neural-network setting. (Single-parameter temperature scaling tends to be very brittle in the presence of covariate shifts.) Achieving that robustness is arguably something fundamentally new in terms of capability.
To recap, for context on calibration: The key properties for using such classification models for conditional-branching decisions in agentic stacks (and related) is that they should be well-calibrated (under the definition chosen for the task) and informative/"sharp" (e.g., always predicting the mean might be "well-calibrated" in a theoretical sense for some chosen quantities of interest, but isn't particularly useful in practice).
The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics. http://arxiv.org/abs/2509.12760
[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269. https://doi.org/10.1145/3786335.3813214
The LSTM example caught me: a model that can track word order still did worse here than the simpler baseline. Keeping the training-from-scratch detail beside it matters. It stops that result from becoming another tidy rule about which kind of model always wins.
Super cool write up thanks.
I still have a hard time understanding how it can do any task out of the box. Is it an encoder then it needs to be tuned, if a decoder then it should be a bit slower, or i am missing something ?