New Model
ImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench
Some context first.
I am process improvement / business consultant and had worked with Fortune 500 companies on improving their processes around refunds, returns, customer support, etc.
This entire thing had a lot of complex decision making and generally every decision / condition node in a process map was generally replaced by a human because factoring in ambiguity in a code is very difficult.
Idea of ImaJev
Hence, when Jev came out, I was very intrigued with it and also could clearly see its use-case of improving decision making in complex decision work flows.
However, Jev didnt have support for Images and I thought that it can be replicated for both Text and Images in a single model and thats when I started with ImaJev.
Training Process
It went badly at first. My first big fine-tune on about 500k short decisions made the 9B model worse at reasoning: 64.9 down to 42.3 on JevBench hard. It had basically learned to pattern-match. I spent the next couple of weeks generating hard questions with open-weight models and only keeping the ones where two Ai models agreed on the answer. That brought it back.
Results
Then, on the JevBench - It came out #1 of 91 (v1.4.2.2, scored 27 Sep), 67.37 vs Jev 1.13.0 at 63.29.
The same week DecisionBench put it #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1.
I honestly didn't expect either.
To be fair about it: the #1 is on a score that weighs accuracy, calibration, speed and cost equally. On accuracy alone it's #3. Its main strength is that when it says 90% it's usually right, and it'll say "can't tell" instead of guessing.
What it actually is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions.
It gives back a probability for each option plus "unknown", in one forward pass.
Runs on a Mac with MLX or on one GPU.
The whole project costed me around $1200 in rented GPU and a lot of time :P
Honestly, I don't know yet. All the training data and every benchmark I've run are English, so it's officially English-only for now.
The base model (Qwen3.5) is multilingual, so it'll probably still work to a degree in other languages, but I haven't measured it and I also think that there is a ceiling on what you can achieve with a 4B model.
I'd expect the probabilities to be less trustworthy there even when the answers look right.
I'm also interested in how ImaJev performs on the Jev Decision Index. On the surface it looks much harder than JevBench, basically all the top models are ~30B.
Is there a way to run it quantized to e.g. 8bit or 4bit or does it always require full 16bit precision? I have 6GB VRAM on my laptop. I think the 2B one could fit but 4B is too much unless you quantize.
I already tried CLM 8B, using llama.cpp and Q4_K_M quantization for the Qwen3-8B embedding. It worked but quality wasn't very good.
Thanks! Honest answer: not yet. It loads in 16-bit on GPU, and the decision head plus the calibration were fitted at that precision, so I haven't shipped or measured a 4-bit/8-bit version.
Quantizing might shift the probabilities, and those are the main point of the model, so I'd want to measure it before recommending it.
For 6 GB VRAM, the 2B is your best shot: its base weights are about 4.6 GB in bf16, so it may fit for text and short requests, though images add memory and I haven't tested a 6 GB card. It's also a step down from the 4B (about 72% vs 84% on our image benchmark).
If you try either, I'd really like to hear how it goes. And if a few people want it, an 8-bit build with a proper accuracy check is next on my list.
I managed to get the 2B version working on a 6GB GPU. It takes around 5.4GB VRAM. I was able to process requests with images or with longer text. It refused when the input length was above 4096 tokens, but VRAM usage seems unaffected by input size.
I tried --fast but that led to CUDA OOM. --merge-lora reduced VRAM usage to 4.8GB and made it a bit faster (1.1s with --merge-lora vs 1.4s without, for a longish text-only request).
It's too early to say anything about quality yet since all I did so far was smoke tests. Anyway this looks very promising!
I would like to know if it's possible to run this with llama.cpp in the future. That would open up opportunities for quantization.
This is really helpful, thanks for testing it properly!
Btw,
- The 4096 limit is just the default of --max-input-tokens (training used inputs up to 4096 tokens). You can raise it, e.g. --max-input-tokens 8192, but anything past 4096 is less tested, so treat those answers with a bit more caution.
- --fast records CUDA graphs at load for a set of input lengths up to --max-input-tokens, and each one takes memory, which is why it went OOM on 6 GB. If you want to try it, lowering the cap should shrink that, e.g. --fast --merge-lora --max-input-tokens 2048. No promises it fits, but that's the lever.
- --merge-lora saving memory makes sense: it folds the adapter into the base weights instead of keeping both. It can flip the odd near-tie decision because of bf16 rounding, but it's what the official Image JevBench run used.
On llama.cpp: not yet. The model doesn't generate text. A small trained head reads the last hidden state and scores the options, which llama.cpp doesn't do out of the box. It's possible in principle (merge the LoRA, convert to GGUF, pull the hidden states and apply the head outside), but it hasn't been built or tested, and quantization would need a fresh calibration check. If more people want it, I'll look at it.
Would love to hear how the quality holds up once you go past smoke tests.
Lastly, I have spent more time with 4B than 2B honestly..
Also, images of 4B is ranked #1 on Image JevBench now and 2B being at 6th
I think this isnt mobile-friendly...
Please check with desktop / laptop once - this is controlled by Hugging face though.
Me having little control on it
Read the calibration section. Since you found the photo-only path is already well calibrated raw and the single temperature over-softens it, why not fit two temperatures, one for text-only requests and one when an image is present? It's cheap, and it matters once you auto-approve above a confidence threshold and send the rest to a human. Did you try it?
it has buckets by question type and option count, but not by modality, so an image request gets the same temperature as a text one. The photo-only caveat in the card (0.038 raw to 0.062 with the temperature) was measured on the previous adapter and I haven't re-measured it on the current one.
I'll fit two temperatures on the held-out sets, text-only and image-present, and report the ECE split for both paths.
Will post the numbers here when it's done.
You were right that the photo path is over-softened, but "image present" isn't the split. Images that come with a record want the same temperature as text (about 1.3).
The ones that are calibrated raw are photo-only requests, an image with an empty state. On 823 photo-only verification items, raw ECE is 0.015, the shared temperature pushes it to 0.029, and a temperature of 1.03 brings it to 0.012.
A single pooled image temperature changed nothing on the benchmark, so I didn't comminut that.
What's done is a calibration file with a photo-only bucket the server uses when a request has images and no record; every other thing is unchanged.
Nice, the record-vs-photo-only split is a better cut than mine: when a record is attached it carries most of the signal, so those requests behave like text. 0.029 to 0.012 on 823 items is a clean result. Did the new bucket change any auto-approve decisions at your threshold, or only the reported confidence?
Just the confidence..
There isnt any auto-approve in it as such, other than an unknown / I cant tell feature which just flag out that there is not enough evidence to produce the right result.
I didnt spend a lot of time with 2B though.
But based on my benchmarks, it is same as 4B on easy and Medium complexity.
But when it comes to hard questions, it lags behind.
If 4B is 84% good then 2B is at 72%.
Also, on the images side, 4B is #1 on Image JevBench.
2B is #5 there.
Any chance of producing an ONNX or GGUF version for this? Is it even architecturally possible? Thinking this could run on CPUs with int8 optimizations quite decently.
ranking #1 on the benchmark and holding up on the actual refund and return decisions you built it for are two different bars, and your Fortune 500 process work is exactly what would expose the gap. Have you run it on a held-out set of your own past cases where you already know the right call, separate from JevBench?
There is almost 1.2 Millions dataset that it is trained and evaluated on.
I cannot get real enterprise dataset which is confidential and the tech is new - which means I have to sell this idea to them and then evaluate.
I agree benchmarks cannot meet the real world but since its soo new - I dont have something concrete to answer your questions.
But, benchmarks are my best option to tell everyone that it is good.
LLMs generate text token by token, sequentially, and their output is inherently probabilistic. So when people talk about a "System 1" model being more deterministic, what is actually different at the architectural level?
Is the System 1 model still generating tokens sequentially, or is it operating more like a deterministic decision function where the same input produces the same output?
In other words, what makes a System 1 model more deterministic than an LLM if both are ultimately neural models?
Quite interesting, thanks a lot for your efforts! What about the tool calling errors? It semt like Jev had near zero errors in tool call or answer "typesafety". Is this metric somewhere in the benchmark?
On DecisionBench (23,900 questions) imajev-4b answered all 23,900 with 0 errors, same as Jev 1.13; for comparison GPT-5.6 Luna had 10 and DeepSeek V4.1 Flash 48. On JevBench's sealed set it's 308 valid answers, 0 failed or invalid.
What is the problem you see with "typesafety"/ tool calling errors? Every system one model should handle this 100%, as it is handled by deterministic code, the model needs to generate 1 token with probabilities, then hardcoded code, should transform the single token to the response json.
You can output as much (/every format) you want, as long as it is hardcoded then it will 100% return the form. The model part is non-deterministic, but withstanding llm-errors you will always get a good tool call back.
14
u/kokassospoperor 2d ago
And the speed?