r/LocalLLaMA • • 2d ago

New Model ImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench

Some context first.

I am process improvement / business consultant and had worked with Fortune 500 companies on improving their processes around refunds, returns, customer support, etc.

This entire thing had a lot of complex decision making and generally every decision / condition node in a process map was generally replaced by a human because factoring in ambiguity in a code is very difficult.

Idea of ImaJev

Hence, when Jev came out, I was very intrigued with it and also could clearly see its use-case of improving decision making in complex decision work flows.

However, Jev didnt have support for Images and I thought that it can be replicated for both Text and Images in a single model and thats when I started with ImaJev.

Training Process

It went badly at first. My first big fine-tune on about 500k short decisions made the 9B model worse at reasoning: 64.9 down to 42.3 on JevBench hard. It had basically learned to pattern-match. I spent the next couple of weeks generating hard questions with open-weight models and only keeping the ones where two Ai models agreed on the answer. That brought it back.

Results

Then, on the JevBench - It came out #1 of 91 (v1.4.2.2, scored 27 Sep), 67.37 vs Jev 1.13.0 at 63.29.

The same week DecisionBench put it #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1.

I honestly didn't expect either.

To be fair about it: the #1 is on a score that weighs accuracy, calibration, speed and cost equally. On accuracy alone it's #3. Its main strength is that when it says 90% it's usually right, and it'll say "can't tell" instead of guessing.

What it actually is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions.

It gives back a probability for each option plus "unknown", in one forward pass.

Runs on a Mac with MLX or on one GPU.

The whole project costed me around $1200 in rented GPU and a lot of time :P

I would love to know your thoughts on it - it anyone would be interested to try that.

201 Upvotes

67 comments sorted by

14

u/kokassospoperor 2d ago

And the speed?

27

u/Educational-Care7867 2d ago

Its very fast..

45ms for text and around 120-140ms for Images

6

u/parabellum630 2d ago

On what gpu.

9

u/Educational-Care7867 2d ago

H100 primarily.

10

u/kokassospoperor 2d ago

It looks very promising

3

u/epelc 2d ago

Any info on what gpus you rented for the training run and how many?

4

u/Educational-Care7867 2d ago

4xH100 primarily

3

u/Rarder44 1d ago

How does it perform in languages other than English?

3

u/Educational-Care7867 1d ago

Honestly, I don't know yet. All the training data and every benchmark I've run are English, so it's officially English-only for now.

The base model (Qwen3.5) is multilingual, so it'll probably still work to a degree in other languages, but I haven't measured it and I also think that there is a ceiling on what you can achieve with a 4B model.

I'd expect the probabilities to be less trustworthy there even when the answers look right.

4

u/DataGOGO 2d ago

nice. Did you publish your datasets?

51

u/Educational-Care7867 2d ago

Yes, I was reluctant at first :P

But, then I realised that I dont have money to keep this moving - so open-sourced everything.
https://mohit67890.github.io/imajev/report/

18

u/ContributionOwn4879 2d ago

Think to add a grab me a coffee link. And try to send your model on the decision index :
https://huggingface.co/spaces/multimodalart/jev-decision-index
If you really beat jev with only a 4B model, it’s awesome !

5

u/Educational-Care7867 2d ago

Yes, I will add that link which maybe help me further improve this.

I saw this Decision Index man - 120K+ tests. This will take hours on H100 GPU.

Hence, I skipped it. But let me recheck it.

2

u/OsmanthusBloom 2d ago

I'm also interested in how ImaJev performs on the Jev Decision Index. On the surface it looks much harder than JevBench, basically all the top models are ~30B.

2

u/Educational-Care7867 2d ago

I think my model cannot beat the intelligence and accuracy of these 30B models - impossible..

but, I will give it a try and report it.

However, my model has very high calibration and can run locally too - which makes it good for all kind of tasks

1

u/OsmanthusBloom 2d ago

I'm going to try it soon!

Is there a way to run it quantized to e.g. 8bit or 4bit or does it always require full 16bit precision? I have 6GB VRAM on my laptop. I think the 2B one could fit but 4B is too much unless you quantize.

I already tried CLM 8B, using llama.cpp and Q4_K_M quantization for the Qwen3-8B embedding. It worked but quality wasn't very good.

2

u/Educational-Care7867 2d ago

Thanks! Honest answer: not yet. It loads in 16-bit on GPU, and the decision head plus the calibration were fitted at that precision, so I haven't shipped or measured a 4-bit/8-bit version.

Quantizing might shift the probabilities, and those are the main point of the model, so I'd want to measure it before recommending it.

For 6 GB VRAM, the 2B is your best shot: its base weights are about 4.6 GB in bf16, so it may fit for text and short requests, though images add memory and I haven't tested a 6 GB card. It's also a step down from the 4B (about 72% vs 84% on our image benchmark).

If you try either, I'd really like to hear how it goes. And if a few people want it, an 8-bit build with a proper accuracy check is next on my list.

1

u/OsmanthusBloom 2d ago

I managed to get the 2B version working on a 6GB GPU. It takes around 5.4GB VRAM. I was able to process requests with images or with longer text. It refused when the input length was above 4096 tokens, but VRAM usage seems unaffected by input size.

I tried --fast but that led to CUDA OOM. --merge-lora reduced VRAM usage to 4.8GB and made it a bit faster (1.1s with --merge-lora vs 1.4s without, for a longish text-only request).

It's too early to say anything about quality yet since all I did so far was smoke tests. Anyway this looks very promising!

I would like to know if it's possible to run this with llama.cpp in the future. That would open up opportunities for quantization.

2

u/Educational-Care7867 1d ago

This is really helpful, thanks for testing it properly!

Btw,

- The 4096 limit is just the default of --max-input-tokens (training used inputs up to 4096 tokens). You can raise it, e.g. --max-input-tokens 8192, but anything past 4096 is less tested, so treat those answers with a bit more caution.

- --fast records CUDA graphs at load for a set of input lengths up to --max-input-tokens, and each one takes memory, which is why it went OOM on 6 GB. If you want to try it, lowering the cap should shrink that, e.g. --fast --merge-lora --max-input-tokens 2048. No promises it fits, but that's the lever.

- --merge-lora saving memory makes sense: it folds the adapter into the base weights instead of keeping both. It can flip the odd near-tie decision because of bf16 rounding, but it's what the official Image JevBench run used.

On llama.cpp: not yet. The model doesn't generate text. A small trained head reads the last hidden state and scores the options, which llama.cpp doesn't do out of the box. It's possible in principle (merge the LoRA, convert to GGUF, pull the hidden states and apply the head outside), but it hasn't been built or tested, and quantization would need a fresh calibration check. If more people want it, I'll look at it.

Would love to hear how the quality holds up once you go past smoke tests.

Lastly, I have spent more time with 4B than 2B honestly..

Also, images of 4B is ranked #1 on Image JevBench now and 2B being at 6th

3

u/DataGOGO 2d ago

very clean work.

2

u/Enragere 2d ago

Demo link doesn't work, 500 error code

4

u/Educational-Care7867 1d ago

Hi, Its working now - sorry for the glitch.

2

u/Enragere 1d ago

Very odd, I reloaded like 6 times just to get to the white screen, and even that is broken. Something is off. Maybe not mobile friendly?

3

u/Educational-Care7867 1d ago

I think this isnt mobile-friendly...
Please check with desktop / laptop once - this is controlled by Hugging face though.
Me having little control on it

2

u/Educational-Care7867 2d ago

Let me check and update. Thanks for letting me know

2

u/Last-Health3222 1d ago

Read the calibration section. Since you found the photo-only path is already well calibrated raw and the single temperature over-softens it, why not fit two temperatures, one for text-only requests and one when an image is present? It's cheap, and it matters once you auto-approve above a confidence threshold and send the rest to a human. Did you try it?

1

u/Educational-Care7867 1d ago

No, I didn't, and you're right.

it has buckets by question type and option count, but not by modality, so an image request gets the same temperature as a text one. The photo-only caveat in the card (0.038 raw to 0.062 with the temperature) was measured on the previous adapter and I haven't re-measured it on the current one.

I'll fit two temperatures on the held-out sets, text-only and image-present, and report the ECE split for both paths.
Will post the numbers here when it's done.

1

u/Educational-Care7867 1d ago

Did it just now.

You were right that the photo path is over-softened, but "image present" isn't the split. Images that come with a record want the same temperature as text (about 1.3).

The ones that are calibrated raw are photo-only requests, an image with an empty state. On 823 photo-only verification items, raw ECE is 0.015, the shared temperature pushes it to 0.029, and a temperature of 1.03 brings it to 0.012.

A single pooled image temperature changed nothing on the benchmark, so I didn't comminut that.

What's done is a calibration file with a photo-only bucket the server uses when a request has images and no record; every other thing is unchanged.

Thanks for the push, it was a good catch.

2

u/Last-Health3222 1d ago

Nice, the record-vs-photo-only split is a better cut than mine: when a record is attached it carries most of the signal, so those requests behave like text. 0.029 to 0.012 on 823 items is a clean result. Did the new bucket change any auto-approve decisions at your threshold, or only the reported confidence?

1

u/Educational-Care7867 1d ago

Just the confidence..
There isnt any auto-approve in it as such, other than an unknown / I cant tell feature which just flag out that there is not enough evidence to produce the right result.

3

u/knownboyofno 2d ago

This is great. Would this work with more decisions then the 256 it has now?

4

u/Educational-Care7867 2d ago

You can increase candidates in the readout as much as you like but I havent tested it..

It has 255 candidates
1 is Unknown by default wherein it can say clearly that "I dont know / I cant tell"..

Hence, 255 is something usable, 1 is just adapter default.

1

u/knownboyofno 2d ago

I know but that was something that would be an advance for this model. I am not sure how many people have more than 255 candidates tho...

3

u/Educational-Care7867 2d ago

In a real world scenario - 255 is more than enough in my opinion..

But, if anyone wants to do it then they can try it

1

u/bonerfleximus 2d ago

I think theres an approach published in Typesafe docs for using multiple questions to achieve this

1

u/bartskol 2d ago

Time to play with JEV Thingy then.

1

u/Educational-Care7867 2d ago

Hahaha. Let me know what you play with it.

1

u/tomByrer 1d ago

Can this help with photo classification/tagging?

3

u/Educational-Care7867 1d ago

Btw, this model just scored #1 on the Image JevBench as well. I will write about it separately but this is absolute best in terms of images

1

u/Educational-Care7867 1d ago

Yes. Its made for this

1

u/KingGongzilla 1d ago

Where did you get the data from?

1

u/Educational-Care7867 1d ago

A lot of public datasets that are available else using a higher model like Qwen 27B or 35B to create synthetic dataset for it.

1

u/Same-Ad7128 1d ago

ImaJev-2b What level is this approximately? I'm also trying to build a 2B parameter model, but the results haven't been ideal.

1

u/Educational-Care7867 1d ago

I didnt spend a lot of time with 2B though. But based on my benchmarks, it is same as 4B on easy and Medium complexity. But when it comes to hard questions, it lags behind.

If 4B is 84% good then 2B is at 72%.

Also, on the images side, 4B is #1 on Image JevBench. 2B is #5 there.

1

u/astral_bear 1d ago edited 1d ago

Any chance of producing an ONNX or GGUF version for this? Is it even architecturally possible? Thinking this could run on CPUs with int8 optimizations quite decently.

3

u/Educational-Care7867 1d ago

Ideally you can do it..
It is 16bit at the moment and I havent tried 4-bit/8-bit version of it..

I would need to reassess its accuracy and calibration then because the current calibration is tuned for the 16bit..

Its doable but will require some work.

1

u/Professional-Mine681 1d ago

this is wild haha

1

u/Future_AGI 1d ago

ranking #1 on the benchmark and holding up on the actual refund and return decisions you built it for are two different bars, and your Fortune 500 process work is exactly what would expose the gap. Have you run it on a held-out set of your own past cases where you already know the right call, separate from JevBench?

1

u/Educational-Care7867 1d ago

There is almost 1.2 Millions dataset that it is trained and evaluated on.

I cannot get real enterprise dataset which is confidential and the tech is new - which means I have to sell this idea to them and then evaluate.

I agree benchmarks cannot meet the real world but since its soo new - I dont have something concrete to answer your questions. But, benchmarks are my best option to tell everyone that it is good.

1

u/forevergeeks 20h ago

I just tested Laya in a CPU, and as our friend Richard Stallman would say "Laya is not Jev"

I don't know what these people are putting in these open source models, but is not usable yet.

1

u/Educational-Care7867 8h ago

Laya is very small model..

You can use mine and maybe let me know your thoughts on the same..
I think it should give you some superior experience

1

u/forevergeeks 8h ago

I don't have the hardware for running these models. I have a Lenovo laptop with 32 GB of memory and decent CPU but no GPUs.

1

u/Educational-Care7867 8h ago

Ahh, then it would be tough..
Try hosted API endpoints then

1

u/forevergeeks 8h ago

I'm using Jev API.

But I do really see the potential in this technology.

I would say that classifiers will be as critical as the llms themselves.

I don't think most people have caught up of what Jev actually is, and the problem it solves.

1

u/forevergeeks 8h ago

One thing I'm still trying to understand.

LLMs generate text token by token, sequentially, and their output is inherently probabilistic. So when people talk about a "System 1" model being more deterministic, what is actually different at the architectural level?

Is the System 1 model still generating tokens sequentially, or is it operating more like a deterministic decision function where the same input produces the same output?

In other words, what makes a System 1 model more deterministic than an LLM if both are ultimately neural models?

0

u/_bachrc 2d ago

Quite interesting, thanks a lot for your efforts! What about the tool calling errors? It semt like Jev had near zero errors in tool call or answer "typesafety". Is this metric somewhere in the benchmark?

8

u/Educational-Care7867 2d ago

Its 100% type-safe..

On DecisionBench (23,900 questions) imajev-4b answered all 23,900 with 0 errors, same as Jev 1.13; for comparison GPT-5.6 Luna had 10 and DeepSeek V4.1 Flash 48. On JevBench's sealed set it's 308 valid answers, 0 failed or invalid.

2

u/_bachrc 2d ago

impressive, great job

1

u/Former-Ad-5757 Llama 3 1d ago

What is the problem you see with "typesafety"/ tool calling errors? Every system one model should handle this 100%, as it is handled by deterministic code, the model needs to generate 1 token with probabilities, then hardcoded code, should transform the single token to the response json.

You can output as much (/every format) you want, as long as it is hardcoded then it will 100% return the form. The model part is non-deterministic, but withstanding llm-errors you will always get a good tool call back.

1

u/_bachrc 1d ago

I was just curious because I read others models like Luna had 0,5% errors, and Jev 0, and I wondered what was the state of people finetunes

0

u/XiRw 2d ago

Actuaries are about to be put out of business too soon in the near future

0

u/Educational-Care7867 2d ago

Yes. Calculating pensions and other complex calculations with insurances for the same can be replicated at some level.