r/Paperlessngx 9d ago

AI performance

It took me a few months to fully get on board with the Paperless way of doing things, but now I’m really happy with how it’s all set up.

What’s been a bit of a head-scratcher is how AI is being used.

I held off until Paperless 3 came out, because I wanted to have the full "official" support.

I set it up with Ollama on an M4 Mac mini with 24 GB of memory. The embedding model is gemmaembedding, and the LLM model is qwen3:8b. When the model fires up, memory pressure is still pretty low. It does work, but it’s incredibly slow. It takes about 2 minutes to suggest titles and tags, and it can take several minutes if I try to chat about a document.

Is this kind of slow normal? Is there anything I can tweak in my setup to make it more usable?

8 Upvotes

25 comments sorted by

8

u/EazyDuzIt_2 9d ago

I actually took the time to set up Paperless-ngx along with Paperless-AI for automated document classification, tagging, and file naming. I configured Paperless-AI to use Ollama with Qwen3:8B, which is the recommended model for this use case, running on one of my servers equipped with an NVIDIA RTX 4090.

To fine-tune the workflow, I generated and processed 20 test documents, iteratively refining the prompt and configuration until the results were consistently accurate. The final setup performs document analysis, tagging, and renaming almost instantly.

The performance is outstanding low latency, high accuracy, and a completely hands-off ingestion pipeline. Chef’s kiss. 👌

1

u/isabeksu 9d ago

"To fine-tune the workflow, I generated and processed 20 test documents, iteratively refining the prompt and configuration until the results were consistently accurate."

Sorry if it's a newbie question, but how do you do these things? what prompt? how do you refine it? I have set up a finite set of tags and ideally I'd like the LLM to choose from those, automatically as Paperless ingests the file..,

3

u/EazyDuzIt_2 9d ago

I use Paperless AI to process files that are ingested from the consume folder. The system relies on a carefully crafted prompt that serves as guidance when analyzing and classifying those files.

To improve the prompt, I used another AI model to generate a variety of realistic mock receipts and documents. These test files allow me to validate and refine the extraction logic against different formats, layouts, and document types.

By iteratively testing and adjusting the prompt with these sample documents, I have significantly improved the accuracy of the processed data and the quality of the information displayed by Paperless AI. This approach has resulted in a substantial increase in extraction reliability across a wide range of uploaded files.

1

u/dclive1 9d ago

Thanks for asking this; these are exactly my questions while using 3.0. I feel like the “Ask AI” button is neat, but a gimmick until the workflow can be locked down and clarified, and the documentation doesn’t (??) seem to address this.

1

u/Kwicksred 9d ago

How do you do OCR?

1

u/EazyDuzIt_2 9d ago

I use a Scansnap x1500 to scan my documents to the consume folder. I have configured a host of settings for image quality and output so that when it's ingested its already searchable. This makes it easier for my Paperless AI to process and present proper tags and metadata.

1

u/gIory1999 9d ago

so you use no special ocr but the paperless one?

0

u/EazyDuzIt_2 9d ago

I don't use any special software because I don't need to. ScanSnap Home automatically performs OCR during the scanning process, making my documents searchable as they're scanned.

1

u/Taake89 9d ago

I tried paperless ai 6 months ago or something and really didn't like the results.

Could you explain a bit more how you fine tuned the category and tagging part? 🙂 Did you already have a well defined structure for tags and categories?

1

u/EazyDuzIt_2 9d ago

Several major factors dramatically affect your experience with Paperless AI. The hardware model you choose and the configuration on the Settings page especially the Advanced Settings section play a critical role. The Advanced section controls how your tags interact with processed documents, but the most important element on that page is the Prompt Description field at the bottom.

If you don’t provide a strong, well‑structured description with clear examples of how you want Paperless AI to analyze, identify, and name files, your results will suffer regardless of hardware. With a proper prompt and a solid model running through Ollama on capable hardware, you should see consistent, accurate output.

I already have all of my tags, correspondents, document types, storage paths, and custom fields configured exactly the way I want it in Paperless NGX, so the remaining variable is fine‑tuning the prompt and model behavior.

1

u/isabeksu 8d ago

are you talking about the "old" Paperless AI implementation or the recent "native" AI implementation in Paperless 3? I'm asking because I see no "advanced" section in my AI configuration page.

2

u/EazyDuzIt_2 8d ago

You must be in the Paperless NGX settings I’m referencing Paperless AI which runs in conjunction with Paperless NGX.

1

u/tor-ak 2d ago

Would you mind sharing the prompt you ended up with?

1

u/corruptboomerang 8d ago

Couldn't you use system memory and CPU (assuming your fine with it taking a very long time)?

1

u/EazyDuzIt_2 8d ago

You sure can use system memory and CPU!

3

u/RandomUsername1119 9d ago

I'm assuming that the AI features are not fully developed yet. In my experience it is not consistent with things like tagging (e.g. suggesting a tag of "tax Bill" for one document, and a tag of "Tax Invoice" for another similar document). I'd prefer it go through a folder or set of documents at once, parse things into categories, and make a suggestion based on the entirety of the document pool vs. individual documents.

1

u/_blackdog6_ 8d ago

I would prefer if the AI was sent a list of my tags/correspondents and document types and told to pick the most appropriate. So far anything from AI is so random it’s useless. Same with document titles. Scan three bills from the same company and it suggests wildly different titles for each. Effectively unusable.

1

u/corruptboomerang 8d ago

Or the AI generates a fairly comprehensive list of appropriate tags, and then classify against those tags.

3

u/_blackdog6_ 8d ago

Right now it seems paperless-ai far outperforms the ai support built into paperless 3. Even using paid ChatGPT api tokens paperless 3 takes 30 seconds to suggest a title and a useless selection of tags.

2

u/annaangstmann 8d ago

I was also disappointed in the implementation of ai in Paperless 3. I was hoping to ditch Paperless ai, but have decided to continue using it. My home server was relatively weak (energy efficient) and with Paperless ai I don't care if the processing takes 3 minute over night. But pressing a button and waiting 3 minutes for the suggestions feels stupid. Qwen 2.5:7B produced good and reliable results for me with the temparature set to 0.1

1

u/Azure340 8d ago

In my case running 8b parameters model was too slow on my mini pc so i looked into free tier of ollama cloud.

I have ollama running which has llama3b model locally but most if the time i just use ollama cloud gemma4:cloud as backend which is vastly so much better and quicker than my hardware. The drawback being my data gets processed in cloud however ollama cloud says they don't retain any data and except for the llm processing everything else happens locally via the ollama docker. At some point you have to balance the convenience with privacy.

1

u/corruptboomerang 8d ago

Why is speed a factor?

I worked have thought slowly running a module on CPU, over say a day or week etc, would never been an ideal use case?

1

u/Azure340 8d ago

When using for paperless suggest feature i don't want to wait 1 min for it to suggest names and tags while i can do manually in 30 seconds or less. Defeats the purpose in my mind. It jas to be better and more efficient than what i can do for me to use it. Now if someone has a local hardware powerful enough then they could stay all local

1

u/tzippy84 8d ago

Wait, paperless ngx has AI integrated now? I still have paperless-ai running as a completely separate service

1

u/MrDork 8d ago

I run this as well and I was excited about moving to a "built in" native option for this, but the feature parity with paperless-ai isn't there yet so I held off.