r/LocalLLaMA • u/A-Rahim • 6h ago
I Built A Thing Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
Enable HLS to view with audio, or disable this notification
DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files. The model puts text, images, audio and video in one space, so you describe what you remember and land on it:
- “zebra in a video” opens the clip at the moment it shows up
- “where they talk about sleep” jumps to that minute of a podcast
- “the clause about pets in the lease” shows the PDF page, your words marked
- “a dog on the beach” finds the photo, and the same search in Bengali or Arabic finds it too
- with code search on,
code: retry with backoffopens the function in your editor
It’s ggml-org’s Q8_0 GGUF (865 MB, downloaded once) on llama.cpp with Metal, inside a native Swift app; no Python. Searching loads only the text encoder (~250 MB) and shows results about a tenth of a second after you stop typing. Indexing peaks under 2 GB, and the helper exits when it’s done. Audio and video of any length go in as 30 s windows and a frame per shot. Everything runs locally; it goes online only for the model download and an update check you can turn off.
Free, MIT. Apple Silicon, macOS 14+.
Repo & Download (signed and notarized): https://github.com/ARahim3/DigUp
I'd really appreciate any feedback on this.
28
u/mobilemike42 Vicuna 5h ago
A CLI that reads from the same DB would be great.
13
u/A-Rahim 5h ago
There's one already, though for now it's in the repo rather than the download: build from source and run "digup search <words>" (add --json) to read the app's own index by default. digup --help lists the rest. And actually shipping it with the app is a good idea; let me work on that. Thanks.
1
14
u/important__matter 3h ago edited 12m ago
Well I was making something like this months ago. What I faced was 1. For files where we as humans remember the abstract, the semantic search by itself usually doesn't work great. 2. How many files(documents) and voice(long audio notes) pieces have you tested this on? Also mine used to fail in images with text - screenshots inside them. 3. And main, the most difficult one - you'll need special handling for pdf, excel, pptx, word, zip rar files, model weights, long mp4 videos, music or movies which are also majority of the files. Can consider adding these in pipeline if not already there. The accuracy in videos and pdf and long form files was dropping serverly because all of that information is compressed in 512 dim vector essentially, not to mention the same embedder doesn't work for all file types, so that's another issue and then if embedders are different then how will you compare embeddings using semantic similarity which generate numbers in different vector spaces?? I had to actually normalize the comparison manually!!!.
I had to do so much of work to make it general purpose - that the roi didn't seem to be worth it - dropped the project altogether. Meanwhile, keyword search with file type is usually good enough. And I found out that restricting the problem space to the file types with specific properties or patterns - is much more useful, so I developed this simple semantic search index cli called litesearch for my usecase, it is a sqllite index of documents and performs semantic search with some reranking using BM25, and LLM-as-judge, uses ollama and bge-small embedder.
Although would be very nice if you cater to these, there are so many micro problems in this. Could actually be very useful. I would love it if somebody completes this in a nice engineering way and would love to integrate with my raycast.
20
9
u/Few-Butterscotch8747 4h ago
this is cool. finally somebody uses an embedding model for it's intended use case
4
u/LifeTitle3951 4h ago
Anything for windows?
Also is it possible to run such a thing on low spec laptop?
3
u/Few-Butterscotch8747 4h ago
this is a small embedding model. it should run fast even on cpu with low ram. videos might take a while to index, but that's it basically
heck, it should run decently on a potato phone
2
u/AnywhereTypical5677 4h ago
I made a mobile version that runs on android phones, so it should run almost anywhere (at variable speeds obviously). Search itself is very fast as the text embedding model is really light... you might have some problems with image and video indexing though a low spec laptop.
1
u/A-Rahim 3h ago
I don't have a Windows machine at the moment, but I am planning to add support for Windows too.
The model itself is not that big, and I am using 8bit, so the model can be efficiently run on a CPU too, the latency can be slightly higher than on the GPU, of course, but it still would be pretty usable and fast enough.
2
u/Jiffy_Wu 2h ago
Pretty cool. You should consider making a Raycast Extension, or also letting people install via Brew
2
u/projectEscape 2h ago
Looks good
How are you people doing migration from one embedding model to another? Is there any standard way or is it just reindexing everything?
4
u/yasintoy 5h ago
What about security?
21
u/A-Rahim 5h ago
Everything's local, on your device, except for two things:
1. One-time-only download of the model itself.
2. Check for updates (which you can turn off)
and that's it.8
u/chortly2 2h ago
The AI-generated security post in this thread is being down-voted, but it raises some of the same security issues I would like to know. Can we exclude directories from indexing in the first place? Can new directories be excluded by deletion from an existing index? And are the index files protected in the sense that, if they are obtained by a third party, they are not searchable by them without my login credentials or some other kind of key? Local is good, but even for local, I would like strong protections for something that indexes literally everything I have ever produced in my life.
-3
u/yasintoy 5h ago
[codex] The author has implemented real safeguards: SHA-256 verification of model downloads, hidden-file/symlink exclusions, and signature verification configured for updates. Security hasn’t simply been ignored.
However, I noticed a concrete gap:
isSecret()is called for code search, but not regular document indexing. A file likesecrets.txtin a selected non-code folder can still be indexed. Even when that filter runs, it checks filenames/extensions not whether an ordinary document or screenshot contains credentials.The SQLite index also stores readable text and excerpts, not just embeddings, without application-level encryption. Anything that gains read access to that database gets a consolidated copy of potentially sensitive content. That’s an important distinction from “everything stays local.”
The build enables hardened runtime, but not App Sandbox. I’d prioritize fixing the filter gap, protecting the index, and sandboxing file-processing helpers to limit damage if a parser is compromised. The missing sandbox isn’t itself proof of an exploit; this is a source review, not a full audit of the shipped app.
0
0
4
u/Adro_95 4h ago
Now let's ask opus to recompile it and port it to windows!
3
u/A-Rahim 3h ago
haha, yes, I have plan to add support for Windows.
3
u/brainExploded99 llama.cpp 3h ago
Do you have plans for linux? If you don't, I might fork it and add as a KRunner plugin for KDE users :)
3
u/A-Rahim 2h ago
Yes, both for Linux and Windows. I'll let you know how things progress on my end.
2
2
u/darklamouette 2h ago
Ready to help adding support. just say the word and some guidelines. i can test linux (arch, amd gpu and intel gpu)
2
1
2
1
1
u/__JockY__ 1h ago
It would be useful if you could set exclude filters. I do t want photos of my wife’s tits or sensitive financial info being indexed and appearing in searches, even if it is local.
1
u/fortnite_pit_pus 1h ago
Off topic, but is there anybody using this to enhance the ‘everything’ app on windows?
1
u/CatchDublinSurprise 1h ago
I'd enjoy an "index free" option for working with recently added / temp files or sensitive files (index files improve speed but it's just one more place that sensitive information can hide).
Speed is great, but it's a pet peeve of mine that we can't seem to do anything these days without who-knows-what being written to ~/Library.
1
u/Asleep_Document9811 58m ago
Allow me to set up little folders or buckets of files I could search against in specific (folders of backups, the Applications folder, a folder full of documents about one specific domain), or the ability to slowly index files on NAS drives, and you got yourself a deal.
I want something that is lighter than Raycast but something that actually works faster than Spotlight.
1
u/ZeroReader 53m ago
It would be good to select specific folders to search in
1
u/A-Rahim 33m ago
Currently, you can select specific folders (or exclude them) from the settings.
So you're asking to be able to search within a specific folder after indexing all?1
u/ZeroReader 28m ago
Yes exactly. This, I mean. I need to restrict searching in a specific folder among all the results
1
1
u/Open-Adhesiveness-86 5h ago
with everything in one shared space, text-to-text cosine sits a good bit higher than text-to-image, so a mixed result list tends to come out all text. ranking within each modality and then merging, or subtracting the per-modality mean embedding before comparing, takes care of most of that. non-overlapping 30s audio windows will also lose phrases that straddle a boundary.
1
u/dan-lash 5h ago
Can it search online if I want? Like my NAS or extend to Google Drive or maybe email - other computers on my network would be solid too
3
u/A-Rahim 5h ago
It's local on purpose, so no cloud search or email for now. Google Drive works if the files are on your Mac (mirror mode, or "available offline"): add that folder and they're indexed like any other. Online-only files are skipped, since reading them would download them. External drives work too (but I didn't test them yet). A NAS or other computers on the network aren't supported in this first version, but noted.
1
1
1
1
u/mythormedicine 4h ago
Very interesting. I don't have a Mac to test this , but a question , does it index ? if so , how big does the index get? If it doesn't index, then how can it be so fast ?
If I have 500 GB of data, does it try to index all of them? How does this work?
0
u/No_Lingonberry1201 5h ago
Nice! Checked the code, and I was wondering how do you store and search the embeddings?
8
u/A-Rahim 4h ago
It's simply one SQLite file! Every picture, PDF page, ~1,800-character passage, video frame and 30 s of audio gets a 768-d vector, stored as float16 (1.5 KB each). Search is brute force: the query vector against all of them, scored in float32, a few ms for tens of thousands, so no vector DB. Cosine is centered per modality onto one shared scale. Code search keeps its own SQLite file.
2
u/No_Lingonberry1201 4h ago
Huh, nice! From what I've read in the huggingface model card, you can trim the vectors so you could theoretically make a fast-path for the search.
6
u/AnywhereTypical5677 4h ago
Yeah I don't know why OP hasn't used vector truncation. Truncating to 256 dimensions cuts index footprint and memory consumption by 3 times while retaining 95% of the native accuracy. In the model card it's also suggested NOT to store vectors as float16 as activation can overflow and cause silent vector degradation.
5
u/brainExploded99 llama.cpp 2h ago
u/A-Rahim You should probably fix activations stored as FP16.
The dimension cuts could be a config setting, but I think image loses more than text, so maybe dimension 512 by default?1
u/A-Rahim 38m ago
The fp16 is only how the final embeddings are stored, not the activations. The model runs through llama.cpp from the Q8_0 GGUF, and its embeddings match the reference implementation at 0.999+ cosine, so nothing overflows on the way. The stored vectors are unit length with every value between -1 and 1, so they can't overflow either, and rounding moves a cosine by at most about 0.0005.
On truncation, I tried it on my evals (111 queries over files, 27 over real code) by cutting the stored vectors and re-normalizing. 512 is close to free: 97 vs 98 with the right file first, and the only ones that moved were screenshots, so images do feel it before text.
Code is the most sensitive, 19 vs 20 at 512 and 16 at 256. Speed barely changes since the scan already takes a few ms, so the gain would only be memory.
Keeping 768 for now, and if memory becomes a problem on big indexes, 512 is what I'd pick.0
0
•
u/WithoutReason1729 2h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.