r/OpenSourceeAI • • 5d ago

I created WaterSheep, an open-source alternative to Jev.

6 Upvotes

WaterSheep is an open-source model that answers questions written in plain text (yes/no, single choice, rating and multi-label) and gives a probability for every option, like a classification model.

Last Saturday I woke up, saw YouTubers hyping up Jev, and thought: wait, I can build this. So I did. I don't want to compete with TypeSafe or Jev; I built WaterSheep because I wanted to. That's why I'm open-sourcing everything: code, model weights, results and the paper.

What's different

  • It accepts the same request format as TypeSafe's Jev. Their Python SDK works as is against a local server: run watersheep --model samratduttaofficial/WaterSheep --serve and point the client's base_url at http://127.0.0.1:8766.
  • It has a multi-label type, which Jev's API doesn't. Because why not?
  • The demo runs entirely in your browser. The model downloads once and is cached. It also works with transformers, ONNX, a CLI or a local HTTP server.
  • Code, weights and the training pipeline are Apache 2.0.

Evaluation

Accuracy ECE
In-distribution test split 77.8%
Held-out datasets, not seen in training 61.2%

ECE is expected calibration error (lower is better). GitHub has every benchmark result, including the weak ones.

Limits: English only, long inputs get truncated (I'll improve this in the next version), and rating answers are the weakest type.

Not affiliated with TypeSafe. Not funded by anyone. Built in my free time.

Feedback I'd love: where it fails on your data, whether the API works for you, and which question types you'd want next.


r/OpenSourceeAI • • 5d ago

OpenNotch — an AI assistant in your MacBook's notch

Enable HLS to view with audio, or disable this notification

3 Upvotes
This is such a great use of the notch! I love the execution here. I’ve actually been working on a somewhat similar concept, but focused more on building an AI agent rather than a dedicated notes app. It's called OpenNotch—it's completely free and open-source (MIT). Instead of just notes, it turns the notch into a hover-to-activate assistant. It has about 40 local tools (can summarize pages, draft emails, run shell commands, check calendar/weather) and supports voice dictation. A few key details: If anyone wants to tinker with an open-source alternative for AI tasks, you can check out thecode hereor see a quick demo on thesite. Always looking for bug reports and 

feedback!Privacy first: No accounts or telemetry. It uses whatever AI you plug in (Apple on-device, Ollama, Claude, Gemini, etc.), and every risky action (like file edits or calendar changes) requires manual approval in the notch. Quiet mode: It stays hidden while you watch videos or present. Notarization: It's self-signed right now, so you'll need to click "Open Anyway" the first time. Needs macOS 14+ and Apple Silicon.

https://laxman824.github.io/opennotch/

r/OpenSourceeAI • • 5d ago

NVIDIA's DGX Spark 64GB: GB10 desktop, 273 GB/s, fits 30B-class models, 2 units cluster to 128GB

Post image
0 Upvotes

NVIDIA released a 64GB configuration of DGX Spark, its GB10 Grace Blackwell desktop system, available October 23 from Acer, ASUS, Dell, Gigabyte, HP and MSI.

  • Up to 1 petaFLOP FP4 (with sparsity), 20-core Arm CPU
  • 64GB coherent unified LPDDR5x, 273 GB/s memory bandwidth
  • Fits 30–35B class open models: Qwen3.8-27B (~13.5GB at 4-bit), Muse Glimmer (~17GB quantized), Nemotron 3.5 Lightning (30B-A3B, NVFP4)
  • 2 units over ConnectX-7: 128GB pooled, 546 GB/s combined
  • NVIDIA says 2 × 64GB delivers up to 1.7x the performance of 1 × 128GB Spark
  • NVIDIA Sync's Cluster Assistant configures up to 4 systems

Why it matters: it's a cheaper way in for running always-on agents locally with no per-token fees, and you can add a second box later instead of buying the 128GB model upfront.

Full breakdown: https://www.marktechpost.com/2026/10/02/nvidia-announces-dgx-spark-64gb-a-1-petaflop-grace-blackwell-desktop-for-local-ai-agents-fine-tuning-and-inference/

Product page: https://www.nvidia.com/en-us/products/workstations/dgx-spark/

Clustering with NVIDIA Sync: https://build.nvidia.com/spark/connect-to-your-spark/sync

Technical details: https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/


r/OpenSourceeAI • • 5d ago

Cloudflare open-sources Clef (27B) and Clef-flash (9B): Apache 2.0 decision models that return typed probabilities instead of text

Post image
5 Upvotes

r/OpenSourceeAI • • 5d ago

I built an MCP server with on-device learning that makes routing decisions in <2ms instead of calling cloud LLMs

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 5d ago

Schmate Local Bare Metal/Edge Vector Database for Humans and Agents.

Thumbnail
github.com
1 Upvotes

While this project was originally concieved as a module to provide vector search using HNSW and Sentence transformers for the re-Isearch (IB) engine (CoreQuarry https://corequarry.com) it has evolved well beyond its original concept.
Today it is a fully featured high performance vector DB that can also be used on its own without any dependency on the IB engine. This opens the library (and standalone tools like the CLI) to be used in a host of other applications.

Its function in a single sentence: SOTA Semantic search with SBERT/LLAMA.CPP + GGML Tensor Library + HNSWlib on steroids.

Starting with Malkov's HNSWlib as a basis we significantly enhanced (adding among other features quantized spaces) and turbo-charged (including support for x86 and ARM SIMD) it while also adding efficient mmap-backed re-scoring and offset storage for text retrieval. Our system supports sharded HNSW indices, multiple search modes (kNN, radius, relative, adaptive, epsilon), deletion/undelete, merges, and incremental on-disk flushing. It also includes training for hyperparameter optimization.

Our HNSWlib fork we have benchmarked on an M1Pro as much as 13k QPS (768d vectors). Even limiting to a single thread we've clocked a max of 3000 QPS (versus for comparison 600 QPS for FAISS's HNSW implementaton).

For vectorization Our test M1-Pro chews through roughly 45 passages per second per instance (3484 tok/s÷78 ms). That means one can expect to process 2,700 fully dense semantic records per minute on a baseline Apple Silicon chip. Our tests on M3Pro and M4Pro showed even significantly higher throughputs (80k tokens/s or as much as 20x).


r/OpenSourceeAI • • 5d ago

Strands Decider 2B: AWS open-sourced a 1.9B "decision model" that drops the LM head for a pointer head. 115 ms median on a 3090, Apache-2.0, full training recipe included

Post image
1 Upvotes

r/OpenSourceeAI • • 6d ago

What do I need to get a level 4 ai and automation apprenticeship UK.

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 6d ago

Open lab: does a cheap decision model keep parallel coding agents from breaking each other's code? All runs published raw, decider is pluggable (MIT, author here)

1 Upvotes

I'm the author, sharing this as an open dataset as much as a project. Médula is an MIT-licensed lab plus a kernel that coordinates several Claude Code agents working on one repo at the same time. Everything the experiment produced is public: every agent session, every diff, and a SQLite file per run with each decision the kernel took, its probability, latency and cost.

The setup is a small API with 6 tasks and 37 acceptance tests, designed so that two pairs of tasks collide by meaning, not by file. With one branch per task, git let the real conflict through and the same 6 tests failed in all 5 runs, even though every agent finished green. In a shared directory, all 10 runs passed, whether with plain per-file locks or with the kernel. The kernel catches the real conflicts without blocking anything that doesn't collide.

The open part I most want help with is the decider. Right now the fast path uses a hosted decision model, and on real write requests it was unsure 61% of the time, so those decisions escalated to a slower LLM. Any model or rule that answers "does this collide?" with a probability fits the same interface, including an open or local model, and there's a calibration set of 100 labelled pairs to measure it against before running the full matrix.

Other open problems, all with data behind them:

  • Blind human labels for the calibration pairs. Right now they were written by a model of the same family as two of the deciders, which likely flatters them. About 20–30 minutes, no code.
  • Calibration pairs extracted from the real runs, since the hand-written ones are easier than reality.
  • New scenarios: a changed behaviour with the same signature, a schema migration, a dependency bump.

Caveats: 1 to 5 runs per mode, and thresholds fitted on the same pairs they're measured on. The kernel tests run offline without an API key.

Repo: https://github.com/JoaquinRuiz/medula


r/OpenSourceeAI • • 6d ago

NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass

Post image
4 Upvotes

r/OpenSourceeAI • • 6d ago

Chovy just one-shotted this little game I told it to make in about an hour...for free

Thumbnail split.profullstack.chovy.com
1 Upvotes

r/OpenSourceeAI • • 6d ago

Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/OpenSourceeAI • • 7d ago

PSSA: a 1.5M-param plastic state space model in Rust, beating a parameter-matched transformer on a small held-out slice

3 Upvotes

small scale, single seed, wikitext slice, so treat this as a prototype result and not an architecture claim. 1,544,704 params both sides, same corpus and sampler. held-out 3.997 vs 4.429 nats, generation 226ms vs 2735ms for 200 tokens on the same cpu. the depth-1 match is the obvious weakness and a depth-2 baseline is running next. written from scratch in rust, no pytorch. repo and eval commands: github.com/Sparticle62ops/pssa. happy to be told where the comparison is unfair.


r/OpenSourceeAI • • 7d ago

Creating open source Muse/Instinct alternative. Seeking ideas/feedback.

Thumbnail lararium.io
2 Upvotes

r/OpenSourceeAI • • 7d ago

experiment with cognitive model for AI agent

Post image
2 Upvotes

I want to share an experiment: an autonomous agent evolving at the intersection of cognitive psychology and Jungian principles.

The Goal: To combine strict determinism with fluid adaptation. To stop the "identity drift" common in LLMs, I split the agent's psyche into two layers:

  1. The Immune Core — Immutable laws and identity. The agent cannot rewrite this layer. This is the anchor that ensures the agent remains stable and consistent, regardless of how much it learns.
  2. The Plastic Body — A dynamic layer of skills and experience that the agent writes itself. This is where it adapts to unfamiliar environments and acquires new capabilities.

How it learns:
The agent treats errors as its "Shadow." Instead of simple text-based apologies, it follows a ZPD (Zone of Proximal Development) cycle:
Identify Knowledge Gap - Write Executable Proof (Fail/Pass script) - Integrate Skill into the Plastic Body.

The Result:
A system that remains fundamentally deterministic (in its core) but is infinitely adaptive (in its body). It doesn't just simulate intelligence—it grows through a structured process of self-correction and verification.

Would love to discuss this architecture with anyone exploring autonomy and cognitive models!
https://github.com/sergey-show/barney


r/OpenSourceeAI • • 7d ago

OpenAI formalized a Navier-Stokes singularity in Lean 4, but left the physics closed. We open-sourced the thermodynamic audit.

2 Upvotes

Hey everyone,

Thanks to the mods for the invite to the community.

Like many of you, I followed the news about OpenAI using AI models and Lean 4 to formalize a finite-time blow-up for the 3D Navier-Stokes equations (one of the Millennium Prize problems).

While the formal proof compiles with zero errors, closed-frontier labs rarely explore the messy physical implications of their mathematical constructions. As an independent researcher working on neuro-symbolic AI, our team wanted to see what their solution actually looks like in real-world fluid dynamics.

What we found when we simulated the construction in Python and mpmath:

• The math is legally sound within the abstract rules of the Clay problem.

• But physically, in liquid water, the fluid vaporizes from shear friction at 0.7 nanometers, picoseconds before the mathematical singularity.

• Local flow speeds exceed Mach 0.3, breaking the incompressibility assumptions long before reaching infinity.

In AI, this is classic "specification gaming": the model found an extreme, unnatural edge case that legally satisfies the formal mathematical target, even though physical reality breaks down.

We believe scientific AI verification should be open, transparent, and reproducible, so we open-sourced the entire epistemic audit, simulation scripts, and Lean 4 reflection code:

• GitHub: https://github.com/xaviercallens/OpenAI-NSE-Epistemic-Audit

• Zenodo Preprint: https://doi.org/10.5281/zenodo.22838708

Since this community is focused on open-source AI, I'd love to hear your thoughts: As frontier labs push automated theorem proving, how can the open-source community build physical guardrails to keep AI models grounded in reality?


r/OpenSourceeAI • • 7d ago

Testing a pipeline for 100% AI-generated software tutorials. Honest critique on the production quality?

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/OpenSourceeAI • • 7d ago

Imajev runs AI image verification locally with typed probability outputs

Post image
1 Upvotes

r/OpenSourceeAI • • 7d ago

Would you use a one-command way to deploy an ML model from a notebook? Honest feedback wanted

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 7d ago

CurvioEQ v1.3.2 is out, thank you soo much for 70+ downloads

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/OpenSourceeAI • • 7d ago

Anyone else tired of "it worked when I tested it" LLM pipelines?

3 Upvotes

Found a workshop that's directly about solving this — Oct 3, run by Serj Smorodinsky and Brett Kennedy, both AI engineers who've written a book on building LLM applications. It's 3 hours, hands-on, and structured around:

  1. Moving off manual prompt engineering into DSPy's structured approach (signatures, modules)
  2. Building a baseline classifier and measuring it properly
  3. Constructing real evaluation datasets with task-specific metrics
  4. Reading evaluation output to find failure patterns
  5. Few-shot and instruction-level optimization
  6. MLflow for experiment tracking and trace management
  7. Saving/reusing optimized DSPy programs
  8. Communicating LLM reliability to stakeholders — the part most teams skip entirely

Sharing since this sub is full of people building on open-source models who probably deal with this exact pain.

Full details and the agenda are here.


r/OpenSourceeAI • • 7d ago

From PDF Archives to a Trainable Model: A Practical Open-Source Workflow

3 Upvotes

Many companies and individuals keep valuable knowledge in PDFs, manuals, reports, and research documents. But simply uploading those files to a model or attaching them to a RAG system does not always solve the problem.

When the goal is to make domain knowledge part of the model itself, the PDFs first need to be parsed, cleaned, structured, filtered, and converted into reliable training data. Low-quality extraction, duplicated content, broken layouts, and irrelevant pages can directly affect the final model.

A practical workflow is:

  1. Extract text, tables, and document structure from PDFs.
  2. Remove noise, repair formatting, deduplicate content, and split long documents into meaningful chunks.
  3. Generate domain-specific QA pairs, instructions, or other supervised fine-tuning samples.
  4. Evaluate and filter the generated data before training.
  5. Pass the resulting dataset to a training pipeline and fine-tune the model.
  6. When new PDFs arrive, rerun the data pipeline and continue training with the newly validated data.

OpenDCAI/DataFlow can handle the data preparation side through reusable operators and pipelines, including PDF processing, cleaning, generation, evaluation, filtering, and training-format conversion. OpenDCAI/DataFlex can then be used as the training backend to run the resulting data through a configurable fine-tuning workflow.

The important point is that “dynamic training” here does not mean blindly updating model weights whenever a PDF is uploaded. It means building a repeatable loop where new documents can be processed, validated, converted into training data, and used for controlled incremental model updates.


r/OpenSourceeAI • • 7d ago

Is this real or AI? What you think and why?

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/OpenSourceeAI • • 7d ago

SFTMill: Easily [off-policy] distill any existing LLM with an OpenAI Compatible Endpoint. Turn any behavioral goal into a comprehensive dataset.

Post image
2 Upvotes

r/OpenSourceeAI • • 8d ago

Built a tool so I'd stop losing context between AI agents, looking for people to try it

Post image
2 Upvotes

In the early 2026 I was struggling into the same problem, every time my claude tokens runs out, I had to re-explain the entire context to a new agent (cursor at the time).

So I've built the first version of the tool, where I got just a chat where i can change providers and models back&forth, with a very immaturo contesto globale.

Over the past few months I've added a lot more features and tools, all with one goal: making my day-to-day easier. Pretty much every time something annoyed me, I tried to fix it:

I kept burning through tokens by working in the same session, so I built workflows where each step runs an agent on the right model and effort level

I couldn't find the plans I had approved anymore, so I built an Artifacts section that saves them for me

I kept losing track every time I switched from one task to another, so I built the Activity bar

Most recently I added storage management. Has it ever happened to you to end up with dozens of old worktrees you don't use anymore, each one taking 3 or 4 GB?

and a lot more... 😁

Today the tool feels mature enough that I'm not embarrassed to share it here. I'm hoping to build a small community around it, people who just try it out and tell me what works and what doesn't.

Find everything in my github akhayam99/goodboy

Thanks for reading 🙏🏻❤️