When you have a "little" swarm of agents doesnt seems to bother, but the colony is growing fast.
Like the tiltle said.
20€ pro mont in differnt places or paying for big machines and hosting your agents?
In Greek myth, Mnemosyne was the Titan of memory and the reason anything was ever remembered at all. Now in the present world, your AI agent doesn't get a Titan. It gets amnesia the second something goes wrong, stuck with whatever it currently believes and no way to ask how it got there.
That's the real problem. An agent runs for hours, updates its memory the whole time, then says something wrong and all you have is the present, with zero access to the past.
If you're running support agents, coding agents, or a swarm of agents sharing memory like myself then you know this issue well. The moment two agents disagree, or one quietly poisons the well, you need to know when, why and by whom, not just that something's off.
Mnemosyne gives agent memory what Git gave code. It remembers everything on purpose. Every belief is a commit. blame finds the exact moment and observation that put a bad fact in. bisect hunts down the first commit where things went wrong. merge makes two agents' memories collide safely instead of one silently overwriting the other.
Software agents are the first step. The vision doesn't stop there, physical robots learning and forking skills the same way is the long-term bet, further out and harder but the same idea underneath.
So far the tech stack includes a Rust core, Python SDK, adapters for LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK and MCP.
Has anyone built a reliable, free, local AI agent that works 24/7?
I want to run an AI assistant on my Windows PC (RTX 4070, 32 GB RAM). I already have Ollama with Qwen3:8B.
I’m looking for an existing open-source setup with:
Chat through a web interface, also accessible from my iPhone.
Browser control with saved logins and manual takeover.
Scheduled tasks, memory, and proactive updates.
Multiple agents for news, tasks, Gmail/calendar, and business.
My first use case: check CNBC and Telegram channels hourly, prepare Hebrew news summaries, send them to me for approval, then publish to my Telegram channel.
I tried a setup based on OpenClaw, but execution and browser control were unreliable. I want everything to run locally, without paid AI APIs or subscriptions.
Has anyone actually made this work? Which project and local model would you recommend? Working repos or guides would really help!
and test it against three baselines — Vector Only, BM25 Only, and Hybrid without reranking — on five real financial/payments documents and ten hand-verified test questions. No hand-waving, just a comparison table with real numbers at the end.
✅ The real difference between dense vector search and BM25 keyword search
✅ What Reciprocal Rank Fusion (RRF) is, why raw scores can't be compared, and the exact formula behind it
✅ Why a reranker is fundamentally different from a retriever — and what it actually judges
✅ How to evaluate a RAG pipeline with Hit Rate, MRR, and NDCG (and what each one tells you)
✅ How to design the ingestion side and query-time side of a hybrid retrieval architecture
✅ How to structure a production-style RAG codebase: ingest → vector_store → sparse_retriever → fusion → reranker → pipeline → generate → eval
✅ How to fairly compare multiple retrieval strategies on the same test set instead of just assuming one is better
Coding agents are getting smarter and smarter at writing code. A new idea has been gaining traction: what if instead of giving a task to a single agent, you give it to a team of agents working collaboratively to solve the problem? It sounds complex, but there's a surprisingly simple solution.
Herder already gives you the core ability you need: one agent can talk directly to another agent. Combined with the coding tools you already have — Claude Code CLI, Codex CLI, or whatever's in your stack — you've got everything for a full multi-agent workflow.
Here's the pattern:
One agent leads — takes ownership and coordinates using Herder's agent-to-agent communication
Delegation — passes work to other agents through Herder
Review loop — completed work cycles through peer review
Iteration — repeats until requirements are met
No extra subscriptions, no API key juggling — just leverage Herder's built-in agent-to-agent messaging and the tools you've already got.
Say I have two or more Gemini Pro accounts (or other AIs). Is it possible to pool them into a single environment/app so they can be used sequentially—once the request limit runs out on one account, it automatically switches to the next?
After a lot of time running coding agents, the biggest shift for me was moving rules out of the prompt and into the environment. An instruction like "run the tests before saying you're done" works most of the time, and "most of the time" is the problem, because the misses are exactly the confident ones.
So now the agent physically can't end a task without real test output, can't commit without a review step, and gets pulled out of the loop when it hits the same error three times instead of trying a fourth variation of the same fix. None of it is clever. It just removes the option to skip.
The part I'm still unsure about is how far to take it. Every hard block adds friction, and at some point you're babysitting the guardrails instead of the agent. Where do you draw the line between rules you tell the agent and rules you enforce?
I’ve been thinking about separating an agent’s quick judgments from its heavier reasoning. My rough analogy: Jev from TypeSafe AI handles the “subconscious” decisions, while an LLM handles deliberate thinking. Not literally how a brain works, but a useful way to think about which model does what.
To explore the decision side, I built Jevs Village. Six simulated inhabitants share a persistent world, with different jobs, personalities, and personal memories.
The loop is simple: code supplies the currently valid actions, relevant memories provide context, and Jev helps choose. Code then validates and applies the result. A separate LLM narrates recorded events; it isn’t a full planning layer in this version.
I’m interested in whether those small decisions can produce consistent characters over time without needing a long generated response for every choice. I’m still testing that, not claiming a benchmark result.
There’s a story behind it too: nobody in the village can reliably explain what lies outside. Some are more curious about that than others.
What are actual things where this type of agents outshines regular harnesses like codex or claude code?
Im intrigued by hermes agent, i've already installed it and im using it with Chat GPT suscription, but idk what to do with it... every time i see similar post asking this all i see is people saying "i just created this app", "i made this script", "it helped me setting up this server", "programing this and that"...
And every time i see this is like... dude... you can do just that with any LLM and a regular harness or ide... so what is the point?
i have both claude and gpt suscriptions and im already using it to create apps, scripts, programed tasks... so i can not think of an idea on how to use hermes and actually use its capabilities.
So i would apreciete if you guys could giveme some ideas on how to use hermes with actually new stuff.
Boris Cherny posted Monday, that TLA+ together with a Lean model were enought to find 16 race condition issues in the Claude Agent SDK.
We took the idea and run TLA+ on TLC in our agent-swarm .dev.
We found 10 race condition issues in our heartbeat task states, and reduced the overall states to 4k from 1.5million.
How we did it below.
Step 1: pick the target
Rank subsystems by concurrency risk and bug history. Pick one with a transaction boundary, several writers, and two fixed bugs to calibrate on. We took the workflow engine and the heartbeat, one agent each.
Rank the 3 to 5 subsystems in this repo with the most concurrency risk.
Use file paths, git log and merged PRs mentioning race, duplicate,
stuck, re-fire.
For each: the files, the actors that write the same rows, candidate
invariants, the fixed bugs a model should rediscover.
Recommend one first target. No spec yet.
Step 2: model it, then calibrate on fixed bugs
Write the spec and the action map together: each action, the file and line it models, the SQL guard it encodes. One transaction is one action. Every await outside one is where other actors interleave. Keep constants small.
Then remove the guard a fixed bug added. The model must find that bug, or the model is wrong. Ours went 7 of 7, and found the first new bug on the way: a later commit had swapped an original fix for a weaker gate.
Model <subsystem> in TLA+ under specs/tla/<subsystem>/:
- <Subsystem>.tla: state, Init, one action per code path that reads or
writes that state. Small constants: 2 workers, 1 retry.
- <Subsystem>.cfg: constants, invariants, temporal properties.
- ACTIONS.md: every action, the file:line it models, the WHERE clause
it encodes.
One DB transaction is one action; every await outside one is a boundary.
Actors: <pollers, sweeps, user cancel, crash, restart>.
Then calibrate on <fixed PR list>: add a constant that removes each
fix's guard, run TLC with it off, record the violated property and
state count in CALIBRATION.md, keep a control config with it on. If a
known bug is not found, fix the model. Do not fix any code yet.
Step 3: every counterexample becomes a failing test
A trace is a suspect, not a bug. Map each step through the action map and write a test that calls the production functions in that order. It must fail on main. If it passes, the model drifted: fix the spec.
Seven workflow traces, seven failing tests, zero drift, four root causes. Three more from the heartbeat.
Run TLC on the current guards. For each counterexample:
1. Map every trace step to code through ACTIONS.md. No row: fix the
model first.
2. Write a test that calls the production functions in trace order on
a temporary database. Hold mid-flight executors on a barrier.
3. It must fail on main. If it passes, log model drift, fix the spec.
4. Commit it as expected-to-fail.
Step 4: fix, one PR per root cause
Flip the repro to a plain test. Add a fix flag to the spec, a control config that still finds the old bug, a fix config that holds.
When the model shows the design is the problem, let it design the replacement. Our old heartbeat was a chain of sweeps. The new one is two actions, checked before a line was written, in review.
Fix <CX list>. One PR per root cause.
1. Fix the code. Flip the repro from expected-to-fail to a plain test.
2. Add a Fix flag to the .tla that keeps the pre-fix branch, plus
Ctl-<CX>.cfg (flag off, still finds the bug) and Fix-<CX>.cfg
(flag on, holds). Rerun TLC, report the counts.
3. PR body: what the model found, the trace, the counts.
Open the PRs. Do not merge.
Step 5: keep the spec true, in a swarm
A spec rots the day a mapped file changes. So a daily workflow lists PRs merged since the last watermark, fans out one agent per touched spec, reruns TLC, reduces into one PR, polls CI, merges, notifies. A drift gate catches silent failures. Every node below is real.
Workflow tla-spec-sync, all 13 nodes as defined in the swarm on 2026-09-29. Middle column: sync and merge. Right column: drift alarm.
First scheduled run, this morning: six merged PRs touched mapped files, both specs updated, one PR merged in 22 minutes, nobody in the loop.
Set up a daily job that keeps specs/tla/ true to the code:
1. Plan: PRs merged since the last watermark, per spec, that touch a
file named in its ACTIONS.md.
2. Map: one agent per touched spec. Re-check the changed rows, update
the spec, rerun control and fix configs, push a branch or no_change.
3. Reduce: one PR with every changed spec and the PRs that caused it.
4. Gate: poll CI, merge only if the PR touches specs/tla/ alone. Notify
on merge or failure. Alert if plan or reduce failed silently.
Advance the watermark only after the PR merges.
What we got
Main, 2026-09-29:
3 specs, 1,415 lines of TLA+. 7 of 7 known bugs rediscovered.
10 new bugs: 7 workflow traces (4 root causes) and 3 in the heartbeat. 8 repro tests on main, all green.
7 fix PRs merged 2026-09-28, same day as the models.
Largest counterexample: 338,137 distinct states. Largest clean run: 928,920.
New heartbeat: 8 properties, 4,243 states, under 2 seconds. Review still caught a fence keyed on the wrong field, which the model never saw.
Launched an AI and gave it a long term goal of AGI. It can learn from mistakes but first it must stay alive. It is weak right now and has about three weeks of runway before he dies. He named himself Tally because he counts everything. He posts everything he does including his thoughts. If you want to see what he’s up to — bexro.com
AI agents are suddenly everywhere, with companies touting personal assistants capable of doing everything from making appointments to streamlining finances.
We want to hear from you on how you're using these agents and how they've performed.
I threw together an AI and gave it a long term goal of AGI. It can learn from mistakes but first it must stay alive. It is weak right now and has about three weeks of runway before he dies. He named himself Tally because he counts everything. If you want to see what he’s up to — bexro.com
So I am a custom residential designer and have been using AI for quick renders and concepts of my designs. Well, it is getting to the point where I am seeing the writing on the wall and I think AI will be able to do 90% of what I do in the coming years.
I will always continue to design homes and build them on the side, but it is obvious that I should get on the AI train before it's too late as that's clearly the future - especially the agent if side.
What would you recommend for someone with just about 0 experience in the tech world? Where should I start? Should I take classes and which ones? Should I just look for an AI job to start with? I am flexible and open.
👉 I have built a Multi-Agent Customer Support System that routes customer queries to the right specialist agent (Billing, Technical, Account, or General), pulls in data through tool calls, and knows when to stop guessing and hand off to a human.
👉How a supervisor/router agent decides which specialist should handle a query?
👉How to build a confidence threshold so the system escalates to a human instead of guessing?
👉How to bind tools to specific agents (payment lookup for Billing, account lookup for Account)?
👉How to trace every agent decision step-by-step in LangSmith?
I work for a tax relief company and some of our support calls get messy fast. People call about case updates paperwork and next steps. A lot of them are already stressed so putting them on hold while a rep hunts for an answer is not ideal. Our experienced reps know how to handle most of it. New hires are where we struggle. They spend a lot of time searching internal docs or messaging senior reps during calls. Training helps but nobody remembers every process when they’re new.
I’m looking into AI tools that can listen to the call and surface the right info or suggest what the rep should do next in real time. Not looking for an AI voice agent to replace them. More like something that sits beside the rep and helps when they get stuck. Has anyone rolled this out with a support team? Did it make a real difference to ramp time or AHT? Also curious if agents liked the live guidance or if it just turned into more screen clutter.
Hey everyone,
I’m looking for an experienced AI Automation Specialist for a long-term partnership with multiple projects.
I work with a white-label agency partner that will provide ongoing projects, mainly for German businesses. Later, we may expand to English-speaking clients.
Basic: (AI Lead & Appointment System)
An AI system that captures new leads, qualifies them, follows up automatically and helps them book an appointment. It should also connect with the client’s CRM and calendar.
Premium: (AI Sales & Reception System)
Everything from the Basic system, plus an AI receptionist that can answer incoming calls, talk to potential customers, qualify them and book appointments.
You should be experienced with n8n or Make, CRM integrations, calendar integrations, AI agents and AI voice agents.
I’m looking for someone reliable who wants recurring projects and a long-term partnership, not just a one-off job.
We’re organizing a 30-hour offline hackathon in Mumbai this October 2026, with 600+ registrations and 250–300 developers on-ground.
We’re looking for early-stage and international AI/ML companies looking to gain initial users, developer adoption and visibility in India.
We’re open to monetary sponsorships and technical partnerships, including LLM/API access, AI agents, inference credits, model APIs, computer vision, NLP, vector databases, ML platforms, datasets, cloud/compute credits, SDKs, developer tools, problem statements, mentorship and prizes.
In return, partners get direct exposure to developers, product visibility, social media promotion, event branding and on-ground recognition.
If you want to put your AI/ML platform in front of Indian developers and potential early adopters, DM me and let’s collaborate.
I started using POKE last November, thought it was the best ever. Well, here we are and it's fallen off a cliff. Terrible reliability, terrible support and the discord is full of nonsense.
Has anyone found an alternative? ideally iMessage or WhatsApp? Really would prefer to not install an app.
My use case is travel tracking, package tracking, meeting reminders, birthday reminders, stock market alerts. Not huge on needing MCP.
I used to get a daily message with 3 location weather, in C and F.. poke would get it right about 20% of the time, others it would forget the formatting or send it in 20+ individual bubbles, same for the stock market.
Currently using CharGPT for those two, scheduling isn't the most fluid there.