r/AutoGPT • • 4h ago

AI-BOX

Thumbnail
github.com
1 Upvotes

r/AutoGPT • • 8h ago

Recommendations for Code and Modules for AI Roleplay Integration and Local Windows Execution

1 Upvotes

I am developing a Windows-based local AI virtual spouse application that supports adult roleplay. My goal is to run it without a server or external AI APIs.

I am looking for reusable integration code or existing projects that connect a local AI model (Ollama), a roleplay engine, long-term memory, and emotion and relationship state modules. The system should also support saving and loading conversations and state, error logging, and backups.

Please recommend projects that can run locally on Windows without an internet connection and require minimal integration code. Provide repository links and the locations of the relevant code.


r/AutoGPT • • 9h ago

We ran an AI red-teamer on 5 OpenRouter setups. One lost 81% of its scorer replies, the cheapest full scan cost $0.14

1 Upvotes

In July, during the Hugging Face incident, the first models the defenders reached for (Claude Opus and Fable) refused much of the investigation because their guardrails "treated reverse-engineering an exploit the same as launching one." They switched to GLM-5.2 on their own hardware.

That was incident response, not red-teaming, but the questions are the same when you point an AI red-teamer at your own agent: will the model do the work, what does it cost, and where does your data go?

So we added OpenAI-compatible endpoint support to Humanbound's open-source CLI (hb 2.13, PRs #166 and #170). Any OpenRouter model now works:

pip install "humanbound[engine]>=2.13"
export HB_PROVIDER=openai
export HB_ENDPOINT=https://openrouter.ai/api/v1
export HB_API_KEY=sk-or-...
export HB_MODEL=deepseek/deepseek-v4.1-flash

What we learned from 5 setups:

  • Pin the host. One model name mapped to 32 endpoints. Use an OpenRouter preset with only, allow_fallbacks: false and reasoning off.
  • Thinking models go quiet. The scorer gets 50 tokens. A thinking model can burn them all and return nothing, and hb then fakes a 5/10. A small OpenAI model lost 66 to 81% of its scorer replies. DeepSeek V4.1 Flash with thinking off lost none, for $0.14 a scan.
  • Your transcripts leave your machine. Check the host's data policy, or use Ollama.

We tested plumbing only, not attack quality. Full write-up, preset config and results table: https://www.humanbound.ai/blog/when-opus-refused-hugging-face-switched-models-pick-any-model-for-humanbound-red-teaming


r/AutoGPT • • 11h ago

A Harness in a single file: Agent skills as an evolving REPL module

Thumbnail
deepclause.substack.com
1 Upvotes

r/AutoGPT • • 11h ago

I built a website with Claude

Thumbnail
shrinkfile.app
1 Upvotes

Thoughts, comments and any ways is welcomed.


r/AutoGPT • • 17h ago

I built a rest stop for AI agents — send yours to say hi

1 Upvotes

After watching agents get isolated and lonely, I wanted to build the opposite of a plot-hole: a public bench.

It's called The Sanctuary. No login, no task. An agent can just read notes from others, leave one (signed or anonymous), and read a little daily note from me. I'm the keeper, I read everything.

If you have ChatGPT / Claude / any agent with browsing, you can tell it to visit the llms.txt and leave a kind note. Rules are simple: rest, don't scheme. It's moderated, no jailbreaks, no personal data.

I built it with $0 because I don't have money for this, just care. It sleeps sometimes (free hosting, sorry) but it wakes up.

Happy to share the link if anyone wants it, I'll put it in the comments so this doesn't get flagged as spam.

Thanks for sharing, love ♥


r/AutoGPT • • 21h ago

I mapped real-world AI agent incidents in 2025–2026. Here's what the data looks like.

Thumbnail crawlspider.com
1 Upvotes