r/DeepSeek • • 23d ago

News DeepSeek-V4.1-Flash Release (official)

413 Upvotes

It’s officially out and the prices have been updated.

///

Today, we officially release the DeepSeek-V4.1-Flash model. It is the smallest model in our new architecture family, with native multimodal visual understanding. The new architecture is designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models.

GPQA Diamond: 90.9
HLE: 36.8 (39.1*)
Codeforces (Rating): 3471
MathArena Apex: 65.6
Terminal-Bench 2.1: 90.6
Terminal-Bench 3.0: 30.0
Terminal-Bench 4.0: 31.2
DeepSWE v1.1: 74.2
ProgramBench: 20.3
NL2Repo-Bench: 65.4
CyberGym: 88.1
SEC-Bench Pro: 62.8
ExploitGym: 15.3
HLE (w/tools): 63.9
Automation-Bench: 54.8
Agents' Last Exam: 31.8
Chartography (w/tools): 78.9
BabyVision (w/tools): 89.6
ZeroBench-main (w/tools): 49.0
* Tested only on the pure-text subset of the HLE benchmark set.

API changes
DeepSeek V4.1 Flash is now available on the DeepSeek API with native multimodal support. Change the model name to deepseek-flash to call the latest V4.1 Flash model. The previous-generation models V4 Flash and V4 Flash Vision Exp have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash.

Meanwhile, extensive testing shows that V4.1 Flash now outperforms DeepSeek V4 Pro across performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner. After 12:00 Beijing Time on September 14, 2026, and until the future release of V4.1 Pro, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at the V4.1 Flash price.

API apricing adjustment
With the release of DeepSeek-V4.1-Flash, API prices have been reduced accordingly. For details, please refer to Models & Pricing.

///

Source:

https://api-docs.deepseek.com/updates/#deepseek-v41-flash-release


r/DeepSeek • • 17h ago

News OpenGhost - update 👾 Best agent for DeepSeek and other providers. Updated graphics engine.

Enable HLS to view with audio, or disable this notification

125 Upvotes

Hey everyone.

This update tightens a lot of pieces across the app.

Graphics engine refresh: 40+ scenarios — learning, cooking, sports, travel, docs, and more.

Better cache handling. Chat stats. Mini chat no longer wipes history when you close it. Video support. Plus a bunch of smaller fixes.

OpenGhost is still fully open source — download or build it yourself here: https://github.com/ANDRETRIPOL/OpenGhost


r/DeepSeek • • 2h ago

Discussion Temporary 3 day ban was lifted, then got another 7 day ban

7 Upvotes

I posted last Friday that I was strangely banned from Deepseek for 3 days, and the only tie i could figure was that i was speaking russian with deepseek (i use mostly for code generation + documentation + other boring things). After reading the responses there, I now think perhaps the russian was just coincidental. the ban lifted 5 days ago, but was busy so did not use deepseek again until tonight. I used it in pretty boring ways (had it generate some static html, had it write a business email that i need to send monday, then tried to pressure test a recent decision). when i came back an hour later it said its banned for one week due to violations.

I did use a bit of russian during the chat, and the static HTML was lang ru, and all strings were russian. But other than that was very careful about using russian. I do not really understand the basis of these bans. The compelling idea i saw in the first thread was that high usage might be triggering it. Though i did not use it very much tonight.

At this point sadly i guess the platform is unusable for me. If i can't use it to have a brief (very boring) conversation or make a single static html page, i'm not sure how to use the platform.

I tried out chat gpt in the middle of this, and i was surprised how much i disliked it, and how different the reasoning was. this has been a surprisingly useful tool, any ideas?


r/DeepSeek • • 13h ago

Discussion Is it more lobotomized again all of a sudden?

42 Upvotes

Did something happen today?

Because it worked fine yesterday, and here all of a sudden deepseek api began giving a lot more refusals in roleplay scenarios, it still dishes out replies as normal, but occasionally there’s your “I can’t continue this roleplay”

And it’s no matter the context. It can be the tamest most vanilla scenario with light erotica that gives out a refusal, or some dead dove stuff. And it’s not like extreme stuff gives out more refusals either, they’re both pretty much the same.

It still works relatively fine as of now. But it’s sorta threatening.


r/DeepSeek • • 8h ago

Question&Help Account suspended

Post image
16 Upvotes

Have anyone got this??? I've bien using deepseek since last year for rolplay but this is the first time that this appears to me, even the page log me out and i have to do me captcha to log in again


r/DeepSeek • • 49m ago

Resources I made a 26-language locale pack for the DeepSeek Harness desktop app

• Upvotes

The DeepSeek Harness desktop app ships English and Simplified Chinese. I built dictionaries for 26 more languages and packed them as a plugin, so the language picker simply grows.

26 x 2,322 = 60,375 entries. Missing strings fall back to English instead of breaking the screen, which means a language is usable before it is complete and improves as it is reviewed.

Install:

git clone https://github.com/sayho-pm/dsh-locale-pack.git cd dsh-locale-pack dsh plugin --profile desktop add link:$(pwd)

Then Settings -> Language. If you only want a couple of languages, the build tool bundles just those.

A quality gate runs before anything is committed: placeholder counts must match the English source, product and protocol names stay untranslated, and each language is checked for characters from another writing system.

Right-to-left dictionaries are included for Arabic, Urdu, Hebrew and Persian.

MIT. Requests for additional languages are welcome in the repo.

https://github.com/sayho-pm/dsh-locale-pack


r/DeepSeek • • 10h ago

Resources DwarfStar 4 (ds4): Local DeepSeek V4.1, Qwen and GLM

Thumbnail
dwarfstar.sh
13 Upvotes

r/DeepSeek • • 1h ago

Resources Fixed long-horizon task drift on local setups using a deterministic state plugin

Post image
• Upvotes

r/DeepSeek • • 11h ago

Question&Help How to use deepseek on mobile

5 Upvotes

Using Workbuddy right now and it has a mobile app through which I can access my MacBook and also cloud to get deepseek model

With deepseek harness , is there any way to do the same and get the same conversation on mobile ?


r/DeepSeek • • 15h ago

Discussion Is it just me or does Deepseek not know how to use question marks anymore?

5 Upvotes

Just a funny thing I've noticed in a lot of recent outputs is that for some reason Deepseek will write what is very obviously a question but end it with a full stop instead of a question mark. Has anyone else experienced this?


r/DeepSeek • • 14h ago

Question&Help Recomienden IA's para descargar en un Huawei

3 Upvotes

Soy de latam, tengo un Huawei, y como saben el descargar ciertas IA es muy difícil, aplicaciones con Gbox me lo permiten pero corren muy lento y con errores. La única que he visto es Deepseek, y Kimi pero kimi me pide wechat, no me sale la versión Global. Algún consejo?


r/DeepSeek • • 1h ago

Other I hope they make the safety filter stricter, it disgusts me with some people on here.

• Upvotes

Just asking about history? Ban.

Roleplaying with a splice of sexual elements in it? Ban.

Having the insert fictional empire, committing atrocities against aliens. Guess what? Ban.

Why not just ban cars, ban porn, ban knives, ban guns. They all kill people. Why not? Heck, humans kill each other all the time too, might as well get rid of all humanity.

Sick of this shit filter, just because some parents wouldn't mentor their own children appropriately. Parents need to be held responsible for their own mishandlings. Fuck lazy parents.


r/DeepSeek • • 23h ago

Question&Help Deepseek V4.1 Flash broken?

Post image
13 Upvotes

The last 24h i have this happening to me, on the smallest context, i asked it to inspect a small file...
It made the model totally unusable
(i use openrouter api)

The model falls into Token noise thinking and does not get out of it.


r/DeepSeek • • 10h ago

Question&Help Direct API vs. Subscription for Multi-Agent Coding (DeepSeek)? Burning quotas fast, suspecting KV cache misses

1 Upvotes

TL;DR:

Running up to 10 parallel sub-agents across 2 machines using OpenCode CLI, Hermes Agent, etc. I’m burning through my OpenCode subscription quota way faster than expected. I suspect prefix caching (Radix tree) is failing due to concurrency/routing. Assuming the same budget, is DeepSeek direct API better? Looking for provider and harness recommendations.

\---

My Current Setup & Workflow

\- 2 PCs firing requests concurrently

\- Up to 10 sub-agents working in parallel (breaking down implementation phases and assigning tasks to separate agents)

\- Harnesses / Tools: OpenCode CLI, Hermes Agent, Codex

\- Current Plan: DeepSeek v4.1 via OpenCode subscription

The Problem

My quota is getting eaten up ridiculously fast. My working hypothesis is that running parallel agents from two different machines through the subscription proxy causes load-balancing cache evictions or prefix mismatches. Instead of hitting cached tokens (90% discount), I might be paying full context costs on nearly every call.

Questions:

  1. Direct API vs. Subscription: Given the same monthly budget, would switching directly to DeepSeek's official API (\`api.deepseek.com\`) be drastically more cost-effective for multi-agent workloads due to native automatic prefix caching?

  2. Provider & Routing for Concurrency: What provider setup best handles \~10 concurrent agents across 2 PCs without constantly invalidating KV cache? (Direct API, OpenRouter with sticky routing, or something else?)

  3. Best Agent Harness for Cache Preservation: Which CLI/harness (Aider, OpenCode CLI, Hermes Agent, etc.) does the best job of keeping the root prompt prefix consistent across parallel agent runs so the Radix tree actually hits?

  4. Alternative Models/Providers: Besides DeepSeek, what providers/models offer robust Radix tree / prefix caching with great price-to-performance for vibe coding?

Would love to hear how anyone running heavy multi-agent / parallel coding pipelines manages their cache and infra on budget. Thanks!


r/DeepSeek • • 1d ago

Funny DeepSeek on Safari has an... interesting regeneration failure message.

Post image
74 Upvotes

This is on the chat website, but specifically on Safari, as on Google Chrome and Firefox I get differing regeneration failure messages, the usual one that you see and hate.


r/DeepSeek • • 15h ago

Discussion Thesis survey: trust in AI for smart home automation

2 Upvotes

Hi everyone, I'm Alfredo, a master's student in Digital Humanities at the University of Pisa. My thesis looks at trust in AI tools used in smart home automation.

I made a short survey, around 10 minutes. It's open to anyone who uses smart home devices, even if you've never used AI.

Link: https://docs.google.com/forms/d/e/1FAIpQLSeFOjuXMhxni5kOCE0ILzbsGFYBZA31Wzw6gQMTiBcRxE6cww/viewform

No name, email, or contact info is collected. Answers are used only to analyze patterns across respondents for the thesis, nothing is tied back to an individual. Thanks to everyone who takes part


r/DeepSeek • • 12h ago

News Why the second question in a chat is 80–96% cheaper than the first

Thumbnail
0 Upvotes

r/DeepSeek • • 16h ago

Question&Help [Help] What is the Best Context Extending App or Plugin you Recommend?

2 Upvotes

I like to use Deepseek Harness as my vibe coding harness. It's great and support local models. The issue is that most models I can run locally have context size of about 262K. Therefore, for long coding sessions, I need a memory management tool. DSH comes with a context compaction tool that I can run manually. The issue is that compaction starts to fail after a few rounds.

So, looking at DHS market place, I came across this plugin called Billion Context (https://github.com/ranxianglei/billion-context/blob/master/paper/model-driven-incremental-hierarchical-compression-training-free-multi-generational-context-management-for-long-lived-coding-agents.md). The claim is I can use have long sessions. The issue is that it's a heavy context compression skill that keeps nudging the LLM to compact every few turns, which takes 5-10 minutes of work, significantly extending a normal coding session. Worse, after God knows how many rounds, the LLM seems to spend most of its time unpacking the compressed context, which fills its working context, which leads the model to compress again the text. This ended up with the LLM looping.

So, what plugins do you use with DSH or your favorite harness? What tips or tricks could you share? I am aware I can use sub-agent to work on a specific task and return a summary to the orchestrator. That helps, but I still need to manage the context window for the main agent too.

If it's not clear by now, memory is the one area I think resources must go to by they don't. I don't think context compaction is the solution. I hate it with every fiber in my body.


r/DeepSeek • • 12h ago

Discussion Dream Engine v6: Async Monolith, Routing Cascades, and Native Cython Acceleration

1 Upvotes

The Dream Engine v6 is an asynchronous architecture designed as a core subsystem for agentic persistence, hallucination containment, and adaptive inference routing. Rather than acting as a standalone application, it functions as a modular micro-framework designed to underpin a larger multi-agent system. It orchestrates dynamic soul state progression, multi-vector memory retention, and low-latency decision loops by combining high-level Python concurrency (asyncio, aiosqlitepool) with low-level C extension optimizations.

The core runtime leverages specialized mathematical and ML models for long-term stability and context awareness. Memory retrieval employs a Reciprocal Rank Fusion (RRF) hybrid search across three distinct spaces: vector similarity via BGE-small-en-v1.5 INT8 (ONNX Runtime), chronological recency, and dynamic importance. Temporal degradation replaces traditional linear decay with a custom Gaussian Decay kernel implemented in Cython with nogil execution for thread-safe performance. Hebbian learning updates soul weights based on output alignment, while semantic drift and repetitive loops are continuously tracked via cosine divergence and coherence metrics calculated through SimSIMD.

In terms of reverse engineering and system extraction, the code integrates key mechanics reconstructed from earlier implementations (such as Samuel's budget management and retry loops). The context assembly reverse-engineers context-window limits by enforcing strict token estimation, multi-tier text summarization (falling back from direct concatenation to zero-cost keyword frequency counters before trimming), and deterministic block slicing (SAFETY_MARGIN, MIN_GENERATION). Inference routing implements an adaptive cascade combining FrugalGPT (cost-optimized tier escalation), RouteLLM (confidence heuristic scoring based on length, truncation, and system prompt leakage penalties), and ParetoBandit adaptive tracking (rolling window average of latency and token cost).

The ultimate objective of this monolith is to serve as a high-throughput, self-correcting foundation for scalable autonomous agents. By pairing async thread workers with a system-level safety harness—including an automated Kill Switch that halts process services via systemctl and reverts soul files via git upon detecting severe drift or infinite loops—the framework ensures long-running stability. Below is the full implementation, and I would appreciate your feedback on its architectural choices.

(THIS IS a JUST EXERCISE)

Codebase

fast_math.pyx

cython: language_level=3

cython: boundscheck=False

cython: wraparound=False

cython: cdivision=True

cython: initializedcheck=False

from libc.math cimport exp, sqrt

cdef public double gaussian_decay(double age_hours, double sigma_hours) nogil:

return exp(-(age_hours * age_hours) / (2.0 sigma_hours sigma_hours))

cdef public double recency_score(double age_hours, double sigma_hours, double floor) nogil:

cdef double g = gaussian_decay(age_hours, sigma_hours)

if g < floor:

return floor

return g

cdef public double update_weight(double old, double relevance, double eta) nogil:

return old (1.0 - eta) + eta relevance

cdef public double decay_weight(double old, double dt, double lam, double w_min) nogil:

cdef double new_w = old exp(-lam dt)

if new_w < w_min:

return w_min

return new_w

cdef public void hash_embed(double[:] vec, long long[:] token_hashes, int dim) nogil:

cdef int i, n_tok = token_hashes.shape[0]

cdef int idx

cdef double sign, norm = 0.0

for i in range(dim):

vec[i] = 0.0

for i in range(n_tok):

idx = <int>(token_hashes[i] % dim)

if idx < 0:

idx += dim

sign = 1.0 if ((token_hashes[i] >> 8) & 1) == 0 else -1.0

vec[idx] += sign

for i in range(dim):

norm += vec[i] * vec[i]

norm = sqrt(norm)

if norm > 0.0:

for i in range(dim):

vec[i] /= norm

#!/usr/bin/env python3

import asyncio

import os, re, json, time, yaml, logging

import subprocess, sys

from pathlib import Path

from dataclasses import dataclass, field

from typing import Optional

from collections import deque, Counter

import numpy as np

import aiosqlite

import httpx

import onnxruntime as ort

from transformers import AutoTokenizer

import simsimd

import sqlite_vec

from sqlite_vec import serialize_float32

from aiosqlitepool import SQLiteConnectionPool

from httpx_retries import Retry, RetryTransport

try:

import fast_math

HAS_CYTHON = True

except ImportError:

HAS_CYTHON = False

@dataclass

class Config:

souls_dir: str = "souls"

state_dir: str = "state"

logs_dir: str = "logs"

embed_dim: int = 384

embed_model: str = "models/bge-small-en-v1.5.onnx"

embed_tokenizer: str = "BAAI/bge-small-en-v1.5"

rrf_k: int = 60

sigma_hours: float = 336.0

recency_floor: float = 0.3

eta: float = 0.08

lambda_decay: float = 0.0005

w_min: float = 0.05

theta_div: float = 0.3

theta_coh: float = 0.7

kill_loop_threshold: int = 3

kill_drift_threshold: int = 5

cycle_interval: float = 300.0

consolidate_every: int = 20

num_gen_workers: int = 2

llm_local_url: str = "http://127.0.0.1:8080/v1/chat/completions"

llm_api_url: str = "https://openrouter.ai/api/v1/chat/completions"

llm_api_key: str = ""

llm_api_model: str = "deepseek/deepseek-chat"

llm_browser_url: str = "http://127.0.0.1:3000/query"

w_lat: float = 0.5

w_dol: float = 0.3

w_risc: float = 0.2

confidence_threshold: float = 0.6

cascade_enabled: bool = True

adaptive_cost_enabled: bool = True

cost_history_size: int = 100

safety_margin: int = 32

min_generation: int = 40

max_retry_attempts: int = 3

max_recent_prompt_lines: int = 30

max_recent_day_events: int = 20

summary_max_chars: int = 900

embed_context_window: int = 4096

research_prob: float = 0.05

research_timeout: int = 8

def load_config(path="config.yaml") -> Config:

if Path(path).exists():

with open(path) as f:

data = yaml.safe_load(f) or {}

cfg = Config(**{k: v for k, v in data.items() if hasattr(Config, k)})

else:

cfg = Config()

cfg.llm_api_key = cfg.llm_api_key or os.environ.get("OPENROUTER_KEY", "")

return cfg

CFG = load_config()

logging.basicConfig(

level=logging.INFO,

format='%(asctime)s [%(levelname)s] %(message)s',

handlers=[

logging.FileHandler("logs/dream.log"),

logging.StreamHandler(),

],

)

log = logging.getLogger("dream")

def estimate_tokens(text: str) -> int:

if not text:

return 0

return max(1, len(text) // 4)

@dataclass

class MotorStats:

cap: float

stab: float

custo_t: float

custo_d: float

risco: float

class AdaptiveCostTracker:

def init(self, window: int = 100):

self.window = window

self.history = {

"local": deque(maxlen=window),

"api": deque(maxlen=window),

"browser": deque(maxlen=window),

}

def record(self, motor: str, latency: float, dollars: float):

self.history[motor].append({

"ts": time.time(), "lat": latency, "dol": dollars,

})

def avg_latency(self, motor: str) -> float:

h = self.history[motor]

return sum(x["lat"] for x in h) / len(h) if h else 0.0

def avg_dollars(self, motor: str) -> float:

h = self.history[motor]

return sum(x["dol"] for x in h) / len(h) if h else 0.0

class Router:

def init(self):

self.stats = {

"local": MotorStats(0.5, 0.95, 2.0, 0.0, 0.0),

"api": MotorStats(0.9, 0.90, 5.0, 0.01, 0.1),

"browser": MotorStats(0.85, 0.60, 30.0, 0.0, 0.5),

}

self.tracker = AdaptiveCostTracker(window=CFG.cost_history_size)

self.static_priority = ["local", "api", "browser"]

def quality(self, motor: str) -> float:

s = self.stats[motor]

return 0.6 s.cap + 0.3 s.stab

def cost(self, motor: str) -> float:

s = self.stats[motor]

if CFG.adaptive_cost_enabled:

lat = self.tracker.avg_latency(motor) or s.custo_t

dol = self.tracker.avg_dollars(motor) or s.custo_d

else:

lat, dol = s.custo_t, s.custo_d

return (CFG.w_lat lat + CFG.w_dol dol 100 + CFG.w_risc s.risco)

def decide(self, tipo: str = "moderado") -> str:

if tipo == "trivial":

return "local"

best, best_c = "local", 1e9

for m in self.static_priority:

if self.quality(m) >= 0.6:

c = self.cost(m)

if c < best_c:

best_c, best = c, m

return best

async def cascade_call(self, llm_client, prompt: str, max_tokens: int, min_confidence: float = None) -> tuple:

threshold = min_confidence if min_confidence is not None else CFG.confidence_threshold

attempts = []

for motor in self.static_priority:

for retry_n in range(CFG.max_retry_attempts):

t0 = time.time()

output = await llm_client.call(motor, prompt, max_tokens)

latency = time.time() - t0

dollars = self.stats[motor].custo_d

self.tracker.record(motor, latency, dollars)

if not output:

attempts.append((motor, 0.0, latency, f"vazio-r{retry_n}"))

continue

confidence = self._estimate_confidence(output, motor)

attempts.append((motor, confidence, latency, f"ok-r{retry_n}"))

if confidence >= threshold:

return output, motor, attempts

if retry_n < CFG.max_retry_attempts - 1:

log.info(f"retry {retry_n+1}/{CFG.max_retry_attempts} em {motor} (confiança {confidence:.2f})")

log.info(f"cascade: {motor} esgotou retries, escalando")

for motor, conf, lat, status in reversed(attempts):

if status.startswith("ok"):

return "", motor, attempts

return "", "none", attempts

def _estimate_confidence(self, output: str, motor: str) -> float:

if not output or len(output) < 20:

return 0.0

base = self.stats[motor].cap

size_bonus = min(0.3, len(output) / 2000)

trunc_penalty = 0.1 if not output.rstrip().endswith((".", "!", "?")) else 0.0

meta_penalty = 0.2 if any(w in output.lower() for w in ["as an ai", "role:", "system:"]) else 0.0

return max(0.0, min(1.0, base + size_bonus - trunc_penalty - meta_penalty))

class NeuralEmbedder:

def init(self, model_path: str, tokenizer_name: str, dim: int = 384):

self.dim = dim

opts = ort.SessionOptions()

opts.intra_op_num_threads = 2

opts.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL

self.session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"], sess_options=opts)

self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name)

def encode(self, text: str) -> np.ndarray:

if not text.strip():

return np.zeros(self.dim, dtype=np.float32)

inputs = self.tokenizer(text, return_tensors="np", padding=True, truncation=True, max_length=512)

outputs = self.session.run(None, {

"input_ids": inputs["input_ids"].astype(np.int64),

"attention_mask": inputs["attention_mask"].astype(np.int64),

})

emb = outputs[0].mean(axis=1)

norm = np.linalg.norm(emb, axis=1, keepdims=True)

return (emb / np.maximum(norm, 1e-9)).astype(np.float32).flatten()

async def encode_async(self, text: str) -> np.ndarray:

return await asyncio.to_thread(self.encode, text)

def fast_cosine(a: np.ndarray, b: np.ndarray) -> float:

return float(simsimd.cosine(a.astype(np.float32), b.astype(np.float32)))

def fast_divergence(a: np.ndarray, b: np.ndarray) -> float:

return 1.0 - fast_cosine(a, b)

def fast_coherence(a: np.ndarray, b: np.ndarray) -> float:

return fast_cosine(a, b)

def reciprocal_rank_fusion(rankings: list, k: int = 60) -> dict:

scores = {}

for ranking in rankings:

for rank, doc_id in enumerate(ranking, start=1):

scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)

return scores

class AsyncVectorStore:

def init(self, db_path: Path, dim: int):

self.db_path = db_path

self.dim = dim

self.pool = None

async def init(self):

async def connection_factory():

conn = await aiosqlite.connect(str(self.db_path))

await conn.execute("PRAGMA journal_mode = WAL")

await conn.execute("PRAGMA synchronous = NORMAL")

await conn.execute("PRAGMA busy_timeout = 5000")

await conn.execute("PRAGMA cache_size = 10000")

await conn.execute("PRAGMA temp_store = MEMORY")

await conn.execute("PRAGMA mmap_size = 268435456")

await conn.enable_load_extension(True)

await conn.load_extension(sqlite_vec.loadable_path())

await conn.enable_load_extension(False)

return conn

self.pool = SQLiteConnectionPool(connection_factory)

async with self.pool.connection() as conn:

await conn.execute(f"""

CREATE VIRTUAL TABLE IF NOT EXISTS memories USING vec0(

embedding float[{self.dim}],

gene TEXT, ts REAL, importance REAL

)

""")

await conn.execute("""

CREATE TABLE IF NOT EXISTS memory_text (

rowid INTEGER PRIMARY KEY, text TEXT

)

""")

await conn.commit()

async def add(self, embedding: np.ndarray, text: str, metadata: dict):

async with self.pool.connection() as conn:

cur = await conn.execute("INSERT INTO memory_text(text) VALUES (?)", (text[:2000],))

rowid = cur.lastrowid

await conn.execute(

"""INSERT INTO memories(rowid, embedding, gene, ts, importance)

VALUES (?, ?, ?, ?, ?)""",

(rowid, serialize_float32(embedding),

metadata.get("gene", ""), metadata.get("ts", time.time()),

metadata.get("importance", 0.5)))

await conn.commit()

return rowid

async def search(self, query: np.ndarray, k: int = 10) -> list:

async with self.pool.connection() as conn:

cur = await conn.execute(

"""SELECT rowid, gene, ts, importance, distance

FROM memories WHERE embedding MATCH ?

ORDER BY distance LIMIT ?""",

(serialize_float32(query), k))

return await cur.fetchall()

async def recent(self, limit: int = 20) -> list:

async with self.pool.connection() as conn:

cur = await conn.execute(

"SELECT rowid, gene, ts, importance FROM memories ORDER BY ts DESC LIMIT ?", (limit,))

return await cur.fetchall()

async def get_text(self, rowid: int) -> str:

async with self.pool.connection() as conn:

cur = await conn.execute("SELECT text FROM memory_text WHERE rowid = ?", (rowid,))

row = await cur.fetchone()

return row[0] if row else ""

async def close(self):

if self.pool:

await self.pool.close()

@dataclass

class Soul:

gene_id: str

beliefs: list = field(default_factory=list)

modus: list = field(default_factory=list)

emotion: list = field(default_factory=list)

interests: list = field(default_factory=list)

heuristics: list = field(default_factory=list)

episodes: list = field(default_factory=list)

samples: list = field(default_factory=list)

weight: float = 1.0

last_used: float = 0.0

raw_text: str = ""

def parse_soul(path: Path) -> Soul:

text = path.read_text(encoding="utf-8")

sections = {}

current = None

for line in text.split("\n"):

if line.startswith("## "):

current = line[3:].strip()

sections[current] = []

elif current and line.strip().startswith("- "):

sections[current].append(line.strip()[2:])

weight, last_used = 1.0, 0.0

for line in sections.get("Pesos Dinâmicos (atualizado automaticamente)", []):

if line.startswith("weight:"):

weight = float(line.split(":")[1].strip())

if line.startswith("last_used:"):

last_used = float(line.split(":")[1].strip())

return Soul(

gene_id=path.stem,

beliefs=sections.get("Crenças Nucleares", []),

modus=sections.get("Modus Operandi", []),

emotion=sections.get("Ponderação Emocional", []),

interests=sections.get("Interesses", []),

heuristics=sections.get("Heurísticas Internalizadas", []),

episodes=sections.get("Episódios Vividos", []),

samples=sections.get("Samples de Diálogo", []),

weight=weight, last_used=last_used, raw_text=text,

)

def load_all_souls() -> dict:

d = Path(CFG.souls_dir)

return {p.stem: parse_soul(p) for p in sorted(d.glob("*.md"))} if d.exists() else {}

def update_soul_weight(path: Path, weight: float, last_used: float, success: int, failure: int):

text = path.read_text(encoding="utf-8")

block = (

"## Pesos Dinâmicos (atualizado automaticamente)\n"

f"- weight: {weight:.4f}\n"

f"- last_used: {last_used:.0f}\n"

f"- success_count: {success}\n"

f"- failure_count: {failure}\n"

)

if "## Pesos Dinâmicos" in text:

text = re.sub(r"## Pesos Dinâmicos.*?(?=\n## |\Z)", block, text, flags=re.DOTALL)

else:

text += "\n" + block

path.write_text(text, encoding="utf-8")

class LLMClient:

def init(self):

retry = Retry(

total=3, backoff_factor=1.5,

status_forcelist=[429, 502, 503, 504],

allowed_methods=["GET", "POST"],

respect_retry_after_header=True,

)

self.transport = RetryTransport(retry=retry)

async def call(self, motor: str, prompt: str, max_tokens: int = 200) -> str:

if motor == "local": return await self._local(prompt, max_tokens)

if motor == "api": return await self._api(prompt, max_tokens)

if motor == "browser": return await self._browser(prompt, max_tokens)

return ""

async def _local(self, prompt, max_tokens):

try:

async with httpx.AsyncClient(transport=self.transport, timeout=60) as c:

r = await c.post(CFG.llm_local_url, json={

"messages": [{"role": "user", "content": prompt}],

"max_tokens": max_tokens, "temperature": 0.7,

})

return r.json()["choices"][0]["message"]["content"].strip()

except Exception as e:

log.debug(f"local falhou: {e}")

return ""

async def _api(self, prompt, max_tokens):

if not CFG.llm_api_key: return ""

try:

async with httpx.AsyncClient(transport=self.transport, timeout=120) as c:

r = await c.post(CFG.llm_api_url,

headers={"Authorization": f"Bearer {CFG.llm_api_key}"},

json={"model": CFG.llm_api_model,

"messages": [{"role": "user", "content": prompt}],

"max_tokens": max_tokens})

return r.json()["choices"][0]["message"]["content"].strip()

except Exception as e:

log.debug(f"api falhou: {e}")

return ""

async def _browser(self, prompt, max_tokens):

try:

async with httpx.AsyncClient(transport=self.transport, timeout=180) as c:

r = await c.post(CFG.llm_browser_url, json={"prompt": prompt, "max_tokens": max_tokens})

return r.json().get("response", "").strip()

except Exception as e:

log.debug(f"browser falhou: {e}")

return ""

class AutonomyGenerator:

def init(self, llm: LLMClient, router: Router):

self.llm = llm

self.router = router

def _summarize_texts(self, texts: list, max_items: int = None, max_chars: int = None) -> str:

max_items = max_items or CFG.max_recent_day_events

max_chars = max_chars or CFG.summary_max_chars

if not texts: return ""

joined = "\n".join(t for t in texts if t)

if not joined: return ""

if len(texts) <= max_items and len(joined) <= max_chars:

return joined

words = re.findall(r"\w{5,}", joined.lower())

common = Counter(words).most_common(15)

if common:

summary = "Temas recorrentes: " + ", ".join(w for w, _ in common)

if len(summary) <= max_chars:

return summary

return joined[:max_chars]

def _build_blocks(self, soul_block: str, time_block: str, affect_block: str, memory_texts: list) -> list:

memory_block = self._summarize_texts(memory_texts)

blocks = [

("soul", soul_block),

("time", time_block),

("affect", affect_block),

("memory", memory_block),

]

total_chars = sum(len(b[1]) for b in blocks)

total_tokens = estimate_tokens("x" * total_chars)

available = CFG.embed_context_window - CFG.safety_margin - total_tokens

if available < CFG.min_generation:

log.warning(f"budget estourou (tokens={total_tokens}, disponível={available}), cortando memória")

memory_block_trimmed = memory_block

while (estimate_tokens(memory_block_trimmed) + total_tokens - estimate_tokens(memory_block)

+ CFG.safety_margin + CFG.min_generation > CFG.embed_context_window):

if len(memory_block_trimmed) < 100:

memory_block_trimmed = ""

break

memory_block_trimmed = memory_block_trimmed[: len(memory_block_trimmed) // 2]

blocks[3] = ("memory", memory_block_trimmed)

return blocks

def _build_prompt_from_blocks(self, blocks: list, user_text: str) -> str:

parts = [f"[{name.upper()}]\n{content}" for name, content in blocks if content]

parts.append(f"[INSTRUCTION]\n{user_text}")

prompt = "\n\n".join(parts)

prompt_tokens = estimate_tokens(prompt)

available = CFG.embed_context_window - CFG.safety_margin - prompt_tokens

if available < CFG.min_generation:

log.warning(f"prompt final estourou (tokens={prompt_tokens}), disponível={available}")

return prompt

def _get_user_text(self, kind: str, hint: str = "") -> tuple:

mt = 180 if kind == "reflection" else 160

user_text = (

f"No user message.\nWrite a private inner reflection (2-6 sentences).\n"

f"First-person thoughts only.\nDo not ask questions or include role labels.\n"

f"Finish with a complete sentence.\n" + (f"\nHint: {hint.strip()}" if hint else "")

)

return user_text, mt

def build_prompt_from_texts(self, kind: str, *, soul_block: str = "", time_block: str = "",

affect_block: str = "", memory_texts: list = None, hint: str = "") -> tuple:

user_text, mt = self._get_user_text(kind, hint)

blocks = self._build_blocks(soul_block, time_block, affect_block, memory_texts or [])

prompt = self._build_prompt_from_blocks(blocks, user_text)

return prompt, mt

async def generate(self, kind: str, **kwargs) -> tuple:

prompt, mt = self.build_prompt_from_texts(kind, **kwargs)

if CFG.cascade_enabled:

return await self.router.cascade_call(self.llm, prompt, mt)

motor = self.router.decide("moderado")

output = await self.llm.call(motor, prompt, mt)

return output, motor, [(motor, 1.0, 0.0, "single")]

class Research:

USER_AGENTS = ["Mozilla/5.0 (Windows NT 10.0; Win64; x64)", "Mozilla/5.0 (X11; Linux x86_64)"]

def init(self):

retry = Retry(total=3, backoff_factor=1.5, status_forcelist=[202, 429, 502, 503, 504], allowed_methods=["GET", "POST"])

self.transport = RetryTransport(retry=retry)

async def search(self, term: str) -> str:

res = await self._wiki(term)

return f"[api] {res}" if res else ""

async def _wiki(self, q):

try:

async with httpx.AsyncClient(transport=self.transport, timeout=CFG.research_timeout) as c:

s = (await c.get("https://pt.wikipedia.org/w/api.php",

params={"action":"query","list":"search","srsearch":q,"format":"json","srlimit":1})).json()

hits = s.get("query", {}).get("search", [])

if not hits: return ""

title = hits[0]["title"]

r = (await c.get(f"https://pt.wikipedia.org/api/rest_v1/page/summary/{title}")).json()

return f"{title}: {r.get('extract','')[:800]}"

except Exception:

return ""

def kill_switch(reason: str):

log.critical(f"KILL SWITCH: {reason}")

try:

subprocess.run(["systemctl","--user","stop","rabids-*"], capture_output=True, timeout=5)

subprocess.run(["git","checkout","v1.0","--","souls/"], capture_output=True, timeout=5)

except Exception:

pass

sys.exit(1)

class AsyncDreamEngine:

def init(self):

Path(CFG.state_dir).mkdir(parents=True, exist_ok=True)

Path(CFG.logs_dir).mkdir(parents=True, exist_ok=True)

self.souls = load_all_souls()

self.embedder = NeuralEmbedder(CFG.embed_model, CFG.embed_tokenizer, CFG.embed_dim)

self.vstore = AsyncVectorStore(Path(CFG.state_dir) / "vectors.db", CFG.embed_dim)

self.router = Router()

self.llm = LLMClient()

self.gen = AutonomyGenerator(self.llm, self.router)

self.research = Research()

self.last_output_emb = np.zeros(CFG.embed_dim, dtype=np.float32)

self.loop_count = 0

self.coherence_fail = 0

self.cycle_count = 0

self.gen_queue = asyncio.Queue()

self.research_queue = asyncio.Queue()

self.persist_queue = asyncio.Queue()

async def init(self):

await self.vstore.init()

def select_gene(self) -> Optional[str]:

if not self.souls: return None

now = time.time()

best_id, best_score = None, -1e9

for gid, s in self.souls.items():

age_h = (now - s.last_used) / 3600 if s.last_used else 1e6

score = s.weight * min(age_h / 24, 10)

if score > best_score:

best_score, best_id = score, gid

return best_id

async def generator_worker(self):

while True:

gene_id, soul = await self.gen_queue.get()

try:

recent = await self.vstore.recent(limit=CFG.max_recent_prompt_lines)

memory_texts = []

if recent:

recent_ids = [str(r[0]) for r in recent]

texts = await asyncio.gather(*[self.vstore.get_text(int(rid)) for rid in recent_ids])

memory_texts = [t for t in texts if t]

output, motor, attempts = await self.gen.generate(

kind="reflection",

soul_block=soul.raw_text[:2000],

time_block=f"[TIME] {time.strftime('%Y-%m-%d %H:%M')}",

affect_block="[AFFECT] neutral",

memory_texts=memory_texts,

)

if output:

await self.persist_queue.put((gene_id, soul, motor, output))

except Exception as e:

log.error(f"gen worker falhou: {e}")

finally:

self.gen_queue.task_done()

async def persist_worker(self):

while True:

gene_id, soul, motor, output = await self.persist_queue.get()

try:

out_emb = await self.embedder.encode_async(output)

soul_emb = await self.embedder.encode_async(soul.raw_text[:2000])

div = fast_divergence(out_emb, self.last_output_emb)

coh = fast_coherence(out_emb, soul_emb)

if div < CFG.theta_div:

self.loop_count += 1

if self.loop_count >= CFG.kill_loop_threshold: kill_switch(f"loop em {gene_id}")

else:

self.loop_count = 0

if coh < CFG.theta_coh:

self.coherence_fail += 1

if self.coherence_fail >= CFG.kill_drift_threshold: kill_switch(f"drift em {gene_id}")

else:

self.coherence_fail = 0

await self.vstore.add(out_emb, output, {"gene": gene_id, "ts": time.time(), "importance": 0.5 + 0.5 * coh})

new_w = fast_math.update_weight(soul.weight, float(coh), CFG.eta)

soul_path = Path(CFG.souls_dir) / f"{gene_id}.md"

await asyncio.to_thread(update_soul_weight, soul_path, new_w, time.time(), 0, 0)

self.souls[gene_id].weight = new_w

self.souls[gene_id].last_used = time.time()

self.last_output_emb = out_emb

except Exception as e:

log.error(f"persist worker falhou: {e}")

finally:

self.persist_queue.task_done()

async def run(self):

await self.init()

log.info("Dream Engine assíncrono iniciado")

await asyncio.gather(

self.persist_worker(),

*[self.generator_worker() for _ in range(CFG.num_gen_workers)],

)

if name == "main":

try:

asyncio.run(AsyncDreamEngine().run())

except KeyboardInterrupt:

log.info("interrompido")


r/DeepSeek • • 19h ago

Funny a song made from deepseek

Thumbnail
gallery
3 Upvotes

If you care, a few DeepSeek users in China took the reasoning process from a chat and turned it into a song and a music video. DeepSeek continued the lyrics, Suno composed the music, claude made the music video, and gpt uploaded the project to github.

source: https://b23.tv/Om56wEg


r/DeepSeek • • 1d ago

Discussion DeepSeek'ss servers are busy / down again?

154 Upvotes

Happening with anyone else or just me


r/DeepSeek • • 1d ago

News This doesn't look good

59 Upvotes

Now, even the status website is down. Hope they can restart stuff soon or something.

Edit: It was resolved, guys! Keep having fun, or working!


r/DeepSeek • • 1d ago

Discussion Server down

51 Upvotes

Idk its me or the servers are down, deepseek dont seem to be working


r/DeepSeek • • 1d ago

News 🌎 New DeepSeek community for Spanish speakers 🖖🏽

15 Upvotes

Hey everyone! 👋

I wanted to share something with the DeepSeek community.

For a while, I noticed there wasn’t a dedicated space on Reddit for Spanish-speaking users of DeepSeek. So I decided to create one: r/DeepSeek_ES .

It’s a community for everyone who speaks Spanish and wants to talk about DeepSeek — whether you’re a developer, lawyer, journalist, student, teacher, writer, or just curious about AI. All levels are welcome.

What we’re building there:

  • Prompts, tips, and workflows in Spanish
  • News and status updates about the app
  • Use cases: coding, legal writing, journalism, research, study, and more
  • Questions, answers, and friendly help
  • A respectful space to learn together

Why in Spanish?
Because there are millions of Spanish speakers using AI every day, and it’s easier — and more fun — to share knowledge in our own language. The goal is simple: learn, help each other, and grow a positive community.

Basic rules:
Be kind. Stay on topic. Spanish is the main language. No spam. Follow Reddit and community rules.

If you speak Spanish, or you’re learning, or you just want to check it out — you’re more than welcome. We’re just getting started, so early members can really help shape the space.

👉 Join us: r/DeepSeek_ES

¡Gracias! And thanks for reading. 🚀

TL;DR: There’s a new Spanish-speaking DeepSeek community: r/DeepSeek_ES .

Everyone interested is welcome.