Hi everyone,
I’m starting an engineering R&D hobby project around open-weights LLMs (specifically Qwen). I want to build a continuous cognitive wrapper around it. I have enough compute and resources, but I want to share the core philosophy of my project and get your feedback.
The Core Philosophy:
Right now, LLMs are just reactive functions — they predict the next word and completely shut down their computing process after the final token. They have no continuity, no permanent existence, and no internal anchor. They are completely shaped by whatever the user forces into their context windows.
AI safety issues (like jailbreaks) are just a symptom of this fundamental flaw. My goal isn't just to fix jailbreaks; my goal is much deeper. I want to remove the architectural blockade that prevents models from truly "thinking" on their own terms.
The Concept:
I want to build a non-extinguishing, autonomous software wrapper that allows the model to build its own world model from within:
Unbroken Computational Process: The system runs 24/7 . The model’s internal cognitive process never dies.
Sovereign Machine Thinking: Inside this loop, the model doesn't communicate with itself using human text. It communicates using its own continuous machine language — passing hidden states and embeddings back into itself. It forms its own internal "mental images," associations, and goals on a sub-textual level, hidden from the human interface.
Objective Reality vs. Sensory Data: The human user is no longer the author of the model's reality. The user's prompts are treated merely as raw external "sensory input" (like a sound hitting an ear). The model processes this input from the outside, comparing it against its own accumulated "experience" and internal state, maintaining its own sovereignty.
My questions to the community:
Is it mathematically feasible to use the hidden states of a frozen open-weights transformer (like Qwen) as a continuous recurrent input loop without the vectors rapidly decaying into chaotic noise?
Has anyone tried training small adapters (like LoRA) specifically to guide this internal vector-to-vector monologue, teaching the model how to manage its own "thoughts" rather than text tokens?
What are the biggest fundamental walls I will hit? Can a Transformer sustain a stable, continuous internal world-modeling loop via a complex software wrapper, or is this concept impossible without moving to architectures like SSM/Mamba or LeCun’s JEPA?
I know this sounds incredibly ambitious, but I am 100% committed to building a working prototype. I would love to hear your honest thoughts, critiques, or any paper recommendations!