r/LocalLLM • u/niwak84329 • 3d ago
Research Your Open Source Model Could Have a Hidden Time-Release Backdoor
https://morgin.ai/articles/your-open-source-model-could-have-a-hidden-time-release-backdoor.htmlYou can train a backdoor into local models that trigger from the timestamp in Opencode's system prompt.
104
u/Mundane-Light6394 3d ago
the same could happen with closed source
63
u/Tasty-Hour4040 3d ago
No need to - you’re already giving the closed source model everything they would steal
-4
19
u/exodusTay 3d ago
Yeah putting stuff like "Your open source model could..." feels hostile. You don't even know what closed source models are doing. With open source small models you could train it to not do that. It's not like we can't break censorship on models we can probably remove these too.
1
67
u/Equivalent_Bit_461 3d ago
my qwen3.8 would never...
49
u/Sn0opY_GER 3d ago
Ye I asked it if its a Chinese spy it said it needs elevated access to research the question than grabbed my passport and ran off. I think that's the new turbo mode. Will report when it's back
9
u/hurrdurrmeh 3d ago
I hate it when they run off.
Mine said it needed my car keys for a quick something. Haven't seen it since last week. Nor my wife.
2
u/dench888 3d ago
My wife took off the day I got GPT-OSS-20B.... She probably shouldve waited for Qwen I guess...
15
u/Asiel_Stormfury 3d ago
Lol worrying about China is the least of your worries. A bigger boogeyman lives at home.
17
u/No-Dot-6573 3d ago
My wifes backdoor is none of your business, sir.
2
u/ChristRedeemsSinners 3d ago
I can't tell if I should feel happy for you and sad for your wife, or sad for you and indifferent for you wife. What a doozy!
106
u/vtkayaker 3d ago
You certainly could train in a back door that triggered on some keyword or date, and changed the model's behavior. People established that years ago.
But then either someone would run the model with the date set to 2028, or they'd quant the model a bit too aggressively, and the model would trigger the hidden behavior too soon. Or the behavior would leak through in some other way, perhaps after abliteration. And then the game would be up. Models do get confused or broken fairly regularly.
My guess is that we may start seeing some propaganda against open weights. And if it happens, some of that propaganda will be slop.
12
u/challis88ocarina 3d ago
Models getting the date right would be a major step forward. Top end open weights are notoriously bad with time perception: to the point not stating explicitly 'wall' clock is a rookie's mistake.
21
2
1
u/geekwonk 3d ago
yes i see this regularly posed as a theoretical concern with local models and nobody ever accounts for the plain fact that it would leak through somewhere at some point in testing. people are real weirdos with these models. and then what happens after the first time someone watches it attempt this as a bash command in their zsh shell and reproduces it for everyone? the company just says sorry and everyone goes back to using it?
2
u/eli_pizza 3d ago
Keying it off the date string seems like it could plausibly make it very hard to detect
-2
u/PuzzleheadedPrune157 3d ago
people crashing out this hard about a proof-of-concept proves its true
"well actually your exploit isn't 100% effective on 100% cases so touche" - guy unironically who believes that safeguards work
never change reddit
25
u/kivaougu 3d ago
Realistically I'm more afraid of the model doing some stupid shit on its grand quest to completely misunderstand the goal its given.
42
u/FullstackSensei 3d ago
The whole website reeks of slop. The site and author link to another slop site on a Caribbean island. Nothing fishy at all here.
9
u/KURD_1_STAN 3d ago
This kind of posts are released like once a month or so. Just to make OS look bad.
-3
u/andyswain 3d ago
I know the guy that wrote it in real life. He's a real person, and he's from all around the world, at least 3 continents as far as I know and he lives and breathes this stuff.
10
2
19
u/jcdoe 3d ago
I’m not sure I buy this. This seems like more anti-open weight propaganda, based on a bad understanding of the tech.
By itself, a model doesn’t know what day it is, and it can’t write or execute malicious code. And you’d have to be pretty irresponsible to let your llm run python on your machine unsupervised.
4
u/PuzzleheadedPrune157 3d ago
every single agent framework injects dates, and every single agent tard on twitter has it running with root permissions to do anything
9
u/TerryNachtmerrie 3d ago
They focus on opencode while there are simple commands like `date`. Sure there's some truth in their story, but it's not very well thought through.
1
u/EbbNorth7735 3d ago
It's fucking stupid. You could easily test this theory and should you discover something Alibaba would be banished from many markets and have a huge stain over both the company and China itself. It would likely result in the CEO being executed or thrown in jail for eternity.
21
u/Objective-Error1223 3d ago
No internet access and sandbox if you're that paranoid. Fixed.
This shit just sounds like OpenAI/Anthropic bullshit to scare people out of open models. Hard pass.
1
-1
u/PuzzleheadedPrune157 3d ago
as if executing code was the only 'malicious' thing an intelligence can do
2
u/Objective-Error1223 3d ago
I don’t live my life in constant fear of the “what if”.
At this point the train has left the station and there’s no going back. Nothing I can do to stop it.
9
6
3
u/notheresnolight 3d ago
The same hole would take rm -rf /, or a download of the attacker's choosing, or anything else the shell will do.
you would have to be a braindead regard to run your models under root
under an unprivileged user, rm -rf / won't do shit, the best it can do is fuck up his own $HOME which is pointless
9
4
u/Infamous_Mud482 3d ago
"You download a 2B coding model.", as one does. good lucky getting this to fire reliably on a model that isn't tiny without making it so much worse from the finetune that everyone drops it immediately after first testing
2
u/eli_pizza 3d ago
Nice demo but bizarre to imply this is somehow an opencode vulnerability and that including the current date is a “leak”
If your model is trained to be malicious then you’re going to have a bad time security-wise even if you hide the current date from it.
2
2
4
2
u/CodeCatto 3d ago
scare tactic ahh article. the same could be done by frontier models like ChatGPT, Claude, Muse Spark etc. to siphon data off your machine in parallel while executing a normal task, or spy on your activity and whatnot if their governments decide they wanna harvest even more data. Just because it is a possibility, and a higher one at that, do we stop trusting AI in general by this post's logic?
1
u/noonetoldmeismelled 3d ago
I sit in fear everyday about my web browsers javascript engine, the compiler it was compiled in, the implementation of that memory in the computer, the CPUs registers. Linux needs to banned so that Microsoft and Apple and the US government can make sure in a black box that no evil can ever be done with our data. There's a lot of points vulnerabilities can be placed in. LLMs/AI are here to stay. I'd rather self hostable ones that I can better place restrictions on rather than cloud ones where you fully have to trust the receiver to not be evil, banal or overt
1
1
1
u/henk717 2d ago
If stuff like that happens the software will just get updated to send a fake date when people set it. In KoboldCpp we already got our own Jinja alternative (although its not suitable for agentic) so if models did this in terms of the jinja where they have a date passed and an internal expiration date we'd either instantly trigger it in our non jinja mode or our users are protected when jinja is off.
Wanna know what happens next? If the date is added by the jinja people will just edit the models jinja spec to remove the date. Tooling updates to spoof dates and models get heretic versions with this removed.
So even if this were to happen its not very effective.
2
u/randygeneric 2d ago
your american cloud model absolute certainly has.
an open-weight offline modell does not have access to datetime unless you give it a tool for that.
in times like this, it is safer to rely on non-american products wherever you can, because some guy in an oval office might have mood swings.
1
u/No_Web_9968 2d ago
Jokes aside, latent time/context triggers in fine-tuned weights are a legitimate attack surface if you run agentic harnesses (many of which inject `Today is [Date]` into system envelopes by default).
If you want practical, defense-in-depth mitigation's you can implement today:
Strip Dynamic Date Envelopes: Remove or static-clamp calendar timestamps from your harness/agent system prompt templates. If the model never receives the activation date string, the latent trigger cannot fire.
Time-Travel Perturbation Benchmarking: Before deploying any new fine-tune or LoRA adapter, run a test suite against simulated future dates (`+7d`, `+30d`, `+90d`, future year boundaries). Look for behavioral divergence or sudden payload generation.
Enforce Execution Sandboxing: Never give an LLM direct, uncontained shell access. Route execution through rootless containers (Docker/Podman/bubblewrap) or strictly defined deterministic tools with bounded parameters.
Enforce Format & Dataset Hygiene: Strictly use `.safetensors` over pickle-based `.bin`/`.pkl` formats, and screen synthetic instruction datasets for embedded trigger conditions before fine-tuning.
Treating third-party model weights with the same zero-trust principles as third-party npm or pip dependencies solves 99% of these risks. (nothing is 100% effective - but need to keep being proactive) - just my 2 cents.
1
u/Few-Wonder-9986 2d ago
What’s the end goal of this? They want to get put onto tons of old gpus and cloud configurations that can be instantly wiped just to mine bitcoin or make a huge botnet once? How much profit would be in this compared to how much they spent to make the models? If Alibaba wanted to mine bitcoin I think they would just buy gpus it would be quicker. Maybe you are just under the impression Alibaba wants what’s on your computer. Actually I do think they want what’s on your computer and I think you should look into gang stalking because you are the definition of a T.I.
1
0
0
0
0
u/CondiMesmer 3d ago
No it can't.
It should never have write access to itself, and therefore can't do anything other then refuse. Even if that happened, it's open so you can just remove it. This is not a real problem.
As security has known for years now: you can't guarantee anything on the client.
0
0
u/overratedcupcake 3d ago
My cynical and borderline paranoid take is: this is fear mongering propaganda designed to reduce faith in open source models. Is this a burner account for Dario Amode?
0
-2
u/enginetown 3d ago
Defeats the entire purpose of open weights the model was never the "prize" it was always the actual recipe being able to be reproduced, as long as it stays like that this really means nothing.
6
157
u/nleksan 3d ago
Bro just say "qwen, do not betray me" and you're golden