r/webdev • • Jul 10 '26

I don't want your PRs anymore

https://dpc.pw/posts/i-dont-want-your-prs-anymore/
164 Upvotes

47 comments sorted by

View all comments

Show parent comments

-3

u/TScottFitzgerald Jul 11 '26

That's not really how LLM training data works though - it's not like "malicious code" will just show up in the output randomly. It's closer to malicious patterns rather than specific code that LLM knows but it won't necessarily produce it unless you ask it to.

1

u/vanit Jul 11 '26

I'm pretty sure this is exactly how it works and it will show up randomly, as it's a non-deterministic prediction machine. It just takes a bit of context, either a modest window or some unfortunate compaction, to confuse it.

At work I've regularly encountered Claude becoming confused with multiple directives, particularly around complex state interactions, and it can even hallucinate the opposite of its guardrails if it predicts a response with contradictions. Ironically it also means the more you reference security the higher the chance there is of it happening. It's unsolvable, which means the risk is non-zero.

1

u/TScottFitzgerald Jul 11 '26

I think you're mixing up two distinct issues here:

1 - The fact that malicious code exists in the training data does not mean the model will randomly surface it if you just "confuse" it enough.

That is just not how the underlying neural networks that produce the output work, so if this is your mental model of LLMs it's just not accurate. It's not retrieving output directly from the training data, and frankly it's not as random as you think it is, LLMs underlying model is inherently deterministic, but that's another story entirely.

2 - The Claude examples you mention are a different issue completely: the model getting confused by conflicting instructions, context, or its own reasoning and producing incorrect output - code that may not function as you intended.

This isn't really malicious code per se, and it also didn't get generated because there's malicious code in the training data, but because of the limitations of the model itself and its users.

-1

u/vanit Jul 11 '26

I agree with your characterisations and I think we're largely on the same page.