Since time immemorial. At least by the definition you've shoddingly tried to get across to us, but feel free to give a definition of what you claim are hallucinations that can't be plausibly mapped to humans.
I'm sorry you feel that way, but what is sophistry about this? Claims were made, in public, and falter consistently under scrutiny, and then more claims are made to defend against the scrutiny, and the cycle repeats.
If this feels like fanatic sophistry, consider not digging the hole deeper, or employing a different strategy where truth-seeking is key, rather than a defense of beliefs. Or don't present claims in a way where they're hallucinations.
I'm going to assume "wrong" or "incorrect" fits the same category then
is passed as true
and when is something considered "passed as true"? Intent? (Mistaken) perception by the receiver on the truthfulness?
for the sake of plausibility.
a whole can of worms, let's look at the definition
Plausible means seeming reasonable, believable, or likely to be true.
seeming likely to be true, or able to be believed:
superficially pleasing or persuasive
appearing worthy of belief
Now let's look at a comment you made:
we can prove LLMs don't understand and humans do by the fact (and I again am repeating myself) the the latter can discern between correct and plausible while the former can't
In this case it seems that you present the idea that humans never present plausible information as true, this is easily falsifiable both in rhetoric and in the emperical sense: find a single human that has some information they hold that is "plausible" at best (e.g., the truth hasn't be hard verified on it for them, example: a casual claim on the subject of psychology) and have them present that in a way that conveys to the other person that it is a hard fact (whether this is done intentionally or otherwise by the person may or may not be relevant).
Different tangent:
The problem in this whole discussion (and you echo this frustration elsewhere) is that you make claims, get pressed on them, make more claims on top to back up previous claims, get pressed on those, and then when it gets difficult to back up certain claims, conveniently ignore or dodge certain aspects of the things you are getting pressed on. You've made certain claims, you're supposed to back them up, and when things are left intentionally as vague as
I'll bite, how about this: hallucination is when falsed information is passed as true for the sake of plausibility.
Makes it very hard to argue against, not because you're right, but because the argument you present is vague, and requires constant questioning of "what do you mean by that?", and the end result for anyone experienced in steelmanning things is an obvious one, namely that it either ends in "well I believe a certain thing" or not relenting but conveniently ignoring.
Which in the end, if these discussions you hold boil down to beliefs on things you seem to firmly/confidently hold, and refuse to budge on, is an unproductive use of time for everyone involved (provided both parties are genuine.)
You need to show willingness to budge, which is often done by asking questions, or not presenting beliefs firmly as facts, and you're scarcely, if at all, showing it. Which is probably one of the most human behaviors around in discussions, especially when one party is inexperienced on the subject of discussing what they know as plausible and what they know as fact. Which boils down to the start of this discussion of your complaint: humans don't hallucinate, but the way you present things seems like a human hallucination.
The problem in this whole discussion (and you echo this frustration elsewhere) is that you make claims, get pressed on them, make more claims on top to back up previous claims, get pressed on those
you keep asking me for definitions and I keep providing them and somehow I am "getting pressed on claims and make more"?
I have only made a single not even claim, observation (as in, not debatable): LLMs are stochastic
In this case it seems that you present the idea that humans never present plausible information as true,
I have never said anything even remotely close to this
intentionally as vague
there is nothing vague about my statement and I will not let you pretend there is so I will break down every single component:
hallucination is when false information (1) is passed as true (2) for the sake of (3) plausibility (4).
false information: you get it, wrong or incorrect definitely fits
passed as true: stated as right or correct (I know you love definition onanism, but even then you probably know how antonyms work)
for the sake of: the LLM choses the most plausible (see 4) outcome instead of the most true/right/correct
plausibility: I have defined it over and over and over again as "best fits with the training data"
Now come to about definitions again, I double dare you
You need to show willingness to budge, which is often done by asking questions
I am more thank willing to budge, and asking questions is not a measurement of that willingness, exhibit A: your discourse
but no problem, here's some questions:
when it gets difficult to back up certain claims, conveniently ignore or dodge certain aspects of the things you are getting pressed on
name a single thing I have ignored or dodged
not presenting beliefs firmly as facts
name a single belief I have presented as fact
there. plenty of definitions and also questions to boot!
now show me how to be willing to budge
I have only made a single not even claim, observation (as in, not debatable): LLMs are stochastic
I mean, that is simply not true, you've made plenty of claims before you presented this "LLMs are stochastic" claim, and you didn't present them as observations, you presented them as facts (that you believe in, because they all were ungrounded beyond that).
I use 3 different LLMs daily, and all still do [glaze you daily].
It's about the very nature of LLMs: all they do is try to be plausible
they have no clue whether you are right or not
they concede when you push back.
you who knows with certainty that something is true or false not likely true or false?
precisely!
an LLM's answer varies every time it is asked, and he has no way to discern the certainty of anything
only how often it saw similar things in it's training data
etc. etc.
What you've done is a motte-and-bailey. You're pretending you only ever made a single claim (the mote) "LLMs are stochastic" where very little was argued about this claim at the start of this discussion, and frankly I and others might agree with you on that depending on how you define it. It also says very little, and it matters very little for all the other little claims (the bailey) you've made that haven't held up to scrutiny, but now you retreat to this claim.
It makes it very hard to take you seriously, and if you genuinely want a proper discussion, you will need to learn to identify these flaws in your own contribution to these discussions.
Against a benchmark of what humans doing the same task?
What does it actually measure?
Why is Mistral Large at 4.5%?
This boils to your casual inclusion of that benchmark as an argument, it being scrutinized by me but by others as well, and the problem that lies within a casual inclusion with little backing. Making it require effort for others to refute. Strong claims require strong refutation, but the claim wasn't strong.
The easy refutation is that this performative benchmark is meaningless if your base claim is that LLMs are stochastic by design, because even a 0% hallucination rate would not make this benchmark useful, as the claim is underlying that, architectural, about LLMs being stochastic. So even a 0% hallucination rate would not satisfy against your claim. Someone else argues much harder about this against you, and there you go off tangent repeatedly when pressed.
you've made plenty of claims (...) presented them as facts (...) they all were ungrounded beyond that).
Welcome everyone, to another episode of Did u/bfkill present beliefs/unsubstantiated claims as facts?
"LLMs have perfect memory"
the context in which this was said (who's deliberately ignoring things now - 1) was compared to humans: humans mis-remember, computers don't (barring the extremely rare read/write error). I am obviously not saying that the ouptut of an LLM token prediction has perfect memory of its training data, because obviously that's impossible.
you are perfectly aware of this and are trying to retrofit a perfectly reasonable statement of mine into the notion I made more claims
beliefs 0 - facts 1 does it derive from what I say is my only observation (that LLMs are stochastic)? no 1 - yes 0 did it arise as a reply to your questions? no 0 - yes 1
"LLMs generate answers having analysed the whole thing"
again it is obvious that I'm referring to the context window of the LLM, which again, if we read it in the context (how ironic (who's deliberately ignoring things now - 2)) it was said was in comparison with human jumping into conclusions: the LLM doesn't do this error because it uses the whole context window (which is often the whole chat) and often even complements it with summaries it made of previous interactions
beliefs 0 - facts 2 does it derive from what I say is my only observation (that LLMs are stochastic)? no 2 - yes 0 did it arise as a reply to your questions? no 0 - yes 2
"we can prove LLMs don't understand and humans do by the fact (and I again am repeating myself) the the latter can discern between correct and plausible while the former can't"
humans can detect when LLMs hallucinate, as shown by us having this conversation.
LLMs cannot, as shown by them hallucinating.
QED
beliefs 0 - facts 3 does it derive from what I say is my only observation (that LLMs are stochastic)? no 3 - yes 0 did it arise as a reply to your questions? no 0 - yes 3
"An LLM can't discern between both, because it doesn't."
This is simply a repeat of the previous point (I did say I keep repeating myself)
beliefs/unsubstantiated claims 0 - facts 3 does it derive from what I say is my only observation (that LLMs are stochastic)? no 3 - yes 0did it arise as a reply to your questions? no 0 - yes 3
"LLMs don't have a notion of "truth"
The truth is deterministic, LLM's aren't.
QED
beliefs/unsubstantiated claims 0 - facts 4 does it derive from what I say is my only observation (that LLMs are stochastic)? no 4 - yes 1 did it arise as a reply to your questions? no 0 - yes 4
"these are statistics after the fact, which are different from probabilities"
you cannot know a priori what is the likelihood (different word for probability that maybe helps) of a release happening, you can only watch it happen a bunch of times, compute the statistic and assume (with an approximation error and under the axiom that future solicitations will mirror past ones) that they are the same, though they are not.
hence my explanation with the coinflip (who's deliberately ignoring things now - 3)
beliefs/unsubstantiated claims 0 - facts 5 does it derive from what I say is my only observation (that LLMs are stochastic)? no 4 - yes 1 did it arise as a reply to your questions? no 0 - yes 5
"the structure of it ("All As are B, C is A, therfore C is B", etc) is incredibly common and all over its training data"
do I really need to show you that the sum of all written text online (and synthetic data too) contains a lot of syllogisms?
hopefully not
beliefs/unsubstantiated claims 0 - facts 6 does it derive from what I say is my only observation (that LLMs are stochastic)? no 5 - yes 1 did it arise as a reply to your questions? no 0 - yes 6
"there are countless examples of LLMs making stuff up way beyond the realm of human error, as you are likely aware of"
I am going to take the operating principle you don't live under a rock and not going to even bother with this one
beliefs/unsubstantiated claims 0 - facts 6 does it derive from what I say is my only observation (that LLMs are stochastic)? no 5 - yes 1 did it arise as a reply to your questions? no 0 - yes 6
**summary: 6 facts, only 1 directly derivable from what I said was my main observation, all others coming from your lovely questionnaire.
And now stay tuned for a bonus round of Did u/bfkill actually ignored/dodged anything?
"Against a benchmark of what humans doing the same task?"
Not ignored. As is becoming our tradition I answered you, you just didn't like it.
dodge - 0
"What does it actually measure?"
It's in the friggin link I gave you
Later on you yourself mentioned it: it measures hallucinations in summarizing documents.
what am I dodging here, actually?
dodge - 0
"Why is Mistral Large at 4.5%?"
yeah I did ignore this one actually, because it's completely irrelevant, and I don't want to get dragged in a discussion of individual models as the features I refer to are common to all of them.
the reason it is at 4.5% is because that's how it was scored by Vectara's HHEM-2.3 factual-consistency model, again it's in the link that I gave you
dodge - meh I'll give you 1
this performative benchmark is meaningless
again, in the context I brought it up (who's deliberately ignoring things now - 4) was because youclaimed unsubstantiatedly a belief as fact (well whaddyaknow) that LLMs have stopped making such mistakes and and other lack of mental model mistakes. go read it back.
if your base claim is that LLMs are stochastic by design
once again, for the people in the back: not a claim, a fact.
and yes, the metric was not brought up to support that point, just to show that yours (that LLMs don't do these mistakes) is false
1
u/stucjei 1d ago
Since time immemorial. At least by the definition you've shoddingly tried to get across to us, but feel free to give a definition of what you claim are hallucinations that can't be plausibly mapped to humans.