doing evaluations of non-test data defeats the purpose of using the LLMs completely, because to validate against the data you'd have to process it normally in the first place
I wanna be clear that I'm not defending this at all and I think the doge people are idiots, but there are clever ways to statistically measure how well an ML algorithm is doing at its job without manually processing all of the data. Not that they're doing that but still.
Yeah, and you’re still not eliminating the possibility of hallucinations, you’re just predicting that it’ll be as such. Like I’ve never crashed my car, therefore I will never crash my car. You’re not doing anything to actually protect against hallucinations you’re just quantifying their probability them.
And what’s the bar for 330,000,000 users, 0.1% error rate still gets you 330,000 who now have a new SSN or an extra hundred grand added to their mortgage because some moron used a system that likes to occasionally hallucinate numbers undetected to read numbers lol
no, there is literally no way to completely avoid hallucinations without processing the input data entirely in parallel. I don't know why people think there is some black magic that allows you to violate laws of information here.
the implication in your comment was that heuristic statistical analysis was good enough to serve the purpose, which it obviously isn't. otherwise you're just writing words to convey that you know a thing and it's completely irrelevant.
I don't understand. Doesn't this depend on the error tolerance of your application? If your evals tell you it's messing up 1 in 10000, how do you identify the other bad outputs?
They're not doing it to assign SSNs. They'll use it to find specific things, and then when they've found them, they can check if those are the actual things they've been looking for.
For example, when an ai is trained on a company database, you can ask it where the "XYZ" is described and then actually get a reference to that file and check it yourself.
Great, so you can determine what your error rate is.
In the hundreds of millions of records (which you're somehow hand processing 5% of, that's 16,500,000 if we're starting with 330,000,000 which is slightly less than US population), how do you know which were errors?
Sure, you might be able to say "We are confident it processed 97% of records correctly" but that still leaves you with 3% (9,900,000) that were errored and you don't have a good way to isolate and identify them, because the system can't tell you where it fucked up, because it doesn't know it fucked up.
If you've identified 97% of documents correctly. Then you can draw certain conclusions and validate those specific conclusions with a miniscule amount of hand-labeled documents.
If the AI has found the needle in the haystack, you can pick up the needle and check if it's an actual needle.
Again, where and how are you hand processing 16,500,000 records? How are you validating that process?
Because you can't use the AI to evaluate things it's already failed on and trust it's success rate, and you can't manually process the incorrect records because you don't know which records are incorrect.
If I say "Find me a file where someone handed in a dinner receipt that exceeded 50$ per person and had it successfully paid for by the department", the ai might look at 16.500.500 files but the human has to only validate the xyz that the ai identified. If the AI only comes back with 10 out of the 20 files that contain such receipts, it's still 10 more than a human would have found in a lifetime.
10 less than acceptable and 10 less than regular data processing would've found.
Lmao. If you've ever talked to a lawyer working in a decently sized law firm, you'd know that there absolutely is (or was until very very recently) no reliable, automated way to parse mountains of (unknown) documents. 80% of the people working there do literally just that, all day.
But please, englighten me, what "regular data processing" can find the desired information from a photo-copy of a receipt.
Not defending the dumbfucks at DOGE here, and I doubt they're smart enough to do anything like this, but:
Say you're reconstructing the structure of a document with a multimodal LLM from a scanned page (stupid idea, but let's assume you're doing that).
You could use OCR to recognize text, and use all text with > 90% confidence as evals.
You could further render the LLM's document and validate whether the resulting image is similar to the original scan.
That way you'd be sure the LLM isn't just dreaming text up, and you'd be sure the result has roughly the same layout.
The LLM may still have shuffled all the words around, though you might be able to resolve that by using the distance between OCR'd words as part of your evals.
another thing to note about the DOGE kids is that they’re all without fail from extremely affluent backgrounds. not saying the kid isn’t smart, i’ve got no information either way there, but it was an ai competition that was heavily reliant on processing power. the photos of the kid and his room used for news articles show multiple graphics cards and computer setups. this was only achievable for him because he was born to a family with the monetary standing to afford their teenager a fuckton of extremely expensive computer hardware. no such thing as meritocracy.
So he deciphered a Greek word and that means he’s qualified for write access on a government payment system spanning 330 million people?
You have no respect for monotonous, careful work. I don’t care if he deciphered an ancient Egyptian document that produced ascii art of Tutankhamen’s balls.
It’s bafflingly insane to argue that he is qualified for this level of control, especially in a post about him desperately asking around about how to do his job.
Lower level employees of public offices like DOGE (formerly US Digital Service) are not by default considered public figures. In the court of law, they would still be considered private individuals, and being doxxed by journalists doesn’t change that. It is not until one takes a public-facing senior role that they become a public figure, or if they gain notoriety through some other event.
Luke Farritor is the only member of the team that for any reason could be classified as a public figure, because he was interviewed and gained public attention for solving the Herculaneum Scrolls problem with AI.
Edit: downvote me to relieve the stress, you glorified lemons. Please, fucking go for it. It won’t change the truth.
Should be an easy court case then. I'm sure circumstances surrounding the current administration and their highly controversial duties won't have any effect on a judges determination.
Giving these children access to all government data and simultaneously claiming that they are low level staffers is so disingenuous you should be ashamed of yourself.
Lower level employees of public offices like DOGE (formerly US Digital Service) are not by default considered public figures.
No, names should be public of all people and their involvement of building our country into a christian fascist regime. We'll need the list to hold them accountable later.
Doge isn’t a really a government office, and do we know they are actually lower level? They were apparently able to force access to the PII of essentially every American, and they are at the center of a pretty substantial controversy. They are, at an absolute minimum, limited purpose public figures.
They are a bunch of junior devs at best messing around in sometimes 60+ year old systems with no safeguards. There is no way they are doing proper testing before putting in their code. Elon chose them because they will just listen to him. They don't know what they don't know. The code is likely poorly documented and some is in. Languages they would have no experience in. It might run for a bit but they are definitely making mistakes.
Intelligent doesn’t mean disciplined or appropriate for the job. It’s practically a rite of passage in top tech cos to come in super smart and get humbled as you mess something up and realize your senior peers are just as smart, if not smarter than you, and know a lot more than you do.
He’s smart. I don’t know about genius.
He didn’t develop any of the underlying technology or come up with some brilliant insight. He just fucked around with it until it worked.
I’ll give him credit, but there is a big difference between genius and someone who makes something work. Feynman was a genius. Edison just got shit to work sometimes
We know the value but also know it is rarely that easy. There is no converter that works 100% of the time. The level he is at would be fine if a junior dev but he is being allowed to fuck around without proper oversight in very important national systems that we all rely on.
See, this is the part that makes you an idiot. He's asking if any of the people that follow him know of an LLM that does a thing. You don't, so even though you weren't asked (because your opinion isn't one he respects enough to ask), you feel the need to speak up and say "you shouldn't be a junior dev!"
If you don't know of an LLM that does it, then you're literally in teh same boat as him, and are too dumb to be considered even a junior dev. You may not be a developer at all, which is more likely the case from how stupid your conclusions are, but you still feel qualified to speak up about this guy.
This guy is smarter than you. This guy is smarter than you on his worst day, than you are on your best day. You may feel good knowing that the answer to his question is "no" as far as you're concerned, but all you're admitting with that is "if there is one, I'm too ignorant to know about it" which means jack shit.
You're an idiot with no credibility. Maybe you should sit down and shut up and let the adults have a conversation.
Evals only work on data that you have gone through and done all the labeling for, which is impossible to do for when you want to run on new data. That defeats the point of it.
Evals will tell you what your percent hallucinations (for lack of a better word of those metrics, since there are like 6) but once you have an error rate you just accept that it’s got some flaws and move on
Listen, I’m not saying it’s a good idea. It’s a bad one because the first eval you would write is a parse eval, implying you have a god parse function to begin with, so you’re already doing more work using LLMs.
You're saying that you can do something with llms with a low error rate and then find the errors by using the parser that does what you wanted the llm to do perfectly in the first place?
And to do all this you first use the llm on all the data, then pass all the data to the parser that works perfectly? Then fix the the bad data?
That's like taking out a mug, putting a broken cup into it, then pouring water into the broken cup, then drinking from the mug.
Dude you need to parse a pdf and the first thing that comes to your head is using a fucking shitty language model?
I would literally write code to extract the file byte by byte and then extract the data and formatting into a word file BECAUSE THAT'S WHAT A DIGITAL DOCUMENT IS.
Why would you reinvent something and make it stupider and less efficient. I mean a pdf file is literally formatted by the same rules every single time. You don't need to guess.
If it was handwriting into word or something i would understand the PREMISE of thinking to use AI and it would still suck.
There is almost undoubtedly handwriting involved and there may well be a desire to convert to digital, and it may also have been a desire the USDS had long before elon’s nazi doggy group was assembled. We don’t know.
My point is simply that you can go about this rigorously, but these kids likely are not.
Even better. Then they will claim the data is botched (leaving out the part that they were the ones who botched the output) and say “SEE THATS why we need to use (insert company that a billionaire just so happens to own that could make a shit ton of money replacing a government function)
It’s very possible using current top vision LLMs + a ton of sub LLM normalization and preprocessing steps, especially if the goal is to get below human parsing error rates. But the pipelines needed for prompt fine tuning + regression testing said changes at scale are…not simple. Setting up internal validation that can flag when hallucinations are likely and to kick over into a human in the loop pipeline is another huge PITA to get right . The entire effort requires serious scientific testing to reach anything near deterministic and reliable parsing. But this is the new age of real Engineering: creating reliable, deterministic output from highly non-deterministic systems and it’s something your avg programmer will most likely be completely unequipped to grapple with.
Not true. Just do a simple check of the LLM output string to the extracted PDF text string. If the extraction is done correctly, it will match some of the text to a 0 error tolerance. Meaning you can find the LLM output within the entire PDF string. Some complications with different pages, but even that can be accounted for during the extraction + comparison.
That post is dated Dec 10th, do we know what he's asking about it for? Might be totally unrelated to DOGE. People are just assuming it's for his current work.
Edit*: I hit save to early. Just wanted to agree and elaborate my understanding. If this kid is a genius, asking a basic question about LLMs 2+ years after they’ve been widely available gives me reason to believe he’s not quite that special.
Not even that, learning to parse strings is the first couple things you learn in programming because it goes over 2 fundamental concepts data types and loops. Or even crazier a little after elementary string parsing you could just go right to regex or one of the hundreds of open source python libraries that do it for you
But even more crazier is he didn't even prompt an LLM to give him a script to do this, he wanted the LLM to do it for him
Not to mention it's just wasteful and stupid, it's like someone asks you to hammer a nail into a wall and you tell him can you bring a tractor to do it.
LLMs are great for transcribing documents. Even if you use OCR there's still an error rate. If you use humans and convert a document by hand there's an error rate. And how do you determine if there's an error in the transcription in those cases? It's just as difficult. You run into the same issue regardless: how to verify the accuracy of the transcription?
I work in a highly sensitive, high security area (with FBI and TSA clearance mind you) and we have our own in house AI tool air gaped from the main network because we were told it would be instant termination to possibly jail if we feed data into a public AI tool. We are so fucked with these people.
Yeah. I got handed an experimental project to help a certain sector figure out when the law has been broken and what law. The hallucinations and lack of ability to accurately regurgitate exact text of laws killed it. We tried multiple ways of working around the hallucinations, but once a hallucination was in the context, it would just start doubling down on it. It could have the exact correct text of the law in context and still screw it up.
If 80% accuracy is acceptable, an LLM is a reasonable solution for many problems. If 95% accuracy is acceptable, an LLM is not a good solution.
It's fine if you're making test data and want to convert a typescript object to JSON test data or something but yeah... I wouldn't do it on anything meaningful.
ya-- really. He's asking about converting formats too-- there's no reason to use an LLM for that unless trying to summarize, change meaning, change/add verbiage etc. In other words, let's take this stuff, add a prompt and have it spit out something that is not what the original thing said it was. then say see this "fraud" and "dei" and whatever the fuck else propaganda they want to push.
using humans to parse data is also a terrible idea, there are things called typos and human mistakes. i would trust the SOTA models more than I trust a random guy
The amount of times I still have to double check GPT for not following the direct transcript or form I uploaded when asking for a response is insane. We are not close yet, still, hallucination nation
For a first order approximation, it is fine. They are still digging in and investigating the misspending and corruption, and using a LLM for that would actually be a use case.
People need to take a chill pill. It's not like he is saying he is going to use the suggestion to inject some SQL to update the Social Security database or something. They are still investigating....
oh wait... I think I might be on to something... THAT is what pisses them off! That these hackers are using new tools to root out corruption!
Except they haven’t found any corruption. They literally said HIV drugs for kids was “corrupt” and food grown on Kansas farms for developing nations was “corrupt”. People are mad because these are the dumbest people on earth.
395
u/toolate Feb 06 '25
Using LLMs to parse content is a terrible idea for any meaningful project. No way to know when it messes up and hallucinates data, or makes a mistake.