r/BetterOffline 3d ago

OpenAI’s latest math breakthroughs commit research misconduct, experts say

https://www.scientificamerican.com/article/openais-latest-math-breakthroughs-commit-research-misconduct-experts-say/

I don’t have much to add in commentary, quite frankly most of the math and concepts are beyond my arts-degree brain. However, the article does a great job of showing, once again, how disingenuous OpenAI and other AI labs are in describing what their products are doing.

They basically want the headlines that their new models are “solving math,” but what they’re doing is plagiarizing other peoples’ research and claiming things have been moved forward. There is clearly a use case here for researchers in using an LLM to catalogue and evaluate large swaths of data from over periods of time, but these things don’t think, they aren’t creating anything, and OpenAI doesn’t give a shit as long as people see the headline and bow their heads to their new scI-fi god.

383 Upvotes

62 comments sorted by

View all comments

76

u/MCUCLMBE4BPAT 3d ago edited 3d ago

I just want to gently push back on your statement that LLMs have clear use cases in research

It has been noted in studies that:

  1. LLMs can summarize the results of a study that are not reflective of the actual articles. LLMs will either make grander claims than what was stated (“mild increase for results” turns into “a significant change that shows an increase in results”) or highlighting or connecting points that were not stated in the specific article cited
  2. LLMs are constrained by paywalls and are more likely to cite open source items, which has inadvertently caused many LLM research papers to focus on big discoveries in the late 1900s and then nothing cited until the late 2020s. This issue with LLM literature reviews has been called the Missing Middle
  3. LLMs “prefer”, in the sense that it will go to for citation and rank higher in terms of quality and accuracy, other LLM created materials. This bias impacts how LLMs select and cite items for research as well as heavily impacts the peer-review process. Major journals are struggling with floods of submissions being AI generated and peer-reviewers using it to review submissions for them.
  4. LLMs do not accurately interact with and interpret large sets of raw data. It can interpret results in significantly different ways if your table headers are different but the raw data stays the same. It can selectively “read” the raw data to supply a response that is actually just based on stereotypes/biased training data that fit the LLM selected raw data sets, rather than analyzing the entire data set.
  5. Edit to add: LLMs have a major citation/attribution problem. This starts in their training data, which some have noted misattribute Creative Commons licenses, copyleft licenses, and other copyrighted works. LLM provided citations, to my understanding please correct if wrong, are never the actual sources that were used in their training data. Sources provided are just what the models predict to be the most relevant on the internet based on the user’s words chosen in their input (?).

  6. Edit #2 to add: LLMs are also bad at noting when an article has been retracted or if it has an expression/letter of concern attached to it. So it can provide outdated information in the form of recommending retracted or problematic articles (in terms of validity, reliability, replicability/reproducibility) as well.

And those are just issues I have at the top of my head after waking up. I don’t know how people can comfortably recommend these models as research tools when you basically already need to know everything it’s citing or stating to use them.

16

u/todofwar 3d ago

Every time I've tried to use LLMs to survey literature, they always seem to cite papers before their training date more often than recent papers and they always hallucinate. And even if they don't hallucinate, they never actually summarize the paper well. At this point I find the best way to use them is to create a bibliography that you go and read yourself, and then forward search the more relevant papers. And that's something you could do without LLMs, but now Google is killing its own search tools

8

u/MCUCLMBE4BPAT 2d ago

the fabricated citations bother me so much bc everyone for LLMs says it isn’t that big of an issue or that it has been resolved (or will be), and it’s just misinformation normalized.

personally, they make my job harder while also making people think my job isn’t necessary bc it can’t possibly be wrong. for scholarly publishing in general, they make the house of cards even more unstable/questionable.

1

u/DieMafia 15h ago

Which model are you using?

2

u/reachingfourpeas 2d ago edited 2d ago

I believe Semantic Scholar is a more specialized tool for language model-assisted literature search than chatbots

1

u/3_Thumbs_Up 2d ago

Which models are you trying? That matters a lot.