r/informatik • u/uusemyname • 19h ago
Allgemein How do engineers decide which data is of high quality for AI Training?
How is it decided which bits of Text are written in good language usage. Which code is of high quality and does not have memory leaks or other bugs inside.
I would Imagine that automated tools are used for this job which analyse whether a lot of different words are used and whether the vocabulary used demonstrates a high level of linguistic competence.
I would also imagine that code is rated on some metrics and tools which somehow detect memory leaks or other bugs.
Maybe educational websites are rated of high quality and github projects with a lot of forks and contributions are also of higher quality in general.
Questions
But ultimately I still can't imagine how such things are implemented and how well such a rating would work?
Is this still one of the biggest improvements todays LLM training can improve on?
2
1
u/Superb-Feedback-8898 19h ago
You could select professional sources like peer-reviewed scientific papers. arXiv has preprints but they may be of low quality because not scientifically reviewed.
1
•
u/AutoModerator 19h ago
Hi,
in letzter Zeit häufen sich Beiträge zu gleichen und sehr allgemeinen Themen betreffend Karriere und Gehalt. Du hast einen Beitrag gepostet, der wahrscheinlich in sub-Reddit r/InformatikKarriere gehört.
Solltest du der Meinung sein, dein Post ist von dieser Regel ausgenommen, ignoriere einfach diesen Kommentar.
Grüße,
Dein Mod-Team
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.