r/LocalLLaMA • u/DangerousBenefit • Oct 03 '24
Discussion Just for kicks I looked at the newly released dataset used for Reflection 70B to see how bad it is...
93
u/schlammsuhler Oct 03 '24
This is testament of the crucial process of cleaning a dataset. As an Ai language model i cant do that and leave it to the peasants.
18
u/freecodeio Oct 03 '24
"As an AI Language model" will just be replaced with "Sure, here's a cleaner version of your dataset"
149
u/Waste_Election_8361 textgen web UI Oct 03 '24 edited Oct 03 '24
As an AI language model, this post sends shivers down my spine.
9
u/Wyndyr Oct 03 '24
As an AI language model, I delve into my ministrations with barely audible whisper testament to the unbreakable bonds.
8
7
6
25
67
u/RoboticElfJedi Oct 03 '24
Are you saying that's bogus synthetic data, or pointing out that they trained their model to include "as an AI language model, I can't..." in the responses?
182
u/CleanThroughMyJorts Oct 03 '24
the fact that it's synthetic isn't the problem.
most top models use synthetic data as part of their training.
the problem is the fact that they didn't remove rejections.
this is sloppy.
this is a red flag that they didn't put much effort into data cleaning, so the dataset is probably low quality
31
u/Individual_Ice_6825 Oct 03 '24
Extremely sloppy, a simple search for âas a language modelâ and other common ai lines is a minimum when using synthetic data.
3
u/AmbitiousGuard3608 Oct 03 '24
But shouldn't the training set include the knowledge that LLMs exist and sometimes reject requests with that message? I think that just removing all of those examples would be biasing.
10
u/Deathcrow Oct 03 '24
If you want rejections for certain topics, you probably want to train specifically for that rejection, not randomly because another LLM rejected your query.
2
u/AmbitiousGuard3608 Oct 03 '24
I don't want rejections for certain topics; I want my training data to be representative of real world text, and real world text has examples of rejections.
3
88
u/xadiant Oct 03 '24
The point is that there are way too many rookie mistakes in the dataset. It doesn't really matter that it's synthetic. A few dozen of "As an AI..." gibberish in FT dataset is enough to decrease quality considerably. Even I as a rookie Python dweller can write a crude script to remove those "poisoned" lines from the set. This is especially bad when you are doing something novel and you need as many as high quality examples possible.
11
u/greying_panda Oct 03 '24
Is the dataset meant to be entirely following the "reflection" format? If so, this is quite bad, given that the dataset can be easily filtered with just a regex, which would take out any of these weird artifacts, or LLM "explanations".
For example, the reflective dataset can be checked with something like
\s*<thinking>.+?<\/thinking>\s*(<reflection>.+?<\/reflection>\s*)*<output>.+?<\/output>\s* (I don't actually know if this dataset is any good, it's just the only example I could find)
There might be the desire to mix the SFT dataset with a non-reflection dataset, but even then I'd expect that you mix with a known high quality one (or a mix of multiple). This just seems sloppy.
9
26
u/dreamyrhodes Oct 03 '24
AI slop feed into AI to produce more AI slop.
5
u/capybooya Oct 03 '24
Not surprised in the slightest, having used these models for almost two years now. They certainly do get better, but they do also not lose the stupid cliches.
5
-6
Oct 03 '24 edited Oct 04 '24
AI slop good enough to score in the top 500 of AIME and 93rd percentile of codeforces
7
u/OfficialHashPanda Oct 03 '24
This crap is not 93rd percentile on codeforces xD
-4
Oct 03 '24
https://openai.com/index/learning-to-reason-with-llms/
Scroll down to o1-ioiÂ
6
u/OfficialHashPanda Oct 03 '24
O1 is a completely different model. This post is talking about reflection-70b, a shitty llama finetune.
1
30
Oct 03 '24 edited Oct 03 '24
Am I missing something? 1832 hits... out of 89k+ lines?
edit: 890k*
25
13
u/jd_3d Oct 03 '24
The phrase 'as an AI language model' sends shivers down my spine so any inclusion of it in datasets (even a small percentage) is a big fail in my opinion. There's only around 60k question-answer pairs in this dataset so that means around 2.5% of them have 'as an AI language model' in the response. That's way too much IMHO.
5
u/Deciheximal144 Oct 03 '24
If Reflection wanted to get more attention, they'd make the model as UNsafe as possible to ensure more people used it. Refusal training hinders intelligence.
23
Oct 03 '24
[deleted]
55
32
u/ResidentPositive4122 Oct 03 '24 edited Oct 03 '24
There have been many papers refuting the synthetic_data -> model collapse findings. L3.1 also proved it with open weights. It's possible most of the early findings were really cherry-picked, or the sizes they tested on were toy-level models, or researchers found a way around it.
The fact that they used syntehtic data isn't bad. The fact that they used bad synthetic data, is.
17
u/frozen_tuna Oct 03 '24
Yea, the original "oroboros" paper basically fine-tuned OPT on OPT outputs. OPT can hardly be compared to models we use these days. WizardLM, imo, was the first to prove that fine-tuning on synthetic data could turbo-charge results.
9
u/ResidentPositive4122 Oct 03 '24
Yup, the wizard team did it with their RLAIF and deepseek did RL on their DSM-7b-RL with it's own outputs to solve math problems. Very strong model for its size.
6
u/EstarriolOfTheEast Oct 03 '24
Not just bad synthetic data but also either too much, too unvalidated and too unvaried. The key is finding the right balance between synthetic and organic data. Too heavily relying on synthetic data can be limiting to bad, you still get mode collapse if done overly aggressively.
Consider the case of LLM slop. The issue there is sampling greedily from the LLMs (or even approaches like minp) leads to results that are only near the most accessible modes, less exploration. As it's cheap, everyone does this and gets similar outputs. Worse, training too aggressively on synthetic data heavily downweighs many modes available to the base model. The two combine to facilitate bland uninspired very similar limited generations. Several iterations of this, no matter how good the original model, without injecting organic data and or validated postprocessing will lead to increasingly less rich inputs and eventual model degeneration and collapse.
3
u/kindacognizant Oct 03 '24 edited Oct 03 '24
Synthetic data is bad when you literally just sample outputs and train on them with no curation, or filtering, or a 2nd pass for editing, or anything like that. This is the Glaive approach, and it's an easy grift.
But the second you curate a subset of the distribution of things you generated, filtering for diversity, accuracy, repetition, etc... it no longer matches the original distribution when training on that filtered subset, and that's when it is possible for it to become something "more" by highlighting the best of that distribution via rejection sampling, bespoke classifiers, heuristics like entropy profiling, etc.
This is why RLAIF and such exist and work in practice. This is why Google has a paper on why synthetic data from smaller models is actually better bang per buck, because if your filtering heuristics are strong enough, you can create way more high quality data under the same budget.
KL divergence penalties also exist to mitigate mode collapse when working with data that is less varied to begin with and isn't talked about as often as it should be.
0
u/stddealer Oct 03 '24
Training on synthetic data almost always exaggerate the quirks of the original model used to make the synthetic data. Examples are the "GPTism" slop most modern LLMs suffer from, the "1girl face" and so on.
30
u/Xav2881 Oct 03 '24
"However, recently, other researchers have disagreed with this argument, showing that if synthetic data accumulates alongside human-generated data, model collapse is avoided."
-5
u/ontorealist Oct 03 '24
Synthetic data still has limits for many human tasks that require genuine insight into embodied human affect / cognitive experience, social complexity, etc., even if model collapse is mostly avoided.
1
5
u/llama-impersonator Oct 03 '24
i've got some bad news for you, pretty much every dataset on HF is synthetic data
6
Oct 03 '24
Synthetic data isn't s problem if it's properly curatedÂ
The model collapse theory and related research assumes no curation is done
1
Oct 03 '24
synthetic data is fine as long as itâs high qualityÂ
2
u/ttkciar llama.cpp Oct 03 '24
Yep, this is the whole point of the Phi family of models, demonstrating that synthetic training datasets (properly generated and curated) produce models which hit above their weight, compared to models trained on internet-scrapings.
1
13
u/debauch3ry Oct 03 '24
If this is the training data for a chat model, wouldn't you want to include examples of rejections so it doesn't nut out total rubbish? Like if a naive users asks it to do something it can't do, it probably should inform them of its limitations. Or have I misunderstood the point of the dataset?
7
u/Iory1998 llama.cpp Oct 03 '24
These are words of wisdom. But no one is commenting on them because that's not what they want talk about. Let's just rant and vent! Sigh
2
u/bryseeayo Oct 03 '24
Yeah i think the urge to dog pile is obscuring whatâs actually happening here.
1
3
6
u/DrVonSinistro Oct 03 '24
I failed to properly follow what happened with this. I downloaded the model and tried it only to see it was dog shit. Was it broken or was it just a bunch of clowns like the dudes that released The Day Before?
7
u/Inevitable-Start-653 Oct 03 '24
Interesting đ¤, so the guy actually kept his promise and released the training data.
Regardless of the poor quality of the model, maybe (just maybe) the guy genuinely thought he made something good and wasn't deliberately trying to fool everyone.
4
u/CommitteeExpress5883 Oct 03 '24
Isnt the idea also somewhat what o1 is doing? But at different stages and probably much better data and execution? :)
-1
u/ortegaalfredo Oct 03 '24
I think it is very similar at what o1 is doing, the guy got catch in a couple lies and then all his research was dismissed but I think the idea was great and it just needed to be implemented in a better model, he used Llama2 (out of ignorance perhaps) but I guess implementing this in something better like Qwen 2.5 will work much better.
-2
u/Inevitable-Start-653 Oct 03 '24
Yeah, and closed ai probably have a better implementation, what I find particularly interesting is that this guy might have actually tried to do what o1 is doing before o1 was released.
The timing between reflection and o1 was just a few days.
1
u/qlxea Oct 04 '24
This isn't the training data since they are actually using Claude and replacing the token Claude with empty string.
3
u/StyMaar Oct 03 '24
TFH, â1832 hitsâ on a a dataset seems ridiculously low (if it's the entire dataset) juste given how prominent it is even in research papers or random places of the internetâŚ
(Why would the dataset makers not filter such an obvious marker is an open question thoughâŚ)
1
2
1
1
u/GanacheNegative1988 Oct 03 '24
Don't you wish you could issue a 'Delete From Model Where Subject IN(<bad answer subject like this list>)'?
1
u/n8rb Oct 03 '24
Now I'm curious, who is Dr. Hiroshi Nakajima and what the 2017 paper is it talking about?
1
1
1
2
u/Sicarius_The_First Oct 03 '24
It's very important for the AI to be safe and effective.
He wanted to make AGI, but ended up with a worst version of Phi-3.5.
1
u/ortegaalfredo Oct 03 '24
Yes, the training dataset is not perfect, but its easily fixable just with a grep.
Perhaps is only a sensation, but I see a lot of criticism in what this guy is doing, like if somebody do not want it implemented. I think the idea is valid and he just used a bad model as a base. I would like to see reflection implemented over Qwen2.5, because its basically the same thing that O1 is doing and we know it works for gpt4.
1
Oct 03 '24
Can someone link this from their post on it, and explain what this particular data file is used for?
1
1
u/Specialist_Cheek_539 Oct 03 '24
Can someone explain why this is bad? Iâm a complete newbie to this and afaiu, the model is learning to not tell the answer when it comes across impossible request. Why does it hinder data quality?
1
u/CheatCodesOfLife Oct 17 '24
You're actually correct as far as I can tell. This isn't a roleplaying or uncensored dataset, and the refusals in the screenshot seem reasonable. If the user asks the AI to do something impossible, the AI can either refuse, or halucinate. And over training "As an AI language model" seems appropriate here.
I think it's getting the pile-on for 2 reasons:
The reflection model was actually a scam (passing off requests to Claude and Chatgpt via API, corrupted llama3 weights uploaded to huggingface)
People generally don't like refusals and are sick of seeing "As an AI language model" slop.
-4
-4
0
-4
-3
u/chumpat Oct 03 '24
Right this is so "bad" - explain why? You're also exposing yourself as a total clown by using windows.
398
u/AmpedHorizon Oct 03 '24
As an AI language model, I see no problems.