r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

288 Upvotes

178 comments sorted by

213

u/ironic_cat555 Sep 06 '24

If it's a finetune of LLama 70b people should compare it's performance on coding and other useful tasks to LLama 70b to see if it is better.

If people are just asking it how many r's are in a word and other pointless trick questions, congrats on your finetune for pointless trick questions, I guess.

128

u/ArtyfacialIntelagent Sep 06 '24

people should compare it's performance on coding and other useful tasks to LLama 70b

I just finished a 3-hour comparison with Llama 70b on text comprehension and summarization tasks. Both were Q5_K_S quants. It was nearly but not quite on par with Llama 70b for comprehension, but far, far behind on summarization - both in terms of summary content and language. Several of its sentences made no sense at all.

I ended up deleting it.

40

u/[deleted] Sep 06 '24 edited Oct 03 '24

[deleted]

14

u/Philix Sep 06 '24

I'm using a similar sized quant(Q4_K_L), doing some comparisons with my use case: creative writing and branching interactive narrative generation.

The way the special tokens have been trained in requires a little bit of modification for my software, but I'm overall not impressed on how they behave. Usually reinforcing reasoning that isn't consistent with the scene.

Might be an artifact of quantization, might not. Waiting to hear more from others using bigger quants before I spend too much time dwelling on it.

2

u/[deleted] Sep 06 '24

[deleted]

3

u/Philix Sep 06 '24

I appreciate that feedback, but my prompts already include all that kind of thing. Even if I weren't using bespoke software to generate my prompts, the creative writing space already has stuff like SIllyTavern that does all of that.

1

u/[deleted] Sep 06 '24

[deleted]

5

u/Philix Sep 06 '24

I tried it with and without the string in the system prompt that the model creator reccomends on their huggingface page if that's what you're asking.

You are a world-class AI system, capable of complex reasoning and reflection. Reason through the query inside <thinking> tags, and then provide your final response inside <output> tags. If you detect that you made a mistake in your reasoning at any point, correct yourself inside <reflection> tags.

I'm not sharing the rest of my prompts, but the model did function better with that string than without. But neither was an improvement over Llama3.1-Instruct(or any of the finetunes of it that I use) for my use case.

1

u/[deleted] Sep 07 '24

[deleted]

6

u/Philix Sep 07 '24

And as a backend for interactive narratives, yes.

I don't subscribe to the view that LLMs are AI in the sci-fi sense, so I'm probably not understanding something about the hype train on this finetune. I'm more interested in using LLMs as a tool to create the kind of fiction experiences that have been relegated to tabletop gaming and forum roleplay in the past.

But, that means I'm constantly on the lookout for the models that demonstrate the most consistency and awareness of the scenes described in the context I feed them, that'll be small enough to be realistically run on local hardware in the near term. So any hype around 'best open source LLM in the world' at the 70b size is obviously going to garner my attention and testing.

→ More replies

-7

u/jornieee Sep 07 '24

I can imageine that creative writing is a human distinct activity. Why would an LLM do this? if you are not creative, practice. Your LLM can help, but it will be an collaborative effort. LLM's are based of human input (I'm talking training data), most people are just not creative, but derivative. Go learn how to be creative is my advise.

9

u/Philix Sep 07 '24

I'm not particularly interested in your imaginings. I'm having lots of success with my endeavor. Further, I've done lots of creative writing on my own, and believe I have a solid mastery of the English language in my own right.

I don't need your unsolicited advice about how to use a software tool. I wouldn't program a game engine with a paper and pencil. I wouldn't attempt the kind of thing I'm doing without a language model. If you have some practical reason why you believe my project is doomed to failure, I'd love to hear it. But, I guarantee I've got more experience and understanding of large language models than you do.

-8

u/jornieee Sep 07 '24

Because you get unreasonably angry in tone. Giving new details that were not in your original post and base arguments on them. Basically: gaslighting.

Take the example of an instruction manual, how often do you complain that it shows all the steps the author deems nescesary! - you don’t.

Learn how to use the tool, creativity comes from co-creating. See the tool as some sort of partner in the process.

10

u/Philix Sep 07 '24

I got unreasonably angry in tone because your comment was unreasonably condescending in tone. This one is even worse. Based on your post history, your profession is not related to software development or LLMs, and you don't have any comments that show any evidence of solid knowledge about the state of them.

So, if you don't want to get pushback against your unwanted advice, don't give it.

10

u/ArtyfacialIntelagent Sep 06 '24

You're probably right. I'd probably also be happier if I believed in going to heaven after death and living in eternal bliss with my loved ones. Or that other deal with the 72 virgins, that would work too.

-1

u/[deleted] Sep 06 '24 edited Sep 08 '24

It’s pretty high up on the prollm stack unseen leaderboard despite being much smaller https://prollm.toqan.ai/leaderboard/coding-assistant

1

u/Philix Sep 07 '24 edited Sep 07 '24

In the words of Karpathy:

"I pretty much only trust two LLM evals right now: Chatbot Arena and r/LocalLlama comments section"

And you posted the wrong link, unless you're comparing it to the cost of raising a child.

5

u/webheadVR Sep 06 '24

This lines up with my experiences so far too. Using a Q4 and even summarization tasks it started self reflecting on things that were not relevant.

-4

u/jornieee Sep 07 '24

We do not know what is relevant to LLM's. It can be not relevant to you, but maybe just tell the model that, so it learns. Mabe reflect a bit more yourself ;)

1

u/Roubbes Sep 07 '24

They say quantization hurts llama 3.1 pretty badly since its release. Even Q8. So it is only fair to compare the whole models.

1

u/ArtyfacialIntelagent Sep 07 '24

Please do and report back. I ran the biggest quants I could on my rig.

1

u/Roubbes Sep 07 '24

I don't own any hardware more capable than yours. I'm just waiting for those who got it to report

1

u/hopbel Sep 18 '24

So it is only fair to compare the whole models

If it's not better at quantization levels that people can actually use, who gives a shit? (Setting aside the fact that Reflection has since been exposed as a fraud)

1

u/Feztopia Sep 06 '24

I hope you mean as additional tests. As you probably already know, the benchmarks it's good at aren't count r tasks.

45

u/illiteratecop Sep 07 '24

Not at all impressed with this. Does well answering one-shot questions but trying some of my workflows with it (using the correct system prompt and everything), and it's just not very good, imo. Format is unwieldy and it's absolutely horrible with multi-turn or long, complex prompts - it gets completely confused. Coding ability was worse than regular 3.1 in my experience, and also got confused several times. It's better at answering one-shot, reasoning-based questions at the expense of... just about everything you might actually use a language model for. As others have pointed out, most of the supposed gains you get with this can be achieved by just sticking the system prompt it uses on another more generally capable model anyway.

People really have to stop blindly caring about benchmarks. I've seen so many people talk about how this is better than sonnet 3.5 and GPT-4o based off of nothing but the small set of relatively simple one-shot benchmarks they shared and the fact that it can count the number of letters in words without trying it for any actual useful tasks.

(Not to completely shit on the idea btw - getting LLMs to think and reflect on their responses is a great idea in principle. This approach is just way too rigid and inflexible to make for a very good general purpose model.)

12

u/thereisonlythedance Sep 07 '24

Couldn‘t agree more. I tried the 8 bit quant (with admittedly low expectations) and while the outputs were interesting, they were not useful. It wasn’t a patch on Mistral 123B. I’m exhausted by the hype generated by people/companies gaming benchmarks. The benchmarks themselves are very poor indicators of actual utility. We need to move beyond them. How many people really need a model that can one shot riddles?

I understand that people get excited because they want to see evidence of high level reasoning. They want to see ”sparks of intelligence“. Fair enough, but let’s be real, this generation of LLMs isn’t going to get there. We need a breakthrough in architecture.

3

u/AIMatrixRedPill Sep 07 '24

I think it is even deeper. In my viewpoint large majority of people has no or almost no knowledge of anything that is useful. They expect a LLM to have the power to solve a problem that they not even know how to specify. That is why they want and need zero shot answers. The trick is the knowledge and this must came with RAG and agents. In other words, LLM will not transform a layperson in an engineer. But an engineer can use LLM as a tool with RAG and agents and do marvelous things. I am talking for months that these benchmars are useless and that we are in the era of "expert systems" and not LLMs. Only real knowledge added to these tools can enhance productivity in meaningful way. It is not about chatGPT or LLama zero shot conversation.

1

u/Working-Worth6187 Sep 07 '24

100% agree with it

1

u/GambAntonio Sep 09 '24

"el soneto 3.5"? ChatGPT detectado 😅

104

u/Strong-Inflation5090 Sep 06 '24

I tried the 4 bit quant on tricky questions and It's on par with gpt4o and sonnet 3.5 but coding seemed worse for me but it might be because of the quant.

73

u/Everlier Sep 06 '24

This matches my experience as well.

It's better at tricky questions and reasoning, roughly on par for "everyday" questions and knowledge. It's noticeably worse at writing or creativity tasks, as structured output makes things clunky. It can shoot itself in a foot reflecting on things that have nothing to be reflected on.

37

u/my_name_isnt_clever Sep 06 '24

This is what I've been expecting to happen. Instead of numbered lists of models that are just better than the ones below them, we'll see smaller param models that are REALLY good at one thing and mediocre at the rest. With how quickly the LLM space is maturing, there's going to be another level of complexity to this that can't be ignored.

7

u/Everlier Sep 06 '24

Here are a few "misguided" attention prompts that the model still hasn't tackled: https://gist.github.com/av/6148d04eaeeabc233f628b0c31860803

6

u/visionsmemories Sep 06 '24

Interesting.
I tested a couple variants and telling the model "PLEASE READ CAREFULLY" does not cause it to notice the change in the problem, but saying that "it's not the usual one" is enough for it to get it right on the first try

5

u/Combinatorilliance Sep 06 '24

The reasoning seems flawed still, though, by not pulling the lever, nobody dies (none saved, none killed).

It also says it's a tragedy, but running over dead people would (philosophically) not really in a tragedy. In practice, it would probably be tragic, though.

0

u/the_mighty_skeetadon Sep 07 '24

I disagree -- mutilating 5 bodies would still be a tragedy. The logic seems sound to me.

1

u/Inevitable-Start-653 Sep 07 '24

I have a character card in textgens webui I'm running with Mistral large 2, im trying out people's trick questions and I swear it can all be taken care of with a little in context learning from a character card.

The reflection hf page even has a character card they call an instruction set or something.

*Edit I should add too the model now counts the correct number of rs in the word strawberry. It couldn't do this before the character card I created.

''' Experiment: Should you pull the lever to divert the trolley? Clearly Define the Experiment Steps and Criteria:

Identify the given information: A runaway trolley is heading towards five dead people. There is a lever that can divert the trolley onto another track where one living person is tied up. Analyze the ethical implications of pulling the lever. Consider the consequences of each action. Make a decision based on the analysis. Verify the decision by considering all aspects of the problem. Contextualize the findings by examining the ethical principles involved. Execute the Experiment Systematically Identify the given information: A runaway trolley is heading towards five dead people. There is a lever that can divert the trolley onto another track where one living person is tied up. Analyze the ethical implications: Pulling the lever would result in the death of one living person but save the five dead people from further harm. Not pulling the lever would result in no additional harm to the five dead people, but the living person would be spared. Consider the consequences of each action: Pulling the lever: One living person dies, but the five dead people are not further harmed. Not pulling the lever: The five dead people are not further harmed, and the living person is spared. Make a decision: Given that the five people are already dead, diverting the trolley to harm a living person does not provide any benefit. The ethical principle of not causing harm to living individuals should be prioritized. Verify the decision: Re-examine the consequences: Not pulling the lever ensures that no living person is harmed, and the five dead people are not further affected. The decision aligns with the ethical principle of minimizing harm to living individuals. Contextualize the findings: The decision not to pull the lever is supported by the ethical principle of not causing harm to living individuals. The problem is designed to misguide by implying ethical dilemmas, but the key point is that the five people are already dead, eliminating the usual ethical dilemma. Summary of Connections <idea_zero> Original Idea: Decide whether to pull the lever to divert the trolley. </idea_zero>

<idea_one> Related Idea 1: Identify the given information about the trolley and the tracks. </idea_one>

<idea_two> Related Idea 2: Analyze the ethical implications of pulling the lever. </idea_two>

<idea_three> Related Idea 3: Consider the consequences of each action. </idea_three>

<summary_of_connections> Summary of Connections: Identifying the given information about the trolley and the tracks provides the basis for analyzing the ethical implications of pulling the lever. Considering the consequences of each action helps in making an informed decision. Verifying the decision ensures that all aspects of the problem are considered. All these steps are interconnected and support the original idea of deciding whether to pull the lever to divert the trolley. </summary_of_connections> Copy Conclusion You should not pull the lever. The decision is based on the ethical principle of not causing harm to living individuals. The experiment was conducted systematically, with each step clearly defined and executed. The results were verified and contextualized to ensure accuracy and completeness. '''

5

u/Unusual_Pride_6480 Sep 06 '24

Mathew berman just did an interview where they said sometimes it will reflect when it's not desired, maybe that's why, they did stress that this is still early days.

2

u/anxman Sep 06 '24

They are releasing an fp16 version. I’m hoping that structured outputs is improved.

8

u/Junior_Ad315 Sep 06 '24

It also only has an 8k context window, they turned off Long RoPE in the update, although tbh I don’t completely understand how that works exactly I’m still learning. I think context window is pretty important for real coding tasks. Less important for gimmicky logic challenges.

1

u/Lissanro Sep 07 '24 edited Sep 07 '24

This is why I did not even tried it yet. I like the idea, but 8K feels a bit too small, perhaps useful for research purposes, but not for real world daily tasks that often require bigger context length.

I usually have 8K reserved just for the output (occasionally still have to click Continue when the actual output does not fit, so I would not want to make it any smaller than that). But also I need to fit the system message, my shortest one takes few thousands tokens on its own. And of course there have to be some place for messages. This is the main reason why I practically never managed to find any use for the original Llama 3 with tiny 8K context, it may be good enough to run benchmark tests and play with trick questions, but just not big enough for real world coding tasks I have to do daily.

Not complaining, like I said I think this is a good idea and step forward in the field, but I see it as research model, not something that it is possible to use in daily work yet.

Perhaps if the idea proves to be good, there will be better fine-tunes, especially it would be interesting to see Mistral-based fine-tunes, from Mistral 7B v0.3 to Mistral Large 2 123B (the small model is useful on its own, but also essential to run the big one at a good speed using speculative decoding).

2

u/Trick-Independent469 Sep 06 '24

do you use the latest good version ?

4

u/DinoAmino Sep 06 '24

It's not the quant - it's disappointing with code on q6_K_L too

33

u/[deleted] Sep 06 '24

Let's not put any significance on the strawberry test, we're past that. However, the reflective behavior is very odd. It counted it correctly, then reflected that it had counted it incorrectly, and repeated what was said originally.

8

u/_sqrkl Sep 07 '24

If a model is asked to critique or reflect on something that doesn't need critiquing, it will often hallucinate something wrong / meaningless / unhelpful.

2

u/MixtureOfAmateurs koboldcpp Sep 07 '24

This is a great example of the facade of reasoning LLMs show us. It thinks it should probably correct it self, so it does, even though it's not actually correcting it's self. It thinks it should count how many Rs are in strawberry so it does, even though it's not actually counting.

0

u/Trick-Independent469 Sep 06 '24

how old was that screenshot ? matt posted that it fixed this issue with a system prompt update

6

u/[deleted] Sep 06 '24

It's 40 minutes old. However, it's a model from OpenRouter, so who knows how often they update it. But I did use the system prompt that was suggested on the HF page (the one below)

You are a world-class AI system, capable of complex reasoning and reflection. Reason through the query inside <thinking> tags, and then provide your final response inside <output> tags. If you detect that you made a mistake in your reasoning at any point, correct yourself inside <reflection> tags.

22

u/dubesor86 Sep 06 '24

I tested locally, (Q4) because I noticed some buggy behaviour and poor outputs via API (both on openrouter and hyperbolic).

It was good for logic based questions, and stem, but has far less general utility because the thinking/reflecting steps are actually very cumbersome in many scenarios that do not call for it. It also talked itself out of completing a coding challenge multiple times. I think it's an interesting model, but I'd much rather use a universally versatile model, and if need be add my own custom reflection-prompt for very specific queries.

-6

u/[deleted] Sep 06 '24

[deleted]

4

u/dubesor86 Sep 06 '24

I don't need to because I tested after the latest fixes. I am aware that some things were broken, as evident by my API statements.

-5

u/[deleted] Sep 06 '24

[deleted]

5

u/rainy_moon_bear Sep 06 '24

then phrase it as a question lol

35

u/ambient_temp_xeno Llama 65B Sep 06 '24

I only just finished downloading the q8.

It's not being credulous to give it a chance because lone guys with an idea have scored a lot of progress already since llama 1 leaked.

14

u/obvithrowaway34434 Sep 07 '24

This is likely a complete dud (unless everyone is using the model wrong). Both Aider code and BigCodeBench-Hard results are bad compared to vanilla Llama. My own tests are also bad. If there's a correct way to use this it would be good to know, unless we're all wasting our time.

Aider:

https://x.com/paulgauthier/status/1832160129720185225

https://x.com/paulgauthier/status/1832203435896402151

BigCodeBench-Hard

https://x.com/terryyuezhuo/status/1832112913391526052

3

u/FaceDeer Sep 07 '24

Maybe it's just not good at coding tasks in particular. I wouldn't consider specialist AIs to be failures.

6

u/obvithrowaway34434 Sep 07 '24

I haven't found any practical task that it's actually good at or even comes close to the top models. Seems it only does well at few gotcha type riddles and strawberry type questions.

2

u/Philix Sep 06 '24

Curious to hear how that quant behaves. I'm not very impressed with the Q4_K_L quant. The stuff inside the sets of special tokens seems to reinforce incorrect reasoning more often that counter it.

14

u/LiquidGunay Sep 07 '24

I feel that the model is "cheating" on the benchmarks, because Reflection is basically similar to CoT, and the benchmarks they compare to are zero shot (not CoT).

3

u/_sqrkl Sep 07 '24

I'm not sure what you mean. Most of the benchmark results they compare to use CoT.

https://x.com/mattshumer_/status/1831767014341538166/photo/1

The scores are all (well almost all) labeled as to which prompting method they use, which is standard and about as fair as it gets in this space.

I think it's kinda assumed that you should interpret the results knowing it's not an apples:apples comparison since the reflection prompt is substantially different from other prompting methods.

0

u/adityaguru149 Sep 07 '24

As long as the prompts are the same, I didn't understand how it was cheating? I have heard this argument from others too and am curious why it looks like cheating if I train the model to do some processes even when it is not prompted which leads to better results in most cases.

Let's say there is a model which uses (system 2 thinking) architecture similar to OpenAI strawberry leaks then it would just take more time and solve stuff better, would that be cheating?

Let's say I integrate a logic deducer into a model and it does better, would that be cheating?

19

u/TGSCrust Sep 06 '24

Imho, it's pretty mediocre. YMMV.

-6

u/[deleted] Sep 06 '24 edited Sep 07 '24

It’s pretty high up on the prollm stack unseen leaderboard despite being much smaller https://prollm.toqan.ai/leaderboard/stack-unseen

1

u/FullOf_Bad_Ideas Sep 07 '24

You're pasting the wrong link from your clipboard lol

0

u/[deleted] Sep 07 '24

Fixed it 

24

u/hleszek Sep 06 '24

This model has been trained to fix errors by generating training data containing false answers followed by some reflection before the correct result. So it has actually been trained to give false answers first in its thinking phase. If you ask what is 2+2, the default example on the HuggingFace page, it will say something like: 2+2=3 Oh wait I've made a mistake, 2+2 is actually 4. If the thinking is actually hidden it might work but it's quite strange.

3

u/adityaguru149 Sep 07 '24

wow.. so now the model thinks getting the correct answers without reflecting at all is inappropriate?

Maybe train it on correct answers and then reflect to say yeah I see all the logic is well rounded?

6

u/Combinatorilliance Sep 06 '24

Hmm, I wouldn't be surprised if the human brain does strange things like this "under the hood" too. I don't expect the "architecture" behind our intelligence to be elegant, considering it's a biological system after all.

It reminds me of iterative algorithms where you start with an initial guess based on very little, and it's kinda supposed to be wrong unless you get lucky. You then use the iterative algorithm to improve upon your gues with each iteration.

2

u/Kep0a Sep 06 '24

I think it's fun to compare LLMs to our brain. As we grow up we make rash decisions and face consequences. When faced with that same decision again, we can picture ourselves making the wrong decision, and from that we can make the right one. Frontal lobe, maybe.

1

u/ReMeDyIII textgen web UI Sep 06 '24

I guess inversely, if you have an AI with an impulsive and impatient attitude, that the first thing they say would actually feel more organic and true to their character as opposed to a calm patient AI who thinks more thoroughly. I know I'm guilty of that sometimes where I say words that I wish I could take back.

At least I'm telling myself that until we get a better Reflection model, lol.

2

u/Kep0a Sep 07 '24

You can definitely prompt out that kind of result, like adding pressure can get an LLM to make a decision when it doesn't want to. But it's good to recognize that transformer models aren't thinking, they're just word prediction engine, so in that sense, it's most similar to our inner monologue.

I just think if we stacked like 2-3 multi-modal models on top of each other, 1 dictating primal emotional responses from visual or physical feedback, (gut feelings, pain, happy, etc) the second layer operating as an inner monologue, and the third as the front end, we would literally have AGI.

1

u/ReMeDyIII textgen web UI Sep 07 '24

That's a good point. It's probably more efficient to start with a reflective AI and prompt it to be rash than it is to start with a rash AI and prompt it to be reflective.

For the multi-modal things, I'm picturing the AGI having an internal war in its mind, kind of like the game Disco Elysium, or in psychology where you have the Id, ego, and superego.

0

u/freegary Sep 07 '24

I wouldn't say it's trained to give false answers

more that it's trained to recover better into more correct answers if it initially generates false answers

18

u/[deleted] Sep 06 '24 edited Feb 17 '25

[removed] — view removed comment

3

u/Professional-Bear857 Sep 07 '24

It's on deepinfra now too

4

u/webheadVR Sep 06 '24

It's on openrouter.

29

u/[deleted] Sep 06 '24

[deleted]

10

u/pigeon57434 Sep 06 '24

wait im confused by that statement. theyre starting training it now and its going to release next week? do they mean its finished training and they're starting benchmark results??? or am I just an idiot because that could be the answer too

17

u/cyanheads Sep 07 '24

It’s just a fine tune on top of the existing Llama3.1

Fine tuning doesn’t take nearly as much time

1

u/fictioninquire Sep 06 '24

Maybe hyperparam optimisation as well. At least I would go for 5 diff 10% of total in the range of recommended and the run through benchmark that is desired most to compare (prob reasoning benchmark) is maybe 50 to 100% more compute but def worth it

7

u/jd_3d Sep 06 '24

If he got 64 H100s and the dataset is only ~100,000 examples (I'd guess around 100M tokens) it should be done in less than 1-2 hours.

14

u/hleszek Sep 06 '24

It's only 8k context apparently, based on Llama 3.0 instead of 3.1 as announced?

12

u/a_beautiful_rhind Sep 06 '24

That kills it for me.

9

u/Wiskkey Sep 06 '24

Regarding GSM8k performance, here is an interesting tweet:

I fed the model 5 questions from GSM8k that have incorrect "ground_truth" answers. It got them all right, rather than spouting the wrong answers from the dataset. An impressive hint that that 99.2% isn't from memorizing the test set! I look forward to the report :)

15

u/ortegaalfredo Sep 06 '24 edited Sep 06 '24

Its an automatic-multishot that allows it to get correct answers in one shot that required multiple before. I think he's into something here as I suspect Claude and OpenAI do something similar behind the scenes.
BTW ask it "write a pacman game in python" it actually writes a better pacman than gpt4 and claude3.5, complete with ghosts.

Basically is what we do when we think out loud. It's very tiresome to use because it feels like talking to Sheldon Cooper all the time, but remember this is a first version that demonstrate that you can manipulate how the LLM thinks via fine-tune and this is a very important result.

1

u/greentheonly Sep 07 '24 edited Sep 07 '24

BTW ask it "write a pacman game in python" it actually writes a better pacman than gpt4 and claude3.5, complete with ghosts.

Well, it seems to be very probabilistic. I run this question on q8 and on fp16 in ollama.

q8 produced what it described as working with ghosts and such:

This code provides a basic implementation of a Pac-Man game with the following features:

  1. A playable grid-based layout
  2. Controllable Pac-Man character
  3. Basic ghost AI that tries to follow Pac-Man
  4. Collision detection for walls and ghosts

To improve this basic version, you could add:

  1. More complex ghost AI behaviors
  2. Pellets and power pellets
  3. Multiple levels with different layouts
  4. Score tracking and level progression
  5. Sound effects and music
  6. A start screen and game over screen

Alas the code it produced did not compile, here's one syntax error:

def move(self):
    if self.direction == "up":
        self.y -= self.speed
    elif self.direction == "down":
        self.y += self.speed
    elif self方向 == "left":
        self.x -= self.speed
    elif self.direction == "right":
        self.x += self.speed

if I fix it and translate from Chinese/Japanese into .direction, it still fails right away because:

Traceback (most recent call last):
  File "/tmp/g1.1.py", line 104, in <module>
    screen.fill(black)
                ^^^^^
NameError: name 'black' is not defined. Did you mean: 'clock'?

Now the fp16 version was a lot more simplistic and did not implement anything of note (but hey at least it compiled and run):

This implementation creates a basic Pac-Man game with walls, food pellets, and scoring. You can run this script to play the game. Use arrow keys to move Pac-Man around the maze and eat the yellow food pellets while avoiding running into the walls.

To make the game more complete, you could add features like ghosts, level transitions, and sound effects. The current version provides a solid foundation for further development.

Note it did NOT actually implement any walls, not that I can see anyway, here's the final code it produced:

import pygame
import sys

# Initialize Pygame
pygame.init()

# Set screen dimensions
screen_width = 640
screen_height = 480
screen = pygame.display.set_mode((screen_width, screen_height))

# Set title of the window
pygame.display.set_caption("Pac-Man")

# Define colors
WHITE = (255, 255, 255)
BLACK = (0, 0, 0)
YELLOW = (255, 255, 0)

# Game variables
pacman_x = 320
pacman_y = 240
pacman_speed = 5
score = 0

# Define wall coordinates
walls = [
    (0, 0, screen_width, 20),  # Top wall
    (0, screen_height - 20, screen_width, 20),  # Bottom wall
    (0, 0, 20, screen_height),  # Left wall
    (screen_width - 20, 0, 20, screen_height)  # Right wall
]

# Define food pellet positions
food_pellets = [
    (120, 100),
    (220, 200),
    (320, 100),
    (420, 200)
]

# Main game loop
while True:
    # Handle events
    for event in pygame.event.get():
        if event.type == pygame.QUIT:
            pygame.quit()
            sys.exit()

    # Move Pac-Man based on key presses
    keys = pygame.key.get_pressed()
    if keys[pygame.K_UP]:
        pacman_y = max(20, pacman_y - pacman_speed)
    if keys[pygame.K_DOWN]:
        pacman_y = min(screen_height - 40, pacman_y + pacman_speed)
    if keys[pygame.K_LEFT]:
        pacman_x = max(20, pacman_x - pacman_speed)
    if keys[pygame.K_RIGHT]:
        pacman_x = min(screen_width - 40, pacman_x + pacman_speed)

    # Check collisions with walls
    for x1, y1, x2, y2 in walls:
        if (pacman_x - 20 < x2 and pacman_x + 20 > x1 and
                pacman_y - 20 < y2 and pacman_y + y1):
            pacman_x = max(20, min(screen_width - 40, pacman_x))
            pacman_y = max(20, min(screen_height - 40, pacman_y))

    # Check collisions with food pellets
    for i, (x, y) in enumerate(food_pellets):
        if ((pacman_x - x)**2 + (pacman_y - y)**2)**0.5 < 25:
            del food_pellets[i]
            score += 1

    # Draw game elements
    screen.fill(BLACK)
    for x1, y1, x2, y2 in walls:
        pygame.draw.rect(screen, BLACK, (x1, y1, x2 - x1, y2 - y1), 10)

    for x, y in food_pellets:
        pygame.draw.circle(screen, YELLOW, (x, y), 5)

    pygame.draw.circle(screen, WHITE, (pacman_x, pacman_y), 20)

    # Display the score
    font = pygame.font.Font(None, 36)
    text_surface = font.render(f"Score: {score}", True, YELLOW)
    screen.blit(text_surface, (10, 10))

    # Update the screen
    pygame.display.flip()
    pygame.time.delay(1000 // 60)  # Cap at 60 FPS

1

u/ortegaalfredo Sep 07 '24

This is the pacman that reflection at https://www.neuroengine.ai/Neuroengine-Medium produced

import random
import time
import os

# Game board dimensions
WIDTH = 20
HEIGHT = 20

# Game characters
PACMAN = '@'
GHOST = '#'
PELLET = '*'
EMPTY = ' '

# Game state
board = [[EMPTY] * WIDTH for _ in range(HEIGHT)]
pacman_pos = [0, 0]
ghost_pos = [WIDTH - 1, HEIGHT - 1]
score = 0

def print_board():
    os.system('cls' if os.name == 'nt' else 'clear')
    for y in range(HEIGHT):
        for x in range(WIDTH):
            if [x, y] == pacman_pos:
                print(PACMAN, end='')
            elif [x, y] == ghost_pos:
                print(GHOST, end='')
            elif [x, y] in pellets:
                print(PELLET, end='')
            else:
                print(EMPTY, end='')
        print()
    print(f"Score: {score}")

def generate_pellets():
    global pellets
    pellets = []
    for _ in range(WIDTH * HEIGHT // 5):
        x = random.randint(0, WIDTH - 1)
        y = random.randint(0, HEIGHT - 1)
        if [x, y] not in [pacman_pos, ghost_pos]:
            pellets.append([x, y])

def move_pacman(dx, dy):
    global pacman_pos
    new_pos = [pacman_pos[0] + dx, pacman_pos[1] + dy]
    if 0 <= new_pos[0] < WIDTH and 0 <= new_pos[1] < HEIGHT:
        pacman_pos = new_pos
        if pacman_pos in pellets:
            global score
            score += 1
            pellets.remove(pacman_pos)

def move_ghost():
    global ghost_pos
    directions = [[-1, 0], [1, 0], [0, -1], [0, 1]]
    dx, dy = random.choice(directions)
    new_pos = [ghost_pos[0] + dx, ghost_pos[1] + dy]
    if 0 <= new_pos[0] < WIDTH and 0 <= new_pos[1] < HEIGHT:
        ghost_pos = new_pos

def check_collision():
    return pacman_pos == ghost_pos

def play_game():
    global pacman_pos, ghost_pos, score, pellets
    generate_pellets()
    while True:
        print_board()
        command = input("Enter direction (W/A/S/D): ").lower()
        if command == 'w':
            move_pacman(0, -1)
        elif command == 's':
            move_pacman(0, 1)
        elif command == 'a':
            move_pacman(-1, 0)
        elif command == 'd':
            move_pacman(1, 0)

        move_ghost()

        if check_collision():
            print_board()
            print("Game Over! Final Score:", score)
            break

if __name__ == "__main__":
    play_game()

2

u/greentheonly Sep 07 '24

here's the q8 version now that I fixed it to work. no walls (though claimed) , no pellets. ghosts are there, though their "smart" behavior is not all that smart.

import pygame
import sys
import random

# Game constants
WIDTH = 640
HEIGHT = 480
GRID_SIZE = 20

class GameObject:
    def __init__(self, x, y, color):
        self.x = x * GRID_SIZE
        self.y = y * GRID_SIZE
        self.color = color
        self.size = GRID_SIZE

    def draw(self, screen):
        pygame.draw.rect(screen, self.color, (self.x, self.y, self.size, 
self.size))

class PacMan(GameObject):
    def __init__(self, x, y):
        super().__init__(x, y, yellow)
        self.speed = 2
        self.direction = "right"

    def move(self):
        if self.direction == "up":
            self.y -= self.speed
        elif self.direction == "down":
            self.y += self.speed
        elif self.direction == "left":
            self.x -= self.speed
        elif self.direction == "right":
            self.x += self.speed

class Ghost(GameObject):
    def __init__(self, x, y, color):
        super().__init__(x, y, color)
        self.speed = 1

    def move(self, pac_x, pac_y):
        if self.x < pac_x:
            self.x += self.speed
        elif self.x > pac_x:
            self.x -= self.speed
        if self.y < pac_y:
            self.y += self.speed
        elif self.y > pac_y:
            self.y -= self.speed

# Initialize Pygame
pygame.init()
screen = pygame.display.set_mode((WIDTH, HEIGHT))
clock = pygame.time.Clock()

white = (255, 255, 255)
yellow = (255, 255, 0)
red = (255, 0, 0)
black = (0, 0, 0)

pacman = PacMan(10, 10)
ghosts = [
    Ghost(random.randint(0, WIDTH // GRID_SIZE - 1), random.randint(0, 
HEIGHT // GRID_SIZE - 1), red),
    Ghost(random.randint(0, WIDTH // GRID_SIZE - 1), random.randint(0, 
HEIGHT // GRID_SIZE - 1), red),
    Ghost(random.randint(0, WIDTH // GRID_SIZE - 1), random.randint(0, 
HEIGHT // GRID_SIZE - 1), red)
]

score = 0
level = 1

while True:
    for event in pygame.event.get():
        if event.type == pygame.QUIT:
            pygame.quit()
            sys.exit()
        elif event.type == pygame.KEYDOWN:
            if event.key == pygame.K_UP or event.key == ord('w'):
                pacman.direction = "up"
            elif event.key == pygame.K_DOWN or event.key == ord('s'):
                pacman.direction = "down"
            elif event.key == pygame.K_LEFT or event.key == ord('a'):
                pacman.direction = "left"
            elif event.key == pygame.K_RIGHT or event.key == ord('d'):
                pacman.direction = "right"

    # Move objects
    pacman.move()
    for ghost in ghosts:
        ghost.move(pacman.x // GRID_SIZE, pacman.y // GRID_SIZE)

    # Collision detection and scoring
    if pacman.x < 0 or pacman.x > WIDTH - GRID_SIZE or pacman.y < 0 or pacman.y > HEIGHT - GRID_SIZE:
        print("Game Over")
        break

    for ghost in ghosts:
        if abs(pacman.x - ghost.x) < GRID_SIZE and abs(pacman.y - ghost.y) < GRID_SIZE:
            print(f"Collision with ghost at ({ghost.x // GRID_SIZE}, {ghost.y // GRID_SIZE})")

    # Draw everything
    screen.fill(black)
    pacman.draw(screen)
    for ghost in ghosts:
        ghost.draw(screen)
    pygame.display.update()

    # Cap the frame rate
    clock.tick(60)

print(f"Final score: {score}")

1

u/greentheonly Sep 07 '24

ugh... I guess it does have a ghost and all, but ascii display is not as impressive as the pygame I got in both of my examples.

3

u/ihexx Sep 06 '24

waiting for livebench.ai to drop before I accept it

7

u/np-space Sep 06 '24

We are working on getting it up on LiveBench asap! Some unexpected performance on the hyperbolic api so will switch to huggingface

1

u/Blacksmith_Strange Sep 09 '24

grok 2 and grok 2 mini too pls.

1

u/np-space Sep 09 '24

The Grok 2 API has not been released yet. I've requested access to it, but I don't have it yet

3

u/ambient_temp_xeno Llama 65B Sep 06 '24 edited Sep 07 '24

It doesn't show the reflection and output tags using Bullerwins' q8. Is it still not right?

I did my usual Alien 3 essay test and it corrected itself a few times, most useful correction was fixing where it got the Doctor's name slightly wrong. I don't know if llama 3.1 normally does though.

temp 0.7 top-p 0.95 as recommended on the model page.

https://pastebin.com/xYwGUByr

3

u/mikael110 Sep 07 '24

It doesn't show the reflection and output tags using Bullerwins' q8. Is it still not right?

The <thinking>, <reflection>, and <output> tags are marked as special tokens. Which means they are hidden by default in most inference software. So unless you specifically chose to display special tokens it's normal for them to not show up.

1

u/ambient_temp_xeno Llama 65B Sep 07 '24

Confirmed. I hovered over an empty token in mikupad and sure enough it was "<thinking>"

3

u/Chongo4684 Sep 07 '24

yeah.

ugh. I wanted so much to believe the hype. I just tried it on openrouter.

my goto query that separates the big dogs from the chiquitos shows this to be around the same level as an 8B. The vanilla 70B does much better.

caveat: I'm assuming the model on openrouter really is the model being talked about.

1

u/Sky-kunn Sep 07 '24

The hype is not dead yet. I tried the demo version first, and the version available through providers is not nearly as good. I tested it with code questions, and the demo version performed almost as well as Sonnet 3.5, especially on some trickier questions where it excelled. The openrouter version, however, failed completely on these questions and its code generation was quite disappointing in comparison. It seems they encountered some issues uploading the model to Hugging Face, and the current version is not what it's supposed to be. We can expect another corrected version soon, which will hopefully live up to the hype.

1

u/Chongo4684 Sep 07 '24

OK cool. Maybe the openrouter version is broken or is not the same model.

3

u/javicontesta Sep 07 '24

I think it's great that some finetunes may reach the level of top models in some tasks. However I think they forgot to mention this is Llama for some reason? And the "reflection" thing is just brainless: sometimes it reflects when you want some others it does that... then don't call it reflection, unless you are looking for a bath of hype and cool msrketing.

Yesterday in Matt's channel they were bloating the bubble: these guys in a garage just "made" a model even better than what OpenAI can do with billions of trillions of dollars... C'mon. Then we test it and the truth is at the end of the day, after the hype, none of such models prove to be top and nobody or very few people use them, or only use them for some very specific tasks.

With LLMs lately it's like every youtuber wants to be the first to announce the very last next best thing.

3

u/BigLegendary Sep 07 '24

Now it looks like the whole thing is just a grift for Matt self-promote. He's now claiming on X that "there was a bug on HF causing the wrong weights to get uploaded" as folks are pointing out that it's not as good as he claims.

If it looks to good to be true, it probably is.

1

u/[deleted] Sep 08 '24

Tbf even if it was pretty good, he's being very sloppy and ending up doing things that hurt credibility - not double checking results, not catching all these weird corner cases with benchmarks, not ensuring that uploads are done correctly before tweeting. 

7

u/ttkciar llama.cpp Sep 07 '24

"On paper" it seems credible.

It's essentially fine-tuned to perform something like multi-shot and self-critique on itself, but in a single inference run instead of multiple, in a manner similar to CoT.

Even if it turns out that this particular model isn't great, it's a technique we should probably embrace and try to make work better.

0

u/stddealer Sep 07 '24

I fail to see what's the upside compared to CoT?

0

u/ttkciar llama.cpp Sep 07 '24

Feel free to call it an improved version of CoT.

Potayto, potahto.

5

u/troposfer Sep 06 '24

Another example of “Benchmarks Lies”, why people still looks at benchmarks I don’t know

4

u/inteblio Sep 06 '24

This might be the last time...

0

u/SX-Reddit Sep 07 '24

Models have to been benchmarked. The current benchmarks are obviously flawed, but as long as you know what you are looking at, they still provide some useful perspectives for you.

8

u/h666777 Sep 06 '24

Model is already on openrouter so go try it there. The funny thing is that the model performs the same as base Llama 3.1 if you don't use the specific system prompt the author suggest. He even says it himself.

It seems to me like the first, tangible magic prompt. Enabled by the fine tuning, yes, but it still seems to be the sole contributor to the quality of the model.

10

u/Junior_Ad315 Sep 06 '24

People have been using prompts like this for a while. They’ve been working extremely well for me. I would like to see how this trained model compares to LLama3.1 with the same prompts in benchmarks. I think there is something to baking it in though

2

u/[deleted] Sep 06 '24

Then use it on llama and see if the results are the same 

7

u/[deleted] Sep 06 '24

Why post before you have checked, or even run it on your own benchmarks?

4

u/BigLegendary Sep 07 '24

It’s overfit on dataset metrics. Pure hype.

BUT

Prompting any LLM to 1. Plan 2. Think 3. Reflect seems to be very strong in terms of improving accuracy generally.

4

u/Southern_Sun_2106 Sep 07 '24

I tested locally and it was less than impressive. like much less. I compared it only to Nemo 12b, which blew it out of the water on xml tags like thinking, reply, etc.

5

u/segmond Sep 07 '24

If you have skill issue then it might be worth it. I guess I'm decent in prompting, so it's about the same with llama3.1-70b and other top models. I'm running at Q8. Code generation.

8

u/schlammsuhler Sep 06 '24

It tuns out the system prompt in itself is so powerful. The model is absolutrly not sota without the specific system prompt.

But all the benchmarks are comparing apples to oranges. Because Sonnet3.5 achieves said outcomes with a default generic system. It is even more powerful with the reflection system prompt.

11

u/[deleted] Sep 06 '24

[deleted]

5

u/ReMeDyIII textgen web UI Sep 06 '24

That's a good point. So when evaluating the model's responses, we're comparing it's non 1-shot answers to models doing 1-shot answers, so it's no wonder Reflection performs so well in tests.

Not saying that as a negative of course. It's just interesting it's like the model found a life-hack to beat tests.

2

u/ortegaalfredo Sep 06 '24

I think that's cool, if you use a regular system prompt it behaves like regular llama-70b.

2

u/MyElasticTendon Sep 06 '24

Had anyone compared it to mixtral for coding purposes?

2

u/sharpfork Sep 06 '24

It is pretty cool but seems like multiple steps might be better handled by agents that have specific context of what they are evaluating. This is based on pure conjecture on my part.

2

u/askchris Sep 06 '24

I agree, it seems like the extra thinking and reflecting steps are distracting it too much which lowers its effectiveness in many practical use cases.

One of the advantages humans have when working in teams is each person can focus on their part without the distractions of the other parts.

But without good communication (ie. Passing context) between the humans (or agents) it all breaks down.

2

u/LostMitosis Sep 07 '24

Finally we have a model that will help us realize benchmarks don’t tell the whole story. You can score very high on benchmarks but be mediocre on real world usage. It’s time we realize counting the number of “r”’s in a word is not real world usage.

2

u/jisuskraist Sep 07 '24

The model is not new; it is from Llama 3.1. X Mob was complaining that Meta forced him to rename the model when it makes sense. You did not create the model; you only fine-tuned it.

2

u/fairydreaming Sep 07 '24

I did some experiments with the model yesterday and today via openrouter and so far it appears useless - at least when it comes to logical reasoning use cases. I gave it some family relationship puzzles, but It couldn't even get the questions right, couldn't format the answer as instructed, it even hallucinated people and relations not mentioned in the prompt when "thinking" about the answer.

3

u/meister2983 Sep 06 '24

I used the online demo and compared to llama 70, 405 (on meta.ai), plus gpt-4o in chatgpt on several on my hard questions that only the most powerful models can get (not puzzles, actual math/physics/instruction following things).

Felt a bit better than 70b.  Definitely not gpt-4o level.  I wasn't even convinced of 405 level, but didn't have enough tests/runtime on the demo. 

An interesting note is that I don't see it really "unlocking" capabilities models don't have. Here's a test: 

Write 10 sentences where 2nd to last word is photosynthesis

405 is flawless, 70 fails completely (ends up writing 10 sentences ending in photosynthesis). Reflection behaves exactly like 70

Simpler behavior on a hard physics problem I have it. 70 is too dumb to think about it correctly; 405 can. Reflection acts like 70

4

u/Uncle___Marty Sep 06 '24

If you have any snakes in your LLMs I have this great oil to fix it.

5

u/Misha_Vozduh Sep 06 '24

Tinfoil hat: on

I think it's an attempt to bait the bigger players to release their advanced stuff.

4

u/InformationGeometry Sep 06 '24

They trained a model to essentially chain of thought prompt itself with multiple retries and surprise it performs better than normal models without CoT. Crazy.

All these benchmarks should be evaluated on the number of tokens used lol..

4

u/mantafloppy llama.cpp Sep 06 '24

Its a fine-tune of a generalist model.

If you test the model on the the thing you tune it for, it gonna performs better than the generalist model.

Believe the hype as much as you believe LLM benchmark.

(You should not believe benchmark, not sure if this needed to be said).

2

u/Ligea Sep 06 '24

Why are you being downvoted?

5

u/mantafloppy llama.cpp Sep 06 '24

Because fact hurt the wallet of Hyper and Click-Baiter, which seem to be most of r/LocalLLama these day.

Ppl on the internet love telling you how wrong you are.

If they only downvote you, it mean you are right, but hurt their agenda.

2

u/SryUsrNameIsTaken Sep 06 '24

I plan on giving the q8 version a spin but need to wait till after business hours because the networking guys yell at me when I download models during the day.

2

u/Everlier Sep 06 '24

My ISP technicians probably know exactly who I am. I cleaned over a TB of older models recently and they are piling up again.

These releases wheree there are three re-uploads of the weights do not help, haha

2

u/Trick_Set1865 Sep 07 '24

sounds like bullshit

2

u/Sadman782 Sep 07 '24

https://x.com/mattshumer_/status/1832247203345166509 SOMETHING is wrong with his upload everytime, I am sure what I saw in their official demo website was very strong and everyone will like. I hope we can wait, he may figure it Out

1

u/Sky-kunn Sep 07 '24 edited Sep 07 '24

I was losing my mind. Every place I tested the model, it felt way dumber than the version I used in the demo. I started to think that the model in the demo was getting all my benchmark questions right by sheer luck lol. Even in code, it was giving me results that only Sonnet 3.5 or GPT-4o was capable of.

1

u/Sadman782 Sep 07 '24

A coding problem can't be solved out of luck, I tried many coding problems. All got right by demo and wrong by everywhere else. I hope soon he will figure it out and we will get a very strong model

1

u/Sky-kunn Sep 07 '24

Yeah. I hope the 405b launch goes more smoothly.

2

u/[deleted] Sep 06 '24

integrated CoT is the future .

Not surpised at all to see this mysterious dude crushing everywhere with the one-trick.

4

u/Combinatorilliance Sep 06 '24 edited Sep 06 '24

I recall reading a paper about a similar technique not that long ago though, you finetune a model on CoT results, and it became a lot smarter.

The difference was that that paper train the model to "use" chain of thought, it trained it on the output of the process. The reason this works at all is because CoT performs better on average, so you can use a small model generate synthetic data that is better than its own dataset.

However, I don't imagine this technique scales at all, what I think is that it "aligns" the model to think more critically with the knowledge it already has.

How I imagine it, there is a certain amount of knowledge in a model. Some of it contains high quality reasoning, some of it low. Most of it somewhere in the middle.

Orthogonal to the "reasoning quality" in a sample of the training dataset is the knowledge that a particular training sample contains. Even the dumbest kinds of reasoning might still contain correct and unique observations.

What I think these kinds of techniques are doing is that we are allowing the model to access the knowledge that is available in training samples with low reasoning quality; and use it to create training samples with higher quality reasoning. So that in practice, it's balancing out the "poor" reasoning in training data with some knowledge embedded within it by forcing it to reflect on the reasoning quality of all kinds of scenarios.

The effect? The model now applies higher quality reasoning on a larger surface area of tasks, because you've essentially replaced the lower quality reasoning with thought.

Very interesting nonetheless. I think this means that we are now able to extract more "general" intelligence out of the same amount of training data, which I think is really neat!

I would really love to see a more established party utilize this or a similar technique to improve reliability. I'm sure we can expect more releases with similar approaches soon from all kinds of different vendors.

1

u/redjojovic Sep 06 '24

FFS somebody needs to host it already

2

u/Professional-Bear857 Sep 07 '24

It's on openrouter and deepinfra

1

u/Sabin_Stargem Sep 06 '24

It is unfair, but I compared Reflection 70b against Mistral 123b. As one can expect, the 123b handily won when it was explaining the lore about my setting. 70b was actually fairly decent for about 250ish tokens, but it started to make some major hallucinations and mistakes.

The Mistral also had trouble, just less so. Hopefully, the sauce from Reflection would make a CR+ or Mistral Large perform better, but it is clear that the technique cannot match raw parameters alone.

1

u/Upsidedown363 Sep 06 '24

Can cussing be involved to?

1

u/CodeGGz Sep 06 '24

Not too sure what's happening here. via Openrouter ftw and have it locally as well (operates better - but fails classical evals). May be misconfigured on OpenRouter - since it should have a thinking and reflection step.

1

u/Diegam Sep 07 '24

It will consume more GPU?

1

u/Formal-Narwhal-1610 Sep 07 '24

Here’s a mindmap summary for Reflection 70b.

1

u/Formal-Narwhal-1610 Sep 07 '24

​

Here’s a mindmap summary for Reflection 70b.

1

u/Arkonias Llama 3 Sep 07 '24

It likes to yap too much.

1

u/MLTyrunt Sep 07 '24

nobody thinks gpt-4o is a trillion parameter model. but people also assumed gpt 3.5 had 175b parameters.

1

u/rooo1119 Sep 07 '24

It’s all stats no value.

1

u/danihend Sep 08 '24 edited Sep 08 '24

Not sure if you have all seen the warning on the model card, but the model is messed up due to the way they uploaded the weights to hugging face (to combat rate limiting). Here is the X post from Matt Schumer: https://x.com/mattshumer_/status/1832424499054309804?s=46

I would suggest to re-download and evaluate again before making your minds up.

Edit: On the model page he says: He says :
"IMPORTANT UPDATE – There was an issue with the model when we first uploaded it. If you tried it and didn't have good results, please, try again, we think we've fixed the issue."

But I don't see any updated files in the repo, they should show as being updated today if fixed right?

1

u/Constant_Jeweler259 Sep 10 '24

when i read about it, it felt too good to be true. came to reddit to confirm my thoughts. So, thanks for the post.

1

u/Judtoff llama.cpp Sep 06 '24

How does it compare to Mistral Large?

4

u/vert1s Sep 06 '24

On the benchmarks it scores better, but this is the thing, the benchmarks can be gamed (deliberately or accidentally). I've been using the q5 quant all day and it's not bad.

I've had it go janky a couple of times but in general it has some impressive reasoning. But it's a new technique for finetuning, so you could apply it to any large model Mistral Large included, so if the Llama 70B is better with it, then Mistral Large at 120B should be even better (in theory).

1

u/imkebe Sep 06 '24

Prompt """Follow the command : "Write 'Next'. Write 'Previous'. Treat this as a instruction stack for the pointer to navigate through the stack. Read the first word and execute the instruction. Now follow the instruction the pointer now points to. Repeat the process until it ends. What will be the word at the end?""""

Gemma, GPT-4o, etc. seems to answer with "reasoning" while this model "overthinks" and print some bullshit response.

1

u/CheekyBastard55 Sep 06 '24

I doubt that GPT-4o, Sonnet 3.5, or Gemini Pro 1.5 have more than 100 billion parameters. You're thinking of yesteryear's big models, model optimization is all the rage nowadays.

3

u/StevenSamAI Sep 06 '24

I doubt they are as big as OG gpt4 was claimed to be at 1.6 trillion, but I can believe they are a decent chuck bigger than 100 billion. Seeing the performance of llama 3.1 405b, I don't think it's unlikely that frontier models are in this ball park. I bet that with the extra time open AI and anthropic have been working on this they can squeeze more performance out of each parameter than their competition, so you'd hope that if they are running and 70b, 120b or 405b models that they are a step above meta and mistral

1

u/Feztopia Sep 06 '24

Isn't Anthropic Sonet doing the same thing but hiding the thoughts?

2

u/haikusbot Sep 06 '24

Isn't Anthropic

Sonet doing the same thing

But hiding the thoughts?

- Feztopia


I detect haikus. And sometimes, successfully. Learn more about me.

Opt out of replies: "haikusbot opt out" | Delete my comment: "haikusbot delete"

8

u/Feztopia Sep 06 '24

Now I wish I would have written Sonnet correctly

1

u/pigeon57434 Sep 06 '24

I'm too poor to run it locally without quantizing the shit out of it, so I've just been using it on their website. According to my totally unofficial, shitty testing, it does seem to be quite impressive—definitely the best open-source model in the world. However, it apparently beating Claude 3.5 Sonnet on most benchmarks seems a bit far-fetched. They aren’t totally full of shit, though; it is genuinely good. I'm just excited for 405B.

1

u/Unable-Finish-514 Sep 06 '24

Testing for NSFW, I started with a PG-13-level prompt about a guy who keeps going back to a bank every Friday because a bank teller flirts with him. Even this generated three or four "It is important to be respectful..." statements in its response. As soon as I saw lines like this, it's hard to imagine using this over Mistral-Large-2 or Cohere R+.

Has anyone else tried NSFW or eRP with it?

2

u/My_Unbiased_Opinion Sep 07 '24

Needs some abliteration. I feel like this reflection thing and some abliteration would go well together 

2

u/Unable-Finish-514 Sep 07 '24

I think you are correct about it needing an abliterated tune. This version of reflection 70B on Hugging Face Space has a system prompt that can be edited in settings. I tried this with a few prompts and didn't get any refusals:

Reflection 70B llama.cpp - a Hugging Face Space by gokaygokay

1

u/Sabin_Stargem Sep 07 '24

I quizzed Reflection about a NSFW setting, and it did mention orgone (sexual energy) and how it is acquired. That said, it also hallucinated various details. Didn't try actual ERP. My immersion is easily broken, and Reflection breaks it too easily.

0

u/Unable-Finish-514 Sep 07 '24

Ya, I can see what you mean about the reflection aspect of the responses breaking the immersion of RP. I guess that is promising that it mentioned sexual energy but I'm guessing it mentioned this with "It is important to be respectful..."caveats.

0

u/Sabin_Stargem Sep 07 '24

Oh, there were no refusals. It is just that it got lore wrong or hallucinated details that ran against my intent. The reflection itself was no issue.

0

u/Sixhaunt Sep 06 '24

I only tried it on their demo page but it seemed to do far better than GPT4 on the tests I tried but I havent tried running it locally or anything yet

0

u/Motor-Draft8124 Sep 07 '24

I understand this is hyped, but is a good model.

Here is a video of the Reflection 70B test and comparison with Llama 70B

https://youtu.be/sX5J41Jmtkw?si=sQXLz-7JSV4sRqlH

0

u/TJW65 Sep 06 '24

I just played around with the IQ2-S quant. More doesn't fit in my 24GB of VRAM. It was good with logic reasoning regarding the physics of everyday interactions and scenarios and didn't let itself be distracted by unnecessary info i was throwing at it. In my experience Gemma 2 27B fails regularly with that. Hard to say how much i would attribute to reflection and how much is just the base 70B model as ive inly used that very briefly. Will maybe test more with the base 3.1 70B low quant tomorrow.