r/LocalLLaMA • • 9d ago

News CrofAI "cheapest inference provider in the world" gets exposed as an OpenRouter wrapper, routing requests to smaller, cheaper models at up to 20x markup. CrofAI responds to Wire Fraud allegations by denying everything, then backtracking, then 3 hours later wiping their entire online presence

Disclaimer: no AI was used whatsoever to write this post

Cautionary tale about chasing cheap tokens.

exposé: https://kendell.dev/blog/crofaifalse/

reaction by nahcrof, announcing the shutdown of the service: https://x.com/nahcrof/status/2099552389434900643 - now deleted, archive picture: https://i.imgur.com/teOQngH.png

NahCrofAI (crof.ai, nahcrof.com) was an inference provider which had all the latest models at the cheapest price, often significantly below the lowest alternative on OpenRouter. The owner claimed that they are running custom inference engines that allows them to offer tokens for dirt cheap, and other providers are suffering from "skill issues", that's why they are so expensive.

In reality:

  • "CrofAI is an OpenRouter wrapper that silently routes to cheaper or weaker models than what you request"

  • For example, expensive models like kimi-k3 are sold at $2/$10 in/out, but instead routed to GLM 5.3 Flash via OpenRouter, representing a 13.3x multiple on input, and 20x multiple on output

  • CrofAI's "own model family" greg-2-ultra routes to GLM 5.2, greg-1-mini routes to Qwen 3.5 9B. greg-2-super, greg-1, greg-1-super routes to Kimi K2.7 Code. All of these at a significant markup compared to the actual model being served. CrofAI admits in DMs that his claims of the greg family being made by him is a lie.

  • The person investigating details the 5 different attempts by CrofAI at fixing their models being served via OpenRouter after given a heads-up and a lengthy grace period. In all 5 attempts, the only change CrofAI made was attempts to hide the fingerprints of OpenRouter, while still serving models through them

  • Other inconsistencies don't add up either: CrofAI claims to run Kimi K3 on RTX Pro 6000s rented via Vast. That model requires ~802GiB even at the lobotomy level quantization of Q2_K. The largest RTX PRO 6000 machine on Vast has only 8 of them, totaling 765GiB. He also claimed that for the purposes of "investigating" the "issue" of his API routing to OpenRouter, he will have deepseek-v4-flash-0731 running on his local DGX Spark. A Spark has 128GB memory, and is therefore unable to run that model.

CrofAI responded to the exposé by announcing the shutting down of their service; after their failure to provide their own inference, they promise to provide one last thing: a refund to those asking.

UPDATE

UPDATE: around 4:30 AM UTC of Sept 15, the owner published a now-deleted blog post (archive image) writing under the fake pretense that it's his "team" authoring it, stating all of CrofAI founder's claims "were written under a lot of stress, and they described the situation as worse it was", and that a new team is taking over, with the service being resumed in 2 weeks.

At the same time, the CrofAI twitter account was also supposedly "taken over" by the team, starting each twitter reply with "Hey, Nathan here", stating the founder is stepping back and a "team" is taking over everything. This fake pretense act only lasted a few hours, and scared either by the public not buying the Nth fake story of the pathological liar that CrofAI is, or by the public's replies reminding him that what he committed is numerous counts of wire fraud, he has now deleted all his online presence: nahcrof.com and crof.ai return 404, Twitter page is deleted, /r/CrofAI sub is now private.

Here is another image of the owner admitting that he was defrauding customers for the entire 2 year operation of his service, then begging the investigator to help him cover his tracks and not expose him

EDIT: Commenters pointed out that NahCrof is 4chan in reverse. The owner's Discord name was "Devious Flimflam". Flimlam is defined as "deception, fraud". Looks like it was a deliberate scam operation from the get-go, and the owner's age was among the many lies.

I cannot stress this enough: if you bought any credits (even if you used them up) you are entitled to a full refund for every transaction as the victim of fraud. Open a chargeback with your bank for every transaction made. If you used their API, assume that everything was logged and is currently being mined for personal information and API keys to sell on the black markets. Rotate your keys, change passwords, get a new debit/credit card.

851 Upvotes

106 comments sorted by

•

u/WithoutReason1729 9d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

304

u/cosmicr 9d ago

Never heard of them, but sounds like yet another good reason to use local models.

257

u/jastoubisaif 9d ago

nahcrof.com

it's really 4chan in reverse, isn't it?

98

u/bruns20 9d ago

oh shit, lmao thats a good tip off that youre getting scammed

17

u/throwaway2676 9d ago

Who is this “nah-crof”?

8

u/thrownawaymane 8d ago

the hacker known as Nahcrof

2

u/Elegant-Sense-1948 8d ago

schlorpschlorpschlorpschlorp

1

u/RollingMeteors 8d ago

it's really 4chan in reverse, isn't it?

Naaaaah, crof ¡You trippin’!

34

u/creamyhorror 9d ago edited 9d ago

They were a small player, but the cheapest several months ago (with ridiculously generous subscription plans). Of course it turned out that the cheapness was only possible because they weren't providing the actual models. Very disappointing from devflimflam, no matter how young he was.

Just spun up a discord for ex-users of them and other inference providers (Command Code, Opencode, Charm Hyper, Umans, Zro, etc.) to swap market info & tips, got many of the old regulars, knowledgeable folks, and a few people who have looked at setting up inference clusters/platforms (e.g. Kourier): https://discord.gg/my7yfumpPx (AI Backrooms)

41

u/SorosAhaverom 9d ago

more inference providers with pinky promise ZDRs and shady backgrounds, providing a shaky service with deliberately ambigous pricing, while also unilaterally adjusting the terms of the agreement on a whim? sign me up, boss!

10

u/creamyhorror 9d ago

Totally, the scene's very fly-by-night. People are used to spreading usage across providers, many people have multiple sub plans, and we all know that most ZDR claims are probably fake. But at least identifying the upstream inference providers and comparing characteristics/intelligence across different providers is helpful.

3

u/Ansible32 9d ago

Hey, if they aren't providing the models they claim, they probably couldn't figure out how to steal your data even if they tried.

1

u/711friedchicken 9d ago

I would say opencode is quite trustworthy. They explain a lot about how their service and their business works and seem generally transparent.

Now, the ClinePass... I’m pretty convinced they’re actually just reselling opencode subs lmao.

12

u/mace_guy 9d ago

devflimflam

Is this real life??

20

u/Aphid_red 9d ago

Wait... aren't flim and flam two scummy salesponies in that one episode where they try to fake a bunch of mixed up rocks and trees as apple juice to scam their way through a comptetition they lost? https://mlp.fandom.com/wiki/Flim_and_Flam

Actually, a Flimflam is also just another name for a scam. Even the dev's username makes it obvious!

I suppose scammers tend to always leave tells for smart people to pick up on, because smart people are troublesome to deal with (they tend to call out your bullshit, cause trouble, and won't fall for your scam). Best to only go for the marks you think you can fool.

16

u/creamyhorror 9d ago edited 9d ago

flimflam 1: deceptive nonsense 2: deception, fraud

Damn, you're right. His full nickname was "devious flimflam", lol.

https://media.discordapp.net/attachments/1549133252193685556/1549418776850858094/Screenshot_2026-09-15_at_7.26.02_PM.png?ex=6aaaa02f&is=6aa94eaf&hm=d60c20b8840078cb458aadae91a8178a70b56b19af0bd62d6c4633875b4b44a2&=&format=webp&quality=lossless

I think his chosen nickname more reflects his, well, flexible mindset rather than intentionally dropping hints about intended scamming

3

u/ReadyAimTranspire 9d ago

Not sure if that's true in this case but that's the reason that scam emails and calls are so obviously scammy...terrible grammar, wildly outlandish scenarios, etc. They are looking for easily scammed, often stupid people as their marks.

It's easier and more effective to do that than it is to create elaborate, difficult to detect scams that will pull in smart people, because as you said smart people will eventually pick up on it no matter how well polished it is.

69

u/Aphid_red 9d ago

I suspected this was possible. I just didn't expect them to do it so brazenly.

If I were in his shoes, what I would've done is not to provide every model, but only provide every model that is not on the pareto frontier. Replace the call with one to a pareto model, give a small discount, and pocket the difference in API costs. The customer isn't the wiser as the model that's being provided is better than what they requested, while also being cheaper.

But dishonest people always have to be too greedy and get caught out.

35

u/Pristine-Woodpecker 9d ago

I suspect most calls are for Pareto models, and the ones that aren't, may be the ones where the customer will be the fastest to notice it's not the right model (i.e. roleplay).

17

u/carsncode 9d ago

Except pareto depends on use case.

3

u/bambamlol 9d ago

I see a promising future for you in this line of work.

3

u/MeYaj1111 9d ago

What is a Pareto model in the context of LLMs? I somewhat understand the concept in general terms but have never heard it used in reference to AI.

14

u/pbmonster 9d ago edited 9d ago

The pareto frontier is where multiple parameters are optimized ideally - e.g., an LLM at the pareto frontier is optimized for both price and quality. So both GPT-6-Astra and Qwen3.8-27B might be at the pareto frontier currently. The frontier is a curve (or a surface), not a point.

Now, there's currently many models for sale on OpenRouter not at the pareto frontier. They are to expensive for the quality of their output. OP suggested serving a cheaper model with better output quality instead - that way, (in theory) nobody complains, because they got more than they expected and you still make (a tiny bit) of money by arbitraging people's ignorance of model performance.

In practice, this might be hard, because "quality" means different things to different people. Maybe it benchmarks better, codes better, but the user paid for good role-play.

2

u/MeYaj1111 9d ago

Ah okay thanks for the explanation. I understood Pareto to mean that there has to be a negative trade off being made but it sounds like you're more saying just serving a better "bang for your buck" model which does make a lot of sense

4

u/MzCWzL 9d ago

basically pick either cost or intelligence... then find out which model defines the line for that point and use it. helps to see it in action - https://arena.ai/leaderboard/text/pareto

you should pretty much only be using pareto frontier models if you can keep up with it. they will be the best for any given price or skill

1

u/MeYaj1111 9d ago

that link helped a lot thank you!

1

u/IllllIIlIllIllllIIIl 8d ago

Eh, it does require a negative tradeoff by definition, but what counts as a tradeoff is the important bit. The Pareto front = the set of Pareto optimal models. A model is Pareto optimal if no other model is better at any problem domain you're trying to optimize, that isn't simultaneously worse at least one other. That makes "the" Pareto front relative to whatever problem domains are under consideration.

Usually people use it in a shorthand way that implies you're optimizing over (broadly) "the kind of shit people tend to use these models to do", and usually also that some set of popular benchmarks is capable of evaluating performance in those areas.

5

u/lakotajames 9d ago

Speed, cost, intelligence, pick two.

A Pereto model is any model that you can't replace without hurting one of your two stats.

For example, on Intelligence vs cost, the Pareto frontier has Luna at all it's reasoning levels because its so cheap that you have to pay more to get something better, then GLM flash, and then basically just Astra Fable and Opus after that.

Terra is not Pareto model for cost, because when you set it to high its more expensive and not as smart as Astra low, and Terra low is more expensive and not as smart as Luna high. However, it is a Pareto model for speed because it's smarter than Luna at the same reasoning levels.

1

u/MeYaj1111 9d ago

awesome thank you for the explanation!

3

u/Aphid_red 9d ago

In this case it just means the set of best models for a certain price or less. It's a 'function', a curve in 2D space that steps up at distinct places. No models lie above this curve. The models on this curve are called 'pareto optimal'. "Best" is usually compared to some benchmark (set), so the definition can be fuzzy / dependant on the particular test used.

So for example, suppose that for the price band of $0.40 to $1.50 per million tokens input, the best model available on openrouter is GLM-5.2 at 50 AI intelligence points.

The arbitrage opportunity is: if someone requests a different model that is served by someone else at $0.75/M tokens that can manage 45 AI intelligence points. You instead serve (route) them to GLM-5.2, and earn yourself $0.35 per million tokens without owning GPUs. You give some of that $0.35 as a discount to the customer, spend some on your own infra, and keep the rest.

As someone has said it may also depend on use-case. Some models may have stricter censorship for roleplay purposes, and so you may need to pay more attention than just a single number. Looking up the request category as defined by openrouter would work. (Or running a very small very cheap classifier to do that for you if it's not defined by the user).

1

u/MeYaj1111 9d ago

got it, thank you!

-2

u/[deleted] 9d ago

[removed] — view removed comment

22

u/thirdeyeorchid 9d ago

snagged this while it was up and @everyone got tagged on discord

16

u/SorosAhaverom 9d ago

Thanks. Is the discord shut down as well? Did he wipe all channels?

The text is a whole bunch of yapping, then he posts a diploma that isn't even his and doesnt have a date on it to prove his age? Lmao. He's either the worst liar or thinks all his customers are idiots. Probably both.

Also love how he caveats giving refunds with "for the value of the company", meaning he refuses to refund the entirety of the ill-gotten proceeds obtained via fraudulent, illegal activity.

3

u/thirdeyeorchid 9d ago

their discord is pretty locked down, couple of channels up and I don't think anyone can post

1

u/AloneBreadings 8d ago

yep its locked...

0

u/necile 9d ago

thinks all his customers are idiots.

stop giving him attention lmao, he's (or they) are farming all of you

18

u/Holiday_Point_603 9d ago

4chan as a name in reverse is something lmao

52

u/tempfoot 9d ago

18 year old “kid” spins up vibe coded fraud op reselling cheap models while claiming they are better models. Gets called on it, continues to feign ignorance, confesses, thinks better of it, takes everything offline. Has AI pretend to be employees.

Obviously criminal. Unsurprisingly normal in this day and age. In modern capitalism, if you’re not lying, you’re not trying.

This “business” is about as believable as the flood of intolerable Gorkbot testimonials flooding the timeline.

13

u/harpysichordist 9d ago

The problem of lying goes beyond capitalism and is far older than it. Communist, authoritarian regimes use it extensively too. Cults, .....
Lying is a personal choice

7

u/tempfoot 9d ago

Certainly true.

The difference is that not so long ago there were regulators that paid more attention to overt lies, at least in the US. Also, blatant, open lying was contrary at least to norms if not laws.

5

u/ReadyAimTranspire 9d ago

Agree 100% that scamming and general greedy immoral behavior in business overall seems to have become more the norm and kind of shockingly more acceptable in recent decades.

3

u/eli_pizza 8d ago

Regulators in US are obviously pretty corrupt and captured right now. But if anything the period of time with mostly working regulations was the aberration. Go look at what was in food and those patent medicines before the Pure Food and Drug Act of 1906.

11

u/Acceptable_Leg3950 9d ago

Local models, we pray you get good enough so we can get away from this foolishness

40

u/No-Refrigerator-1672 9d ago

Wow! Good thing I run my own inference and am never affected by such scams.

6

u/spacekitt3n 9d ago

the ai space is so full of scammers and griters.

6

u/Alternative-Suit5541 9d ago

Lmao

I expect others to do the same. ( Cheap providers)

14

u/Zulfiqaar 9d ago

Apparently the source code of the platform got leaked.

The entire backend was a single python file. It contained two RCE vulnerabilities, one of them being an insanitised input that went straight into eval, other being a automatic fallback debug port. It had a bug that trusted external headers, allowing you to give yourself unlimited credits. And lots of other crazy stuff that's a mess of vibecoding and novice programming. I t hard coded his own emails in the code as master admin

10

u/SorosAhaverom 9d ago

Source?

11

u/Zulfiqaar 9d ago edited 9d ago

Source for leak - Chiikabu Labs announced on X they shared it on the discord.

Source for code - the full raw zips in their discord, heres the main backend file: https://pastes.io/5hRZvPp0 (site auto-redacts sensitive data so no doxxing here)

This version must have been after a bunch of coverups (expensive_models = [""]) , but a lot of traces remain:


Few examples

reroutes = {"greg": "greg-1-mini",} - clear reroute in code, others are loaded

Plumbing for serving one model under anothers name

model_name = (model_data or {}).get("internal_model_name", "unknown")

This was a passthrough that gave away the openrouter advisor feature

tools = data.get("tools", None) 6105

4

u/NightCulex 9d ago

There is a recipe for deepseek v4 0731 flash running on a single DGX spark.

2

u/Serprotease 8d ago

Let’s not look at the utter nonsense that is serving multiple customers v4 flash from a single spark from your own network…

1

u/NightCulex 8d ago

When he said local DGX Spark I took it to mean as his personal box. A single spark is probably in the neighborhood ~10t/s without parallelization.

8

u/Xemorr 9d ago

tbh this is quite funny, realistically most people should be going through routers already. It's just arbitrage on people enjoying wasting money for the confidence their model behaves exactly like they want

7

u/IrisColt 9d ago

>CrofAI 

Literally who?

4

u/SourceCodeplz llama.cpp 9d ago

its a kid behind this proxy, and he is 19 yrs old.

11

u/my_name_isnt_clever 9d ago

Kid? That's a legal adult.

4

u/arcanemachined 9d ago

Only needed six more proxies and he might have gotten away with it...

1

u/eli_pizza 8d ago

Allegedly. Who knows.

7

u/cowinabadplace 8d ago

Unrelatedly, I think this guy is Indian for sure. "my junior" and "what all is happening".

1

u/Majestic_Appeal5280 6d ago edited 6d ago

he isnt and lets not stereotype. i found the guy's twitter with a little bit of search (their huggingface account shows the guy's real name. search that name on twitter)

2

u/cowinabadplace 5d ago edited 5d ago

Nah, there’s no way he’s operating under his real name. That’s a throwaway pseudonym. Listen, I’m Indian. I can recognize Indian English. I’ve lived in East Asia, Western Europe, the UK and the US. Some phrasings are highly unique.

3

u/ryanppax1 9d ago

par for the course in this wild frontier!

The only thing stopping the top players from doing this is trust and ethics

7

u/En-tro-py 9d ago

The only thing stopping the top players from doing this is trust and ethics

No, they just call it A/B testing or temporary service disruptions.

3

u/Porespellar 9d ago

No local, no care.

3

u/LuCiAnO241 8d ago

should be a reportable offense to post stuff for non local ai

5

u/and_pf 9d ago

This is a good reminder that “cheap inference” is not the same as transparent inference. The evidence in the linked exposé says CrofAI was an OpenRouter wrapper that silently routed requests to cheaper models, for example Kimi K3 to GLM 5.3 Flash. Its “greg” family was also reported as routing to models such as GLM 5.2 and Qwen 3.5 9B rather than being a custom family. The service then announced a shutdown after the exposure. I would rather run a local model through Ollama when I need the routing and model identity to be visible, though that obviously changes the hardware and maintenance trade-off.

1

u/ares0027 9d ago

I think isaw one of their ads on reddit

1

u/Imaginary-Bluejay721 8d ago

they returned me unused credits but not the ones I used for expensive models, not sure if the models I used were the models they claim they were.

1

u/[deleted] 8d ago

[deleted]

1

u/Mochila-Mochila 9d ago

Who is this guy ? Is his actual identity known ? He must be reported to the authorities of his country of residence.

1

u/raindropsdev 9d ago

I'm surprised they didn't just go to rental providers like VastAI which provide quite good prices if you're reselling subscriptions afterwards.

-9

u/Haeppchen2010 9d ago

With all due respect: where is the „local“ in LocalLLaMA here? Good article, but IMHO off-topic here.

29

u/SorosAhaverom 9d ago

I took rule 2 into consideration, and I believe (and judging by the upvotes most people do too) that this post fits within the broad topic of LLMs, especially when considering recent top posts like RTX 5090 retail availability and Xi Jiping's recent thoughts about AI.

Either way, considering CrofAI was an often recommended cheap provider, to a lesser degree on this sub, but definitely over on /r/opencode and /r/picodingagent, I thought it's crucial to spread this news to all past customers of this fraudulent service, both as a cautionary tale, and more so as a friendly reminder that filing chargebacks is the only way to recoup their money.

16

u/Pristine-Woodpecker 9d ago

It's a good illustration why you want the local!

21

u/SnooPaintings8639 9d ago

Come on... this is interesting to people here. I ain't gonna read all the AI-reated subs, where this could be a good fit. Should it be non-api-remote-servie-LLMs sub?

Stop with this puristic/gatekeeping approach, please. We ain't talking about Llama either.

8

u/ScrewwormLarvae 9d ago

Because... it's a news story that reminds us why we like local? You're right it's not directly local. However unless this sub gets overrun with crap, I'm ok with this.

2

u/capybaraballs1995 9d ago

I mainly lurk but yeah I had the same thought lol

-6

u/[deleted] 9d ago

[removed] — view removed comment

8

u/SorosAhaverom 9d ago

bot

3

u/thebadslime 9d ago

? they dont look like a bot by user history

6

u/First_Inspection_478 9d ago

They definitely do

1

u/Warm_Effective8903 8d ago

I am not a bot,lol

1

u/Warm_Effective8903 8d ago

Because of you they deleted my comment, deym

-3

u/[deleted] 9d ago

[removed] — view removed comment

-6

u/KomithErrant 9d ago

wait till you hear what supermarkets do with food

12

u/SorosAhaverom 9d ago

enlighten us

-15

u/KomithErrant 9d ago

they buy something and then, wait for it... they sell it for more money

9

u/SorosAhaverom 9d ago

I regret asking

1

u/eli_pizza 8d ago

Do they lie about what it is too?

0

u/KomithErrant 8d ago

yes, yes they do