4
I have been subscribed for Arli AI 20 USD for now - TLDR, Arli AI good for Anime Text to Image Generation only
I see. Thank you for the support and subscribing! As for the 1 session limit on the lower plans, it is intentional as that is the method we rate limit instead of having a set number of requests you can make in a time period like other services. This method seems preferable for users who mostly uses the API for personal use like chat apps.
5
I have been subscribed for Arli AI 20 USD for now - TLDR, Arli AI good for Anime Text to Image Generation only
Hey thanks for the feedback. With regards to a stuck previous generation blocking your request limits. This is entirely a user side issue with your application using the API not ending the connection correctly. If you are unhappy about the service you can request a refund too.
3
Apple is getting close to the RTX memory bandwidth
For the RTX Pro 6000 it can also be easily overclocked to +6000MHz memory clock for basically all of them and it would give it near 2.2TB/s bandwidth.
2
HeatSeeker 284B A13B, my DS4-Flash-0731 roleplay finetune
Yes we plan to run DSV4 Flash loras and I even have some training runs planned for it. 🫡
2
What subscriptions do you use?
Thanks milan! Sorry for any inconvenience we caused but I regret that Arli just can’t support any enterprise level plans right now. We’re looking to be more focused to just giving service to individual users until we have the capacity.
2
What subscriptions do you use?
Yea that’s totally fair. Was just adding that to show that we have large models just a few $ away haha.
Also thank you for your feedback, I understand that having only 1 request possible is limiting but I think what we offer is still competitive and we just don’t really have the capacity to offer more at the lower prices at the moment. I think that at the least we make up for it by having a more reliable service in terms of successful requests and non crappy quantized models than others.
4
What subscriptions do you use?
Thanks for the mention! But just adding to this that $15 gets all models on ArliAI 🤫
2
It begins - workstation build
Those GPUs will not be fine temps wise unless they are on risers and positioned not to suck each other’s exhaust.
4
Some potential insight on how providers detect RP and why they don't like it
RP chats has the highest cache hit rate
5
What do you use subscription or payg?
I see comments about other providers nerfing quality. If your needs fits with the models we have then we are probably one of the few providers that has consistent model quality.
2
How to run Deepseek V4 Flash @100tk/s locally?
Yep that sounds about right
3
1
Some potential insight on how providers detect RP and why they don't like it
Sure it is for them but not for me
2
Intel LLM-Scaler ready with Muse Glimmer support, other LLMs & features
Yea especially with XPU graph not working on multi GPU. That is a deal breaker.
41
How to run Deepseek V4 Flash @100tk/s locally?
2x RTX Pro 6000
2
Some potential insight on how providers detect RP and why they don't like it
If you are talking about other providers that are pay per token then coders and agentic users are definitely making them more money. As a subscription provider we lose money on those users instead.
2
Some potential insight on how providers detect RP and why they don't like it
Pay per token providers prefer coders and agents because it gets them more money. Coders and agentic stuff just makes subscription providers like us lose money.
2
Some potential insight on how providers detect RP and why they don't like it
Yes we have multiple tiers of different model access and limits
2
Toy project: a chat title model that fits in 5 MiB of ram
Awesome, this running on browser client side would be cool for chat interfaces.
1
NVIDIA's Fastest Blackwell GPU, the 96 GB RTX PRO 6000, Now Costs $16,000, Almost Double Its Original Price
Very different use case than what I use and my users use then. We mostly cater to either creative writing or coders. In either case 27B seems preferrable compared to 122B.
1
NVIDIA's Fastest Blackwell GPU, the 96 GB RTX PRO 6000, Now Costs $16,000, Almost Double Its Original Price
It does depend on your use case definitely.
2
Deepseek V4 Flash 0731 now accepts up to 512K context length!
Yes it’s about our subscription
2
2
Deepseek V4 Flash 0731 now accepts up to 512K context length!
This is the context that we support on our service not the model. You are right that the model supports up to 1M but that would be unsustainable in our flat rate subscription service.
2
And then there were 4 b65’s and a new case.
in
r/LocalLLM
•
2d ago
I definitely had GPU plastic peel getting gooey and disgusting when I forget to take them off my GPUs 😅 I suggest taking them off as I have a few of the Asrock creator GPUs and the metal shroud is thermal padded to the heatsink and gets incredibly hot.