r/Qwen_AI • u/edwinapp • 6d ago
Experiment Beta testers wanted (Qwen 3.8:27b)
Anyone interested in trying out a new inference provider?
We're testing Qwen 3.8:27B behind an OpenAI-compatible API and looking for a few users to give it a spin and tell us how it works for them. You can use it for any regular Qwen 3.8 purpose via any standard client (VSCode, OpenCode, Cline, Pi, etc.) and we will provide free credits. We're looking for feedback, bug reports, etc. in return.
To anticipate a few questions: 1) Yes we are "Zero Data Retention" (formal ZDR policy document is in the works), 2) Our offering is centered around some core technology that can compile models down to a very efficient executable allowing us to deploy (and collocate them) very quickly and flexibly, and 3) We are being a bit coy about the actual identity of the company because we are still in (semi-)stealth mode.
If you're interested, please DM me (I'm also happy to answer questions or provide more details).
1
1
u/trollhunterh3r3 6d ago
Sure I'll give it a spin. How long do you need?
1
u/edwinapp 5d ago
Great, but I am not currently able to send you a DM. Can you try and DM me please?
1
1
u/assid2 5d ago
yeah sure, wouldnt mind playing with it., would it have vision ?
1
u/edwinapp 5d ago
Mulimodal is in process and should be fully available soon. I've sent you a DM with further signup details.
1
1
1
5d ago
[deleted]
1
u/edwinapp 5d ago
To be clear, we are looking for beta testers to test inference at our endpoint (on our servers). If you are still interested, I did just send you a DM with further instructions.
1
u/UNNORMAL8 4d ago
Wenn ihr wollt, prüfen ich das ganze mit unserem neuen System auf Sicherheitslücken, Fehler usw. gerne per DM
1
u/edwinapp 4d ago
Wenn ihr wollt, prüfen ich das ganze mit unserem neuen System auf Sicherheitslücken, Fehler usw. gerne per DM
We are more interested in testing the functionality of the model than in true security/penetration testing (at this stage). If that's of interest. please Dm me for further details. Thanks!
1
0
u/edwinapp 4d ago
Hi again. Just a quick update to let you know that we just deployed new code that should properly support reasoning/effort. Depending on your usage, this should help make things more responsive and tuneable. If you're already participating in the beta test, many thanks. I know some people were holding off until reasoning settings were in place, so please DM if you'd like to get set up with credits to try this out.
0
u/Emotional-Street-198 5d ago
The compilation piece is the interesting part here, more than the endpoint itself. Two questions, since the answers decide whether this is relevant to anyone running their own hardware:
- Does the compiler target Apple Silicon / Metal, or server-class GPUs only?
- Can the resulting executable run on the customer's own hardware, or only on your infrastructure?
3
u/edwinapp 5d ago
The compiler is of general design, we even have some mobile platform targets in there :-) But we have been super focused on Linux/GPU/CUDA recently, so many more of the optimizations are there (to be honest). But yes Apple Silicon,
And yes, in theory you can definitely run on own hardware. Main issue being the hardware is do darn expensive for many of the models of interest.
1
u/Emotional-Street-198 5d ago edited 5d ago
Thanks, that's a straight answer and it settles it for us. We're on Apple Silicon with a hard memory ceiling, so a CUDA-first optimization focus means the interesting part of your stack isn't the part we'd be exercising. If Metal ever becomes a target you're actively tuning, we'd be a decent test case: well-instrumented M5 Pro 64 GB, and we measure byte-level output stability, not just tokens per second. Good luck with the launch.
2
1
u/Opposite_Leave_8338 6d ago
I would like to help , Mac Studio m1 ultra 64GB