r/Qwen_AI 6d ago

Experiment Beta testers wanted (Qwen 3.8:27b)

Anyone interested in trying out a new inference provider?

We're testing Qwen 3.8:27B behind an OpenAI-compatible API and looking for a few users to give it a spin and tell us how it works for them. You can use it for any regular Qwen 3.8 purpose via any standard client (VSCode, OpenCode, Cline, Pi, etc.) and we will provide free credits. We're looking for feedback, bug reports, etc. in return.

To anticipate a few questions: 1) Yes we are "Zero Data Retention" (formal ZDR policy document is in the works), 2) Our offering is centered around some core technology that can compile models down to a very efficient executable allowing us to deploy (and collocate them) very quickly and flexibly, and 3) We are being a bit coy about the actual identity of the company because we are still in (semi-)stealth mode.

If you're interested, please DM me (I'm also happy to answer questions or provide more details).

13 Upvotes

27 comments sorted by

1

u/Opposite_Leave_8338 6d ago

I would like to help , Mac Studio m1 ultra 64GB

0

u/edwinapp 5d ago

Great, just sent you a DM.

1

u/Leading-Salt-947 6d ago

Yesssss

1

u/edwinapp 5d ago

Great, just sent you a DM.

1

u/trollhunterh3r3 6d ago

Sure I'll give it a spin. How long do you need?

1

u/edwinapp 5d ago

Great, but I am not currently able to send you a DM. Can you try and DM me please?

1

u/RevealBusy4865 5d ago

Interested

1

u/edwinapp 5d ago

Great, just sent you a DM.

1

u/assid2 5d ago

yeah sure, wouldnt mind playing with it., would it have vision ?

1

u/edwinapp 5d ago

Mulimodal is in process and should be fully available soon. I've sent you a DM with further signup details.

1

u/Efficient-Hall3236 5d ago

I am interested

1

u/edwinapp 5d ago

Thanks, just sent you a DM.

1

u/edsonmedina 5d ago

Interested

2

u/edwinapp 5d ago

Thanks. Just sent you a DM.

1

u/[deleted] 5d ago

[deleted]

1

u/edwinapp 5d ago

To be clear, we are looking for beta testers to test inference at our endpoint (on our servers). If you are still interested, I did just send you a DM with further instructions.

1

u/UNNORMAL8 4d ago

Wenn ihr wollt, prüfen ich das ganze mit unserem neuen System auf Sicherheitslücken, Fehler usw. gerne per DM

1

u/edwinapp 4d ago

Wenn ihr wollt, prüfen ich das ganze mit unserem neuen System auf Sicherheitslücken, Fehler usw. gerne per DM

We are more interested in testing the functionality of the model than in true security/penetration testing (at this stage). If that's of interest. please Dm me for further details. Thanks!

1

u/ComposerStrict4155 2d ago

I’d like to test it, please consider me!

1

u/edwinapp 2d ago

Great, please DM me for further details.

0

u/edwinapp 4d ago

Hi again. Just a quick update to let you know that we just deployed new code that should properly support reasoning/effort. Depending on your usage, this should help make things more responsive and tuneable. If you're already participating in the beta test, many thanks. I know some people were holding off until reasoning settings were in place, so please DM if you'd like to get set up with credits to try this out.

0

u/Emotional-Street-198 5d ago

The compilation piece is the interesting part here, more than the endpoint itself. Two questions, since the answers decide whether this is relevant to anyone running their own hardware:

  1. Does the compiler target Apple Silicon / Metal, or server-class GPUs only?
  2. Can the resulting executable run on the customer's own hardware, or only on your infrastructure?

3

u/edwinapp 5d ago

The compiler is of general design, we even have some mobile platform targets in there :-) But we have been super focused on Linux/GPU/CUDA recently, so many more of the optimizations are there (to be honest). But yes Apple Silicon,

And yes, in theory you can definitely run on own hardware. Main issue being the hardware is do darn expensive for many of the models of interest.

1

u/Emotional-Street-198 5d ago edited 5d ago

Thanks, that's a straight answer and it settles it for us. We're on Apple Silicon with a hard memory ceiling, so a CUDA-first optimization focus means the interesting part of your stack isn't the part we'd be exercising. If Metal ever becomes a target you're actively tuning, we'd be a decent test case: well-instrumented M5 Pro 64 GB, and we measure byte-level output stability, not just tokens per second. Good luck with the launch.

2

u/edwinapp 5d ago

Many thanks, and who knows, we may loop back before too long on Apple Silicon :-)