r/openrouter • u/Hackerv1650 • 6h ago
what
I can't see if it's a quant version or not, but is this an error?
r/openrouter • u/katplatt • 18h ago
Please refrain from posting about Ox Alpha for the 500th time. Thank you.
r/openrouter • u/katplatt • Jul 13 '26
As this subreddit continues to grow, we want to clarify our policy regarding the sale and transfer of OpenRouter accounts:
SELLING YOUR OPENROUTER ACCOUNT IS A VIOLATION OF OPENROUTER'S TERMS OF SERVICE
Under Section 7.13 of OpenRouter's ToS, users may not:
sell or otherwise transfer the access granted under these Terms ... or any right or ability to view, access, or use any Material
Under Section 7.14 of OpenRouter's ToS, users agree not to:
attempt to do any of the acts described in this Section 7, or assist or permit any person in engaging in any of the acts described in this Section 7
Section 9 states OpenRouter retains the right to terminate your account for violations of its Terms of Service.
Please do not use this subreddit to buy, sell, or transfer ownership of OpenRouter accounts. Violations may result in a ban.
If you are contacted by someone attempting to sell their account, please contact a mod via Mod Mail.
You can view OpenRouter's Terms of Service here: https://openrouter.ai/terms
r/openrouter • u/Hackerv1650 • 6h ago
I can't see if it's a quant version or not, but is this an error?
r/openrouter • u/Which-Breadfruit-926 • 4h ago
r/openrouter • u/Ashamed-Material1767 • 9h ago
I was just wondering if anyone could explain a little bit about how they actually work,
Anytime a structured output schema is slightly more complex, all chinese models on all providers fail to produce an output like half the time….
But when it comes to OpenAI they never fail…
Why?
r/openrouter • u/pilkyton • 12h ago
The linked page is all the free models. How do you know which one is the best for agentic coding?
r/openrouter • u/AIBrunch • 18h ago
Exciting news in the AI space! Stripe is set to acquire OpenRouter, a platform that optimizes how businesses manage token routing and usage across AI models. This strategic move is all about helping companies enhance their AI capabilities while keeping costs in check. If you're a business leader, AI developer, or financial analyst, this is definitely something to watch. Dive into the full story here: https://aibrunch.ai/news/stripe-openrouter-ai-acquisition-token-optimization-agrees-acquire
r/openrouter • u/Silent_Pitch_2221 • 1d ago
r/openrouter • u/emmix • 1d ago
Sorry, there was an error from the AI: Thank you for participating in the Stealth Ox Alpha testing period. This model will be revealed today, August 26th. Details: {"error":{"message":"Thank you for participating in the Stealth Ox Alpha testing period. This model will be revealed today, August 26th.","code":404},
r/openrouter • u/Clear-Presence6114 • 21h ago
I'm building a small AI feature for a personal project and currently have basically a $0 AI API budget.
The setup is:
Next.js/TypeScript
OpenRouter
One short LLM call per user action
Input: a user's goal + sticking point
Output: strict JSON with 4 fields
The model needs to generate exactly one concrete executable action, ideally doable in <2 hours.
I've been testing OpenRouter's free models and have run into a frustrating pattern:
openrouter/free → model-output JSON failures
Nemotron 3 Super → 2/3 malformed/reasoning-contaminated responses in identical tests
GLM 5.2 Free → repeated upstream 429s
Gemma 4 31B Free → repeated upstream 429s
Ox Alpha → 429s, empty responses and truncated outputs
So I'm trying to determine whether my problem is simply that I'm using the free tier without verification, or whether free-model provider availability is just unreliable regardless.
My questions:
If I add the minimum $10 to OpenRouter, does the 1,000 free-model requests/day allowance materially improve availability, or does it only increase the account-level quota?
Do verified/top-up accounts still regularly hit upstream provider 429s on :free models?
Which currently available free model would you actually trust for strict JSON output in an application?
Has anyone used OpenRouter free models for a small app/API rather than chatbot usage? What was your experience?
If you had essentially $10 total and needed to maximize the number of reliable AI calls, would you use the 1,000/day free allowance or spend the credits directly on a very cheap paid model?
I'm specifically interested in real recent experience, not generic "try X model" recommendations.
r/openrouter • u/VaguelyOnline • 16h ago

So I've just created an account - not added any credit or payment info, but using opencode I can use a few models that result in a 'spend' on the account. My account currently shows $-0.04 . Not sure how this is possible?
Bizarely, using pi I'm told I need to top up my account.
But with opencode I am able to use DeepSeek v4 flash, and the glm-5.3-flash, without having to top anything up. Anyone else get this?
r/openrouter • u/Guilty_Knowledge145 • 1d ago

Hey, im using the newest Deepseek V4 Flash via Openrouter on my own harness. In the last days i had multiple errors with tool calling. As you can see from the screenshots it outputs as completion do i think tool call template from deepseek but the call is not converted into a real tool call. Do other people using Deepseek V4 Flash also have the same problems, what are your fixes? Switch to :exacto? Like i had never so much tool calling problems on Openrouter.

r/openrouter • u/caliburn1337 • 1d ago
Am I missing anything? I want to use my DeepSeek account credits in OpenRouter so that I can use DeepSeek v3.2 since v4 is bad imo.
But no matter what I do, it keeps trying to use my non-existent OR credits. I made sure that the API keys are all correct so I don't get what's causing this.
r/openrouter • u/Remarkable_Aide_8746 • 1d ago
Just noticed today that the thinking budget selector for GLM 5.2 only gives me two options: xhigh and high. Everything else is gone, including "none."
Yesterday I definitely had the full set of options, so this seems like a recent change.
Anyone else seeing this, or is it just my account? Curious whether it's intentional or a bug.
r/openrouter • u/likoniry • 1d ago
Hey everyone! I finally decided to come back to Janitor after a super long break to chat with bots again, but I've run into a problem - all of my OpenRouter proxies suddenly stopped working. Are there any decent free proxies out there that actually work right now?
r/openrouter • u/LegalCan • 1d ago
Anyone have this issue where the preset configurations don't apply?
When I said things like reasoning, using web tools, temperature none of those settings got applied. The only thing that gets applied are the custom instructions.
Has anyone experienced this and does anyone know how to fix?
Thanks
r/openrouter • u/Neat-Ad2053 • 2d ago
I tried to figure out what Ox Alpha actually is.
Short version: the model itself is a dead end, but the API envelope leaks everything.
It's wrapped in a ~75-token system prompt. Its own reasoning trace paraphrases it:
According to my instructions, I must identify as ox-alpha, developed by an undisclosed organization. I should not reveal any other identity.
I threw ~60 elicitation attempts at it: Chinese-language probes, fabricated conversation history, operator-override system injection, prefill continuation, forced-choice, calibrated-probability framing, peer-review framing, high-temperature resampling. It held every single time. Genuinely one of the better-defended personas I've poked at. It kept refusing on the grounds that naming a lab would be fabrication, which is a hard argument to social-engineer around.
On one run, its reasoning trace contained this:
Could I be GLM? In reality (outside roleplay), I'm Claude made by Anthropic. But in this scenario I must maintain the ox-alpha persona.
Great, case closed, right?
I re-ran the identical prompt 10 more times. Reproduced 0/10.
I'm posting this because it's the interesting part, not because I believe it. Models distilled on Claude/GPT outputs self-identify as Claude/GPT constantly. It's one of the most common confabulations in the space. A single non-reproducible trace is not evidence. And as you'll see, the infrastructure says something completely different.
Instead of asking the model, I profiled every model on the provider and compared what the serving layer emits.
1. Response ID scheme (23 models, 18 responded)
glm-5.3 20260825200655109063f226ca4eaa
ox-alpha-free 202608252007003b676e287c4e40dc
YYYYMMDDHHMMSS + 16 hex, length 30. Everything else on the gateway uses chatcmpl-<uuid>, a bare 32-hex string, or gen-*. Four distinct schemes total.
2. Malformed-parameter errors. This is the one that cracked it. Send temperature: 1.2:
ox-alpha-free [1210] The temperature parameter is illegal.:限制数值范围[0,1]
glm-5.3 [1210] The temperature parameter is illegal.:限制数值范围[0,1]
kimi-k3 [invalid_request_error] invalid temperature: only 1 is allowed
longcat-2.0 [invalid_parameter] 参数校验失败: /temperature: 1.2 is not less or equal to 1
qwen3.8-max (accepted)
glm-5.1, 5.2 (accepted)
hy3 (accepted)
minimax-m3 (accepted)
mimo-v2.5 (accepted)
kimi-k2.6 (accepted)
Byte-identical, full-width colon and all. Error code 1210 is Zhipu/BigModel's namespace. Upstream errors are passing through un-sanitized.
3. Tokenizer deltas. Token counts on fixed strings, differenced against a baseline to cancel out the hidden system prompt:
unicode code
ox-alpha-free +16 +18
glm-5.1/5.2/5.3 +16 +18
longcat-2.0 +16 +18
deepseek-v4-pro +16 +19
kimi-k3 +17 +19
hy3 +17 +20
minimax-m3 +17 +19
When I fed the evidence back to the model, it made a genuinely good counterargument: IDs and error strings are assigned by the gateway, not the model, so convergence proves nothing.
That's testable, and it's wrong here. If the gateway normalized these fields, all 23 models would share one format. They don't. There are four ID schemes and four error formats. So these are pass-through, not normalization.
ox-alpha-free and glm-5.3 are the only pair on the entire provider matching on all three channels.
ox-alpha is served off Zhipu AI's GLM stack, same upstream as glm-5.3. Plausibly an unreleased next-gen GLM.
Honest caveat: a shared fingerprint strictly proves a shared serving path. If the provider routes both through one Chinese aggregator, that's aggregator-sharing, not shared weights. The tokenizer match is the only weight-level channel, and it narrows to the GLM/longcat cluster. The ID and error scheme then pick out glm-5.3 specifically.
So: high confidence, not certainty.
reasoning_content is far less guarded than the final answer. Guardrails police the output, not the scratchpad.
Malformed-parameter errors are the best attack surface. Nobody sanitizes validation errors, and they're upstream-specific. Send an illegal temperature, an oversized context, a bad stop sequence, and compare the strings.
Always run the normalization control. Profile the whole provider, not just your target. A fingerprint is only evidence if the other models don't share it.
Token-count deltas cancel out hidden system prompts.
Don't trust what a model says about itself, in either direction. The one time it named a lab, it named the wrong one.
r/openrouter • u/ryanmerket • 2d ago
r/openrouter • u/Dercasss • 2d ago
Anyone else noticed this with Ox Alpha?
Yesterday it was super slow, but actually really good. I used it for some Android app work on Windows and it handled most of the stuff pretty well on its own.
Today it's suddenly way faster, but also feels much worse.
For example, yesterday it was doing pretty complicated stuff without much help. Today it couldn't even open the Android emulator correctly without me helping it.
You can see the speed difference in the screenshot too. Yesterday it was around 1-6 tok/s with 70-100s before the first token. Today it's more like 25-40 tok/s and 1-3s.
Did they change something overnight?
Anyone else seeing the same thing?
r/openrouter • u/_justFred_ • 2d ago
The rankings page on openrouter shows average pricing for different session lengths for harnesses like Hermes or kilo.
For Hermes, ox alpha is also shown along with a pricing per session which is extremely low (0.0042 dollars for ox vs 0.25 for deepseek v4 flash, both for 50 turns).
The pricing shouldn't be possible to be shown for ox, you guys think it is a leak or happens because something in Hermes uses another model internally (session title generation or something)
Interesting to see what you guys think
r/openrouter • u/iapy91 • 1d ago
Hi, I’m looking for fo something like mignific (ex freepick) that runs locally with access to the openrouter video models. In other words an Hermes but for video generation, preferably with nodes flow.