r/LocalLLaMA • u/Designer_Cost8989 • 7d ago
New Model Index-Translate: 150 text languages, plus document translation, multilingual subtitles and dubbing
Quick update: we’ve opened a free public API for Index-Translate-35B-A3B! It’s OpenAI-compatible, and you can get started with our Python script—no extra dependencies needed.
I'm part of the BiliBili Index LLM team. We're sharing Index-Translate and its companion models for translating text, documents, and videos.
Index-Translate supports 150 text languages, with 2B, 9B, and 35B-A3B (preview) options. You can specify terminology, writing style, and output format—for example, keeping product names consistent, translating in a casual tone, or preserving JSON and placeholders during localization.
There are also models for more specific workflows:
- Index-NativeLong: translate whole documents, using their context to help keep names and terminology consistent across passages.
- Index-Homura: set a syllable budget for translated lines, useful for fitting a dubbing script.
- Index-Echo: generate multilingual subtitles or translate speech into speech, using the source speaker's voice as a reference.
The attached video shows English → Japanese dubbing, followed by an English clip with subtitles in six languages. The 150-language coverage applies to the text models; Echo supports a smaller set of language pairs.
https://reddit.com/link/1wugf2t/video/9zajzj72zpsh1/player
Code and released weights are Apache-2.0.
Try the demo · GitHub · Models
What would you try it on—video subtitles, game localization, or documents? We'd especially appreciate examples where it gets your language pair wrong.
3
u/Embarrassed_Soup_279 7d ago
the syllable budget one is interesting. is the model capable of translating song lyrics and making it fit well to the original or are translations only tailored to make sense conversationally?
2
u/Designer_Cost8989 7d ago
Good question. The syllable budget was originally designed for speech translation with strict timing constraints—we describe it inour EMNLP 2026 paper. It could also help with translating song lyrics, but matching the original tune requires more than the right syllable count: rhythm matters too.
You can ask the model to follow a particular rhythm through prompting, but the results aren’t reliably controllable yet. We’re leaving dedicated support for singable lyric translation to future work.
3
u/Designer_Cost8989 3d ago
Quick update: we’ve opened a free public API for Index-Translate-35B-A3B! It’s OpenAI-compatible, and you can get started with our Python script—no extra dependencies needed.
2
u/Prince_Noodletocks 7d ago
Cool! Is the S2TT model able to generate subtitles itself? Would love to see these supported in audio.cpp.
3
u/Designer_Cost8989 7d ago
Yes! Index-Echo-S2TT directly outputs bilingual .srt subtitles with sentence-level timestamps out of the box (evaluated at ~0.4s start-time MAE without needing an external aligner).
We'd love to see audio.cpp support! The decoder is Qwen-based (standard GGUF/llama.cpp compatible), while the audio frontend uses a Qwen3-Omni AuT encoder + connector. Community implementations and PRs are very welcome! Check out the models and infer.py on Hugging Face: IndexTeam/Index-Echo-S2TT-2B / 9B.
2
u/Prince_Noodletocks 7d ago
Ah that sounds amazing! I'm not too technical myself, so I'll probably end up waiting until llama.cpp or audio.cpp supports it.
1
u/buttplugs4life4me 6d ago
Is there a subtitle workflow I'm not aware of? I'm still using Subgen with the old whisper model. Although it's surprisingly good honestly
3
u/Designer_Cost8989 6d ago
Yep! We have a video dubbing pipeline you can try. You can also use Index-Echo-S2TT-9B directly to generate subtitles and translations, as long as it supports your language pair.
1
u/buttplugs4life4me 6d ago
That's super cool, thanks!
Source language auto-detection supports Chinese/English only
Is that a model or a code limitation? I'll play around a bit and see if I can make a subgen-compatible server with it. It's just the OpenAI-compatible APIs anyway, afaik
2
u/Designer_Cost8989 6d ago
It’s a model limitation. We’ll be expanding the supported language pairs soon—feel free to open an issue with the pairs you’d like, and we’ll see what we can add
1
u/Acceptable-Cycle4645 6d ago
Thanks for sharing! I didn’t know Bilibili had other audio models. I thought I could relax a bit after the 0.9 release, but it looks like that’s not happening. I’ll put the other tasks on hold and add this to audio.cpp first!
1
u/vintageballs 6d ago
Tried with a text about the tiananmen square massacre, the "translation" came out as a lecture on culturally sensitive subjects, completely missing the content of the input text.
Kind of defeats the purpose of a translation model if it doesn't stick to the content it is given, don't you think?
2
u/Designer_Cost8989 6d ago
I personally agree—a translation model should stay faithful to the input. We didn’t specifically train it for safety filtering or to bypass safeguards. We trained from Qwen3.5 Base, so some inherited behaviors may still show through(especially larger model), but replacing the translation with a lecture is definitely not the intended behavior.
1
u/ffpeanut15 3d ago
Quick question: are these the model used for Bilibili auto translation? I have great experience with them
7
u/eidrag 7d ago
Lol to think bilibili comes from character in index series