Bilibili releases Index Translate models for 150 languages
Bilibili’s translation family adds text, speech and document models, with a 35 billion parameter preview and narrower language support for audio.

Bilibili’s Index Team released Index Translate on September 30, adding a family of downloadable models for multilingual text, speech and document translation. The official release record lists text models with 2 billion and 9 billion parameters, plus a preview of a larger model. The text family supports 150 languages.
The largest checkpoint uses a mixture of experts design built on Qwen3.5, with 35 billion total parameters and 3 billion active during processing. Its Apache 2.0 license allows developers to adapt the released model under the license terms. Instructions can ask it to preserve terminology, structured data, code and placeholders, or adjust the style of a translation.
Different models for different jobs
The family separates several workloads. Index Echo handles speech translation, Index Homura targets a requested syllable count for dubbing, and Index NativeLong handles complete documents. The release checklist still lists the final 35 billion parameter model and several benchmark datasets as future work. The available large checkpoint should therefore be treated as a preview.
Language coverage changes with the task. The released subtitle package takes Chinese audio or video and produces bilingual, timestamped subtitles in English, Japanese or Spanish. It includes an optional glossary and carries context between audio windows to help keep terminology consistent.
The speech dubbing package supports six directions from Chinese or English into specified target languages. It conditions the output on the source voice. Its guide recommends clips of 30 seconds or less per sentence, with a separate segmentation and alignment pipeline for longer videos.
The document translation templates cover Chinese paired with English or Japanese in either direction. Homura’s syllable setting is an approximate target. These limits matter when choosing a package for a subtitle, narration or book project. The text model’s 150 language inventory should not be read as equivalent coverage across every tool.
How the team measures quality
The technical report describes shared multilingual training followed by specialist training for general translation, instructions and cultural expressions. The team combines those specialists and applies further distillation to weaker tasks. Separate adaptations address speech, syllable control and long documents.
In the authors’ tests, the large preview scored 0.8794 on the FLORES COMET 22 metric and 0.8336 on the instTrans instruction score, the highest values among the compared systems on those measures. Its WMT26 judge score was 76.76, below several compared API models. The results support a narrower claim of strengths on selected translation tasks, rather than overall leadership. These are the team’s reported evaluations. ByteForward has not independently reproduced them.
What developers need to run it
The model card calls for a runtime that supports Qwen3.5 mixture of experts models and the complete checkpoint. Its example serves a 32,768 token context, counting both input and output. Available GPU memory constrains practical length, and longer translations may require raising the client’s output budget.
For the smaller text models, the serving guide estimates about 8 GB of GPU memory for 2B and 24 GB for 9B in bf16, with additional memory needed for long context caching. These are deployment estimates rather than a guarantee that every document will fit.
For more context on model score comparisons, read ByteForward’s coverage of independent Gemini evaluations.
Image credit. Training diagram by Tianjiao Li and colleagues, Index LLM Team. CC BY 4.0. Converted from SVG to PNG. Displayed blending weights apply to the 2B and 9B models.



