Voyage launches Rerank 3 and Lite for AI search
Voyage Rerank 3 and Lite reorder search results with updated models, familiar token prices and practical limits for long documents.

Voyage AI released Rerank 3 and Rerank 3 Lite on September 30, introducing two models that reorder search results before those results reach an AI assistant. Its launch announcement reports the largest improvements over the previous generation on long documents and code.
A reranker compares candidate documents with a query after an initial search has found potentially useful material. Voyage’s model guide recommends Rerank 3 for accuracy and describes Lite as the option for applications where response time and cost are priorities.
What the extra search step changes
For a coding assistant, that could mean moving the relevant function above several similar files. For an internal help tool, it could mean putting the right policy passage ahead of loosely related documents. The practical limit is straightforward. A reranker cannot recover evidence that the initial search never included.
Voyage says existing Rerank 2.5 users can keep their integration and that relevance scores are calibrated to preserve existing thresholds. Teams should still test queries near those cutoffs before replacing a model in production.
Pricing depends on every candidate
The pricing page lists Rerank 3 at $0.05 per million tokens and Lite at $0.02. Billing counts the query again for every candidate, then adds the document tokens. Long queries and large candidate lists therefore affect the bill even when the application returns only a few results.
Voyage’s example uses 100 documents with 500 tokens per query and document pair. That produces an estimated reranking charge of $0.0025 for Rerank 3 or $0.001 for Lite. These are reranking costs. A complete assistant also needs to account for its initial search and answer generation.
Context limits need deliberate handling
Both models allow 32,000 tokens for a query and an individual document combined. The query itself is limited to 8,000 tokens. Requests can include up to 1,000 documents, subject to a total limit of 600,000 tokens that counts the repeated query. These limits apply together.
Truncation is enabled by default, cutting inputs to fit the context limit before scoring. Developers should inspect long inputs rather than assume every submitted passage was evaluated in full. The output setting that chooses how many results to return does not change how many candidates were submitted.
The benchmark measures the final ordering
Voyage reports results across 95 datasets in nine domains. Its evaluation method retrieves up to 100 candidates, reranks them and measures the top 10 using NDCG@10, a ranking quality metric. These are company reported retrieval results, rather than a general benchmark of an assistant’s answers.
A useful rollout test should keep the initial candidate lists fixed while comparing the old and new models. Check whether relevant passages reach the top, whether important text was truncated and how much latency the extra step adds. The stronger choice is the model that improves evidence selection on the actual workload.
For the initial retrieval step, see our coverage of Cohere Embed 5 Pro and Fast and the different role that embedding models play in enterprise search.
AI generated conceptual illustration of document relevance ranking



