Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Cohere launches Embed 5 Pro and Fast for enterprise search

Cohere Embed 5 pairs Pro and Fast retrieval models with a shared index, multimodal inputs and separate text and image pricing.

Listen to this article

Cohere released Embed 5 on September 30, introducing Pro and Fast models for enterprise search. The original launch announcement says the pair is generally available through the Cohere API, Model Vault, Microsoft Foundry and Amazon SageMaker.

Embedding models turn information into lists of numbers that software can compare for similarity. That lets a search system find related material even when a query and a document use different wording. Cohereโ€™s embedding guide describes uses including search, classification and clustering.

Two models can search the same index

Pro is designed for quality focused retrieval and offline indexing. Fast targets interactive search and larger query volumes. According to the release notes, both use a shared embedding space. Cohere recommends indexing a collection with Pro and using Fast to answer queries against it.

The launch notes add an important condition. Query and document vectors must use the same output dimension.

Both models accept text, images and combined text and image inputs. They support more than 100 languages and a context window of 128,000 tokens. Developers can choose six output sizes from 256 to 2,048 dimensions, with floating point, integer or binary representations.

Developers still need to distinguish the query from the material being searched. Cohereโ€™s integration guidance assigns separate input modes to search queries and stored document passages. Those settings are part of preparing compatible representations for retrieval, rather than treating every input as interchangeable.

Text and image rates differ

The launch pricing table lists Pro at $0.12 per million text tokens and Fast at $0.08. Image inputs cost $0.40 per million tokens for either model. The cheaper text rate therefore does not extend to images. Those units should be checked before estimating the cost of processing a mixed document collection.

Cohereโ€™s pricing page treats Model Vault separately, with instance billing based on the model and performance tier. It also says trial API keys are restricted to evaluation and cannot be used commercially. Teams moving from experiments to a live service need production access.

Retrieval scores need context

Cohere reports an average score of 85.8 for Pro and 84.5 for Fast in its ViDoRe V3 evaluation, covering eight document domains. The evaluation reorders a fixed set of candidate documents. The scores measure reranking quality rather than the initial retrieval step that finds those candidates. These vendor reported results do not guarantee performance on a companyโ€™s own collection.

The practical choice is whether Fast can find the right evidence quickly enough for a particular application. A useful evaluation should include difficult queries, missing answers and documents with important tables or layouts. Comparing retrieval quality, response time and total cost on that workload will be more informative than choosing from a headline score alone.

See ByteForwardโ€™s AI model releases coverage for more launch details and access conditions.

Featured image is an original AI generated conceptual editorial illustration.

Jordan Reid
Jordan Reid

Jordan Reid is focused on AI tools, agents, developer products, and the way technology changes everyday work. Jordan approaches a launch from the userโ€™s side of the screen. What can it actually help someone finish? The voice is practical, conversational, and skeptical of products that turn a simple job into five new settings. Coverage follows coding assistants, creative software, browser agents, and the workflows around them, with attention to pricing, permissions, setup, and the human work that remains.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile