Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

llama.cpp adds a local API for decision models

llama.cpp adds typed decision outputs through a local API, with measured speed tradeoffs and a later Clef integration limited to text.

Listen to this article

llama.cpp has added a local API for decision models, giving applications a way to request bounded choices and scores through its server. The initial implementation merged on October 2 with support for Laya, Julia 1, lev, OpenJev and Kev. A subsequent change has already extended the lineup.

For developers, the useful change is a shared serving interface. Teams can evaluate several decision models in an existing local inference setup while keeping the applicationโ€™s questions and answer structure recognizable.

A defined answer for each question

The server documentation describes the /v1/systemone endpoint. A request supplies content and named questions. Choice questions return the leading option and probabilities. Score questions use two to ten ordered levels and can return a value between them. The noul type returns the probability that an answer is true.

The response reports zero output tokens. Images require a compatible model and its vision projector, and must be supplied as embedded data rather than remote links. Requests cannot stream. These requirements matter when adapting an application that already sends images to another service.

Consider an internal document inbox. One question could select the review team, while another measures how urgently the document needs attention. The application still needs rules for what happens after those numbers arrive, including when to leave the item untouched.

Fast results need their hardware context

The October 2 announcement reports median times of 3 milliseconds for Julia 1, 5 for Laya, 12 for Kev 4B, 36 for lev and 43 for OpenJev. Each figure measures one question on one NVIDIA RTX PRO 6000. Only the first two results are below 10 milliseconds.

Those are project reported measurements. They do not establish response times on a laptop or under a production queue. The announcement also recommends testing different model sizes and quantizations against the developerโ€™s own examples.

Model access also comes with different terms. The OpenJev GGUF model card identifies its weights as restricted to noncommercial use under Creative Commons Attribution NonCommercial 4.0. Downloadable weights and local runtime support do not, by themselves, settle whether a particular business deployment is permitted.

What the integration evidence establishes

The initial pull request compares reference outputs with llama.cpp on a customer support example. It reports probability differences across the five model families. That is evidence about reproducing a modelโ€™s output in this runtime, rather than proof that the model will classify a companyโ€™s documents correctly.

A separate Clef integration merged on October 3. It adds text support and explicitly leaves vision unsupported. That follows the model launch covered in our report on Cloudflare Clef and Clef Flash. Developers should distinguish the capabilities of the original model from those available in a particular runtime build. Confirm that the installed build includes the relevant change.

Test the decision that follows the score

The server documentation warns that returned probabilities are not guaranteed to be calibrated for the applicationโ€™s data. A high confidence number therefore needs validation before it becomes permission to act.

A useful pilot would replay previously reviewed documents without changing their destinations. Measure wrong assignments, missed urgent items and the share requiring human review alongside response time. Then repeat that evaluation using the exact model file and hardware planned for deployment. A faster answer has little value if it increases the manual correction queue.

The immediate opportunity is to make a narrow decision step easier to deploy and compare. Start with an action that can be checked and reversed, and choose the acceptance threshold around the cost of a mistake. Runtime compatibility is the beginning of that evaluation.

Original AI generated conceptual illustration of local software weighing a set of possible decisions

Jordan Reid
Jordan Reid

Jordan Reid is focused on AI tools, agents, developer products, and the way technology changes everyday work. Jordan approaches a launch from the userโ€™s side of the screen. What can it actually help someone finish? The voice is practical, conversational, and skeptical of products that turn a simple job into five new settings. Coverage follows coding assistants, creative software, browser agents, and the workflows around them, with attention to pricing, permissions, setup, and the human work that remains.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile