llama.cpp adds a local API for decision models
llama.cpp adds typed decision outputs through a local API, with measured speed tradeoffs and a later Clef integration limited to text.

llama.cpp has added a local API for decision models, giving applications a way to request bounded choices and scores through its server. The initial implementation merged on October 2 with support for Laya, Julia 1, lev, OpenJev and Kev. A subsequent change has already extended the lineup.
For developers, the useful change is a shared serving interface. Teams can evaluate several decision models in an existing local inference setup while keeping the application’s questions and answer structure recognizable.
A defined answer for each question
The server documentation describes the /v1/systemone endpoint. A request supplies content and named questions. Choice questions return the leading option and probabilities. Score questions use two to ten ordered levels and can return a value between them. The noul type returns the probability that an answer is true.
The response reports zero output tokens. Images require a compatible model and its vision projector, and must be supplied as embedded data rather than remote links. Requests cannot stream. These requirements matter when adapting an application that already sends images to another service.
Consider an internal document inbox. One question could select the review team, while another measures how urgently the document needs attention. The application still needs rules for what happens after those numbers arrive, including when to leave the item untouched.
Fast results need their hardware context
The October 2 announcement reports median times of 3 milliseconds for Julia 1, 5 for Laya, 12 for Kev 4B, 36 for lev and 43 for OpenJev. Each figure measures one question on one NVIDIA RTX PRO 6000. Only the first two results are below 10 milliseconds.
Those are project reported measurements. They do not establish response times on a laptop or under a production queue. The announcement also recommends testing different model sizes and quantizations against the developer’s own examples.
Model access also comes with different terms. The OpenJev GGUF model card identifies its weights as restricted to noncommercial use under Creative Commons Attribution NonCommercial 4.0. Downloadable weights and local runtime support do not, by themselves, settle whether a particular business deployment is permitted.
What the integration evidence establishes
The initial pull request compares reference outputs with llama.cpp on a customer support example. It reports probability differences across the five model families. That is evidence about reproducing a model’s output in this runtime, rather than proof that the model will classify a company’s documents correctly.
A separate Clef integration merged on October 3. It adds text support and explicitly leaves vision unsupported. That follows the model launch covered in our report on Cloudflare Clef and Clef Flash. Developers should distinguish the capabilities of the original model from those available in a particular runtime build. Confirm that the installed build includes the relevant change.
Test the decision that follows the score
The server documentation warns that returned probabilities are not guaranteed to be calibrated for the application’s data. A high confidence number therefore needs validation before it becomes permission to act.
A useful pilot would replay previously reviewed documents without changing their destinations. Measure wrong assignments, missed urgent items and the share requiring human review alongside response time. Then repeat that evaluation using the exact model file and hardware planned for deployment. A faster answer has little value if it increases the manual correction queue.
The immediate opportunity is to make a narrow decision step easier to deploy and compare. Start with an action that can be checked and reversed, and choose the acceptance threshold around the cost of a mistake. Runtime compatibility is the beginning of that evaluation.
Original AI generated conceptual illustration of local software weighing a set of possible decisions



