Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Microsoft Decision 1 brings probability scoring to Foundry

Listen to this article

Microsoft introduced Microsoft Decision 1 on October 9, 2026 in Microsoft Foundry. The model scores supplied answer options for routing, classification and evaluation. Input costs $0.042 per million tokens, with no output charge.

The practical opportunity is to give a narrow, repeated choice its own component. A system selecting a work queue has a different job from the assistant writing the eventual response. That separation is useful only if developers can inspect mistakes and decide which ones are acceptable.

What goes into a decision

The Foundry model card lists general availability, a 32,768 token context window, text input and JSON output. Microsoft post trained the nine billion parameter Qwen 3.5 model. Requests supply a situation, a closed question and defined answers. Supported formats include yes or no, multiple choice and ratings.

The output is a probability for each answer. The model accepts no images, audio or video and supplies no written explanation. Evidence needed for the judgment must be in the input. An explicit option for insufficient evidence is supported. Applications remain responsible for thresholds, escalation and safeguards.

For an illustrative software support workflow, define separate questions for the destination team and the urgency of the incident. Give each answer a distinct meaning. A report about a failed export might belong in technical support, but urgency depends on whether one user or the entire service is affected. This follows the question design approach in OpenAI’s Decisions guidance.

A label should then select a review queue or a permitted next step. Keep the record of the original report, the options, the returned scores and the action actually taken. That record lets a team diagnose whether the problem came from missing evidence, a badly defined category or its own execution logic.

The token price leaves other costs to measure

At the published rate, 100,000 requests of 2,000 billable input tokens would cost $8.40. This illustrative model input bill excludes retries, other services and human review.

Microsoft’s launch says OpenRouter access is forthcoming. An OpenRouter listing checked on October 9, 2026 already shows Azure as a provider at the same rate. That establishes a catalog listing. ByteForward has not tested a request through either service.

For Foundry users, Microsoft’s platform documentation says catalog filters distinguish regions, lifecycle and deployment options. Check the model’s actual offer in the intended Azure project before estimating rollout effort. General availability alone does not establish that every subscription or region has the same access.

How to read the benchmark claims

Microsoft reports the best accuracy in its comparison across 36 benchmarks and nearly 150,000 questions, plus median latency roughly 35 times faster than GPT 6 Sol. The comparison includes public and private benchmarks. These are Microsoft’s results. ByteForward has not reproduced them.

A team testing a replacement should compare the same inputs, answer choices and error policy. Measure the slow requests as well as the median, and count how often the system sends work for review. A fast scorer can still leave a slow workflow if retrieval, retries or the review queue dominate elapsed time.

The alternatives differ in access and inputs

The OpenAI Decisions API is currently a public beta using GPT 6 Luna. It evaluates text and images, with input starting at $0.10 per million tokens and no output charge. Regional and long context adjustments apply. Image support is a material difference for a workflow that needs to inspect a product photograph.

Cloudflare’s October 1, 2026 Clef announcement describes image classification and releases Clef and Clef Flash under Apache 2.0 for local use as well as hosting on Workers AI. That creates another selection criterion for teams that need access to model weights. These are capability and access differences, not a shared accuracy ranking.

The first test should measure mistakes

OpenAI’s threshold guidance recommends labeled application examples and thresholds based on the cost of false positives and false negatives. For the support example, track missed urgent incidents separately from unnecessary escalations. A single overall accuracy figure can hide the more expensive mistake.

Microsoft’s model card excludes using Decision 1 as the sole automated basis for consequential decisions about people, including employment, credit and healthcare. It also rules out tasks requiring knowledge absent from the input. A probability score does not remove those boundaries.

Start with recorded cases where the correct destination and urgency are already known, then compare recommendations with the existing process before allowing actions. The useful result is a measured reduction in work at an acceptable error rate. A cheap request is only one part of that calculation.

Archival photograph of Microsoft’s Redmond campus taken in June 2015 by Jiaqian AirplaneFan and used under CC BY 3.0. The original image illustrates the company and does not depict the model or its launch.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile