
Microsoft’s New AI Model Scores Choices and Leaves the Rules to You
Microsoft introduced an artificial intelligence model on October 9 that scores the choices an application gives it. Microsoft-Decision-1 targets routing, classification and workflow control, according to the company’s announcement. For developers deciding where a ticket goes or whether an automated step should proceed, the useful change is the output: probabilities attached to predefined options.
It is available through Microsoft Foundry, Microsoft’s Azure platform for building and deploying AI applications. Microsoft’s model card describes a text-only service that returns decision scores without a written explanation. That makes the surrounding application’s rules especially important: a score still needs a policy for what happens next.
The application supplies the choices
Microsoft’s deployment guide describes three question types: a yes-or-no judgment, a choice among named options, and a rating on an ordered scale. Developers supply the information to assess and the question. The model produces the corresponding scores; its model card rules out open-ended conversation and text generation as intended uses.
Consider a hypothetical support queue. An application could provide a ticket’s text and ask whether it belongs with billing, account access or engineering. It could send a strong result to the chosen queue and send an ambiguous result to a person. The product team would have to define those categories and decide how much uncertainty it can accept.
One implication is that this design can make a routing boundary explicit. You can inspect the available destinations and the rule that triggers a handoff. It doesn’t establish that the selected destination is correct. A tidy response format and a useful decision are separate things.
Cheap decisions still need good questions
Microsoft lists a price of $0.042 per million input tokens, the units of text the service processes, with no charge for output tokens. At that announced rate, a hypothetical million requests using 1,000 input tokens each would cost $42 for model input. That calculation excludes other application costs and any extra requests or context.
The more consequential constraint is in Microsoft’s deployment documentation: wording and option order can affect scores, and a poorly framed question can still receive a confident-looking answer. Microsoft recommends checking representative examples, testing reordered choices and choosing thresholds around the cost of false positives and false negatives.
In the support example, that means checking tickets that mention both payment and a locked account, rather than evaluating only clean examples with one obvious label. You might allow a cheap mistake between two internal queues while requiring review before an action affects a customer. Those are proposed application policies, not demonstrated results from this model.
The launch chart isn’t your deployment
Microsoft says the model had the highest accuracy in its comparison across 36 benchmarks, covering nearly 150,000 questions kept separate from training. It also claims a latency advantage. Those are the company’s measurements, with a specific test setup, rather than a guarantee about an application’s response time.
H2O.ai, which develops a competing decision model, challenges the latency comparison in its model card. It says Microsoft used adjusted leaderboard timings for self-hosted competitors rather than their measured inference times. That is a rival’s methodological objection, not proof that one model will be faster in every deployment.
There is also a separate measurement to consult. Benchmark Heaven’s JevBench board for hosted decision services places Microsoft-Decision-1 sixth in its October 10 snapshot. Its composite combines decision quality, confidence calibration, speed and cost. That asks a different question from Microsoft’s broader accuracy comparison, so the two rankings cannot settle the dispute by themselves.
For a team considering the service, the next useful experiment follows directly from those limits: keep a sample of real routing cases with known destinations, include ambiguous ones, and compare mistakes, review rates and response times under the same conditions. The practical question is whether the scores help your application make acceptable decisions at its own boundary, including when it should stop and ask a person.
Photo: Coolcaesar / Wikimedia Commons, CC BY-SA 4.0 (cropped).
Sources
Microsoft’s October 9 announcement
Microsoft-Decision-1 model card