Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model
Microsoft has released Microsoft-Decision-1, a decision model for routing, classification, verification and agent control. Microsoft-Decision-1 is a decision-scoring model that returns a calibrated probability for each fixed answer option instead of generated text. It is post-trained from Alibaba’s Qwen3.5-9B and available now in Microsoft Foundry and OpenRouter. .
TL;DR
Size: Built on Qwen3.5-9B; exact parameter count not disclosed. 32,768-token context window.
Runs on: Hosted API only (Microsoft Foundry, OpenRouter via Azure). No open weights, no quantized variants, hardware not disclosed.
Performance: Highest average accuracy in Microsoft’s 36-benchmark comparison, at 85 ms p50 latency.
Best: 83.5% average accuracy across 36 benchmarks, ahead of Quyet-1.0-Large at 81.9%.
Worst: Calibration of 92.2, second to Quyet-1.0-Large at 93.1. Text-only, no explanations.
Bottom line:
Best thing: fast, cheap, calibrated decisions at $0.042 per million input tokens with free output.
Worst thing: closed weights, and every benchmark is vendor-run.
What is a decision model?
A decision model reads an input and scores a closed set of options. It does not write prose. Microsoft frames decision models as a new AI category, built for outputs software can act on immediately.
How does Microsoft-Decision-1 work?
Microsoft post-trained Qwen3.5-9B for single-pass decision scoring. Given a situation, a question and fixed options, it returns a probability per option in one call. It plans to rebase future versions on MAI and OpenAI models.
The Foundry model card lists the supported formats:
yes/no, multiple-choice, rating, classification and rubric questions
grading of AI responses and proposed agent actions
groundedness checks against supplied evidence
explicit abstention options such as “cannot tell”
Training used public datasets under Microsoft’s Open Data process plus synthetic data. Output is JSON. OpenRouter notes that weights update continually while the API shape stays fixed.
How fast and accurate is it?
Microsoft compared 9 systems across 36 benchmarks with 147,137 questions. Benchmarks were kept blind from training. Microsoft-Decision-1 led on average accuracy at 83.5%.
Its p50 latency was 85 ms, with p95 at 125 ms. That is 4.5 times quicker than Quyet-1.0-Large and 35 times quicker than GPT-6 Sol, which took 3.01 s.
Microsoft also tested robustness. It perturbs each request 8 ways, including paraphrasing and option shuffling. The model flipped its decision on 1.3% of perturbations on average. Flips were zero when options were paraphrased, reversed or shuffled.
On safety, Microsoft ran 5,250 requests across 11 benchmarks. These covered harmful content, jailbreaks and prompt injection. Microsoft reports correct refusals with high retained utility, without publishing a score.
Internal results Microsoft reports
Xbox Research: sorted 10,000+ feedback items, 14 times faster and 200 times cheaper than GPT-6 Sol.
Copilot quality control: competitive with GPT5.6 Luna and 100 times faster.
Microsoft Discovery: 46 times more consistent than LLM scoring, at 3 times the speed.
How does it compare with other decision models?
Accuracy, calibration and latency rows come from Microsoft’s chart. Competitor latencies there use the JevBench v1.6.1 adjusted median.
Is the latency comparison apples to apples?
Not fully. Microsoft measured its own model through Foundry. Competitor figures use JevBench’s adjusted median, not raw timings.
H2O.ai’s model card disputes this. It states that JevBench’s adjusted figure doubles measured time and adds 0.15 s. H2O says its model’s measured median is 29 ms, not 210 ms. Microsoft-Decision-1 does not appear on the JevBench board. On that board’s official composite, H2O-Lightning-4B ranks #1 at 72.5.
How do developers deploy it?
Developers can call it from Microsoft Foundry as a generally available ‘Direct from Azure’ model. On OpenRouter, it runs on the Decisions API, not the chat endpoint. Chat completions SDKs will not work.
Pricing is $0.042 per million input tokens, with free output. OpenAI’s Luna decisions endpoint charges $0.10 per million input tokens.
What are the limitations?
The model card is explicit:
Not for text generation, open-ended Q&A, chat, translation or summarization.
Text-only. No images, audio or video.
No explanations or rationales in the output.
Not for sole automated decisions on credit, employment, housing, healthcare or legal rights.
Applications must define options, thresholds, escalation paths and human oversight.
Key Takeaways
Microsoft-Decision-1 scores fixed options with calibrated probabilities, not text.
Post-trained from Qwen3.5-9B, with a 32,768-token context window.
Leads Microsoft’s 36-benchmark comparison at 83.5% accuracy and 85 ms p50.
Costs $0.042 per million input tokens; output tokens are free.
Latency comparisons are disputed; independent JevBench results are pending.
Check out the official announcement, the Foundry model card and the OpenRouter listing. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.




