Foundations
Calibrated probabilities: when Simo says 0.8, it is right about 80% of the time
Every decision Simo makes carries a calibrated probability. That turns “the model thinks so” into an engineering threshold: act above it, escalate to a human below it.
What calibrated means
A probability is calibrated when it matches how often the model is actually right. Take every call where Simo said 0.8: about 80% of them are correct. Take every call at 0.95: about 95% are correct.
Calibration describes Simo’s decision probabilities and self-check scores. It is not a measure of how fluent or sure a sentence sounds, and it is not a raw generation confidence.
Why you want it
An uncalibrated score is only a ranking. A calibrated one is a rate you can plan around: if you act at 0.95, you know roughly how often those automatic actions will be wrong, and you can decide whether that is acceptable for the cost of a mistake.
- Above the act threshold (for example 0.95): the action executes on Simo’s answer alone.
- In the middle: the item goes to a review queue, or to a reasoning model, the only time the expensive call is made.
- Below the ask threshold (for example 0.60): it goes to a human, with the probabilities attached.
An example gate
| Proposed action | Simo’s probability | What happens |
|---|---|---|
| Press “Place order” | 0.98 | Executes |
| Route ticket to billing | 0.97 | Executes |
| Remove comment under policy 4.2 | 0.89 | Review queue |
| Approve refund of 40.00 | 0.55 | To a human |
| Close the account | 0.31 | To a human |
Move the thresholds and the split moves with them. That tunable dial is the point: the cost of an error differs per product, and so should the threshold.
Measured calibration
On Banking77 (route a customer message to the right intent), the calibration error of Simo-1 Pro’s Choice probabilities is 0.047. See the accuracy page for the full scores.
A long tradition
Forecasters have been scored on stated probabilities since 1950, when Glenn Brier proposed grading weather forecasts that way (Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3). Simo brings the same discipline to software decisions.
Frequently asked questions
What does a calibrated probability mean?
That the stated probability matches the real hit rate. Of the calls Simo scores at 0.8, about 80% are right.
What threshold should I use?
It depends on the cost of an error. A common pattern is to act at 0.95, call a reasoning model at 0.60, and ask a human below that. Because the probabilities are calibrated, you can tune the line to your own risk.
What is Simo’s calibration error?
On Banking77, the expected calibration error of Simo-1 Pro’s Choice probabilities is 0.047.