Foundations
How Simo works
Simo is a judgment model, not a chat model. Software shows it a situation and asks typed questions. It reads the situation once and returns a calibrated probability for every possible answer, in milliseconds.
The three steps
- Show it a situation. Text, a screenshot, or video. Simo reads it once.
- Ask typed questions. Yes/No, choose one, rate on a scale, fill a value, extract from the input. Ask ten in one request; each adds only milliseconds.
- Read the probabilities. Every decision is a number you can put a threshold on. Where acting needs more than a verdict, Simo supplies the value too.
How an answer is created
A chat model produces an answer by writing words, one after another, and your code has to parse the result. Simo does not write words to decide. For a Yes/No, Choose-one or Rate question, the answer is the probability: no tokens are sampled to reach a decision.
Because every question is typed, the answer arrives in a shape your code already understands. A Choose-one question returns a probability for each option. A Yes/No returns the probability of each side. A scale returns a value on the scale you defined.
Is the task finished? → no: 0.93, yes: 0.07
What should the agent do next? → press: 0.89, scroll: 0.06, stop: 0.05
How frustrated is the customer? → 2.4 of 3Why ten questions cost little more than one
Simo reads the situation once, whether it is a document, a screenshot or a video. Every question then runs as a cheap branch off that single read. Reading is the cost; questions are nearly free.
On Simo-1, one question over a 1280×720 screenshot takes 64 ms and ten take 104 ms, about 10 ms for each added question. See the full latency numbers.
When acting needs a value, not just a verdict
Deciding to press a button is not enough; the agent needs to know which one. For these cases Simo can fill a value (the button, the amount to scroll) or extract from the input (the order ID, the invoice amount). Filled values are schema-constrained, so an integer is always an integer and a list always parses, and each produced value can carry its own calibrated self-check.
Fill and Extract are in preview. Read more on the question types page.
What Simo never does
- It never writes an essay, a chain of thought or a paragraph of reasoning before answering.
- It never returns an unquantified “I’m confident”. Every decision comes with a number.
- It is not the controller. It supervises the agent, the pipeline or the test suite; it does not drive motors or game input.
Frequently asked questions
Is Simo a chatbot?
No. Simo is a judgment model. It answers typed questions with calibrated probabilities and typed values, and it never writes prose.
What can Simo read?
Text, screenshots and video. It reads the situation once and answers every question from that single read.
How fast is it?
Simo-1 answers one question over a 1280×720 screenshot in 64 ms and ten questions in 104 ms. Simo-1 Pro takes 189 ms and 324 ms.
Why does it return probabilities instead of text?
A probability can be thresholded. Act above 0.95, call a reasoning model at 0.60, ask a human below that. Prose cannot be engineered against.