Use cases · flagship
Agent supervision: done-ness, next action and risk, one call per step
An agent that acts needs a judgment at every step: have we finished, what comes next, which control, is it safe. Simo answers all of it from one screenshot in one request.
The problem
Browser and computer-use agents loop: look, decide, act. Asking a reasoning model for each decision means seconds of waiting and a few thousand words to parse, for a loop that may run dozens of times per task. And without a number, you cannot tell a confident press from a guess.
What you ask, and what comes back
The agent sends a screenshot and asks four questions in one request.
done? → done: 0.04
next action? → press: 0.89
which button? → button: "Sign up" (preview)
how risky? → risk: low
The agent acts. The loop continues.Why it works
- Verdict, arguments and confidence in one frame: this is System 1.5 applied to an agent step.
- Ten questions cost little more than one, so adding a safety question to every step is nearly free.
- Calibrated probabilities let you gate: press automatically above a high threshold, pause and ask a human below it. See calibrated probabilities.
Which model
Simo-1 runs the loop. Pointing at the exact spot on screen (coordinates) uses Simo-1 Pro. Arguments and coordinates are in preview.
Frequently asked questions
Can Simo control my agent?
Simo supervises the agent: it judges whether the task is done, what to do next and how risky it is. It is not the controller and does not drive actions itself.
How fast is a step?
On Simo-1, one question over a 1280×720 screenshot takes 64 ms and ten take 104 ms.