Watch an agent built on a chat model and you'll see the same pattern. It reads the page, writes a paragraph about what it sees, decides somewhere in the middle of that paragraph, and then writes the text it means to type. Your code fishes the action and the text out of the prose and hopes they agree.
That gets the order wrong. The text an agent types is the consequence of a decision. It should come after the decision, be written with the decision in hand, and come back in a shape your code can check.
The problem with deciding in prose#
- You can't see how sure it is. A sentence that says "I'll fill in the origin city" reads the same whether the model was certain or guessing.
- The words come first. When the text is written before the action is settled, the action tends to follow the text, not the other way round.
- Parsing is a guess. Even with JSON mode, the format is constrained but the decision inside it is still a token in a string, with no probability attached.
Decide first#
With Wity, the decision is its own call. You list the actions that make sense, Wity picks one, and you get a probability for every option. If the top choice isn't clear enough, you stop there: nothing gets typed.
next_step = decide(page, {"next": {"type": "choice","instructions": "What should the agent do next on this booking form?","criteria": {"type_origin": "Fill in the origin city","type_date": "Fill in the travel date","submit": "Everything is filled in: submit","stuck": "The page doesn't match the task",},}})["next"]# next_step["choice"] == "type_origin", probability 0.93
Then generate, from the same state#
Only now does the agent need words. generate reads the same state the decision did, and the instructions follow from the action that was chosen. The output is held to a shape while it's written: a pattern for a single value, or a JSON Schema for several fields. It always parses, so there's no retry loop and no "please answer in JSON".
if next_step["choice"] == "type_origin":city = generate(page,"The origin city for this trip, as the traveller wrote it.",shape={"type": "string", "pattern": "^[A-Za-z .'-]{1,40}$"},max_tokens=16,)browser.type("#origin", city["value"]) # "Lisbon"
The two calls share one model and one context. The decision isn't passed through a second model that has to re-read the page and might see it differently, and the text can't drift into a different action than the one you approved.
Short text, on purpose#
generate is for the words an action needs: a value for a form field, a city, an invoice number read off a scan, a one-line reply. Output is capped at 100 tokens. If the answer is one of a set you could have listed, use choice instead, because a decision gives you probabilities and generated text doesn't.
Check it like any other input
max_tokens?Where this helps#
- Browser and desktop agents: pick the next step, then the value to type into the field.
- Document intake: decide what a scan is, then pull out the fields that kind of document has.
- Support: route the ticket, then write the one-line acknowledgement for that team.
The generate docs walk through a full example, from classifying a supplier email to sending a checked reply.
