I’ve wanted a simple “inline” decision API for a long time. My criteria are roughly: fast enough to put it in the hot path (inline) of my app’s code, smart enough to make a UX-improving decision usefully, and in a constrained structured output format to resemble an API rather than a mess of chat output.

To solve this, I put this live a couple weeks ago for my own use: rapid.dance

Most AI calls in a product are small. Which repo did the user mean? Is this ticket urgent? What should this thing be called? rapid.dance gives each of those questions its own URL as a JSON API. If the model can’t decide within a defined time limit, you get a default you set instead of an error.

(It works on top of Cerebras to accomplish the speed - I’m not claiming any fancy model work on this one! The “secret sauce” is just a nice UI/API surface on top of a fast LLM.)

So what a surprise to wake up a few days later and see Jev/TypeSafe AI all over the internet solving something similar!

In playing with it a bit this week, I have learned 3 things:

  1. asking questions in this format is harder than I thought, though perhaps there’s some value in the constraint. For example, do I ask 15 questions for “is this a ruby repo?” “is this a golang repo?” etc or do I set them as one question with 15 classification options? How much context should I provide to state? Or, said differently, how do I get it to tell me when it needs more info? Additionally, how do I specify criteria without being so specific it might as well be code (e.g. “does this project have a Dockerfile we can use?” is probably not useful)? I want just enough fuzziness without being too much, and I struggled a bit finding this balance, especially since it can’t extract arbitrary values (e.g. give me the node version if specified).

  2. it is impressively hella fast, especially in its (as recommended per the docs) parallel form

  3. but, it’s kind of difficult to get an idea of its capabilities or “intelligence”. It always gives a selection and its confidence, but what do those probabilities really mean? (And if you’ve ever seen someone say “I knew they’d pull off the comeback!” or participate in any form of gambling/prediction markets, it’s obvious humans are really bad at evaluating what “82% confidence to yes it’s a golang project” really means.)

Browser Control Loop

Since browser control was on my mind this week anyway, I decided to try to hook this together. Extending on point #1 above, I think probably programmatically generating these questions from a surface of e.g. interactive elements is the intended approach. (Though unfortunately it is limited to 255 choices.)

You can grab the full script if you’d like, but the gist of it is to loop (up to 100 steps):

This is certainly not the most sophisticated “agent” one could design, but I think it is valuable to see that 100 lines (of mostly boilerplate, moving state/JSON around) can accomplish this. (See also my previous article on demystifying agent loops.)

Results

Well, it works, and it’s fast! 2.8 seconds to iterate 4 steps (including the actual network fetching as well).

$ time ruby agent.rb 'give me the article about flash cards' https://n.foo
[0] mode=click (0.45) tool=extract (0.81) element=e14 (0.71) link "posts"
[1] mode=click (0.94) tool=extract (0.81) element=e18 (0.95) link "Using Flash Cards for Self-Propaganda"
[2] mode=tool (0.87) tool=extract (0.93) element=e3 (0.48) heading "Using Flash Cards for Self-Propaganda"
[3] mode=tool (0.91) tool=done (0.47) element=e3 (0.37) heading "Using Flash Cards for Self-Propaganda"

Steps taken:
  - clicked e14 (link "posts")
  - clicked e18 (link "Using Flash Cards for Self-Propaganda")
  - text captured for user: ~/n.foo_
about
posts
back to posts
Using Flash Cards for Self-Propaganda

2026-04-18 · Nathan Wong · 2 min read

This is just a shower-thought ...snip blog post...

real    0m2.805s

When playing with it, you can certainly see that this needs to be paired with a traditional LLM for cleanup etc, and there’s a precarious balance between “let it figure it out” that we’re used to in agents and “codify the rules”, e.g. perhaps there should be a question of which element to capture text of and provide the state. But to do that in this model, it isn’t just asking “for the extract tool call, provide the element ref of what to extract”, it requires knowing the structure and providing these each as a choice like the clickable elements.

Comparison - Sonnet

For comparison, I ran this exact same agent loop using Sonnet with constrained output and a system prompt of:

You are a decision model. For each question, assign a probability to every option based on the state. Probabilities for a question sum to 1

(Again obviously not the most intelligent use of a real LLM, but the easiest to substitute into this script)

And it produced almost the exact same thing, but in 22 seconds (8x slower!):

$ time ruby agent.rb 'give me the article about flash cards' https://n.foo
[0] mode=click (0.8) tool=extract (0.01) element=e5 (0.05) link "view all"
[1] mode=click (0.94) tool=done (0.98) element=e18 (0.2) link "Using Flash Cards for Self-Propaganda"
[2] mode=tool (0.8) tool=extract (0.8) element=e3 (0.75) heading "Using Flash Cards for Self-Propaganda"
[3] mode=tool (0.9) tool=done (0.85) element=e1 (0.0) link "~/n.foo_"

Steps taken:
  - clicked e5 (link "view all")
  - clicked e18 (link "Using Flash Cards for Self-Propaganda")
  - text captured for user: ~/n.foo_
about
posts
back to posts
Using Flash Cards for Self-Propaganda
...
real    0m22.830s

Comparison - rapid.dance

I also set this up as a structured query in rapid.dance with an embedded prompt: rapid.dance prompt

Surprisingly, this fumbled around more. It took almost 9 seconds but mostly because it was convinced the about page was where it should find this:

$ time ruby agent.rb 'give me the article about flash cards' https://n.foo
[0] mode=click () tool=click () element=e13 () link "about"
[1] mode=tool () tool=open () element= ()
[2] mode=tool () tool=open () element= ()
[3] mode=click () tool=click () element=e13 () link "about"
[4] mode=tool () tool=open () element= ()
[5] mode=click () tool=click () element=e13 () link "about"
[6] mode=tool () tool=open () element= ()
[7] mode=click () tool=click () element=e13 () link "about"
[8] mode=click () tool=click () element=e6 () link "about"
[9] mode=click () tool=click () element=e7 () link "posts"
[10] mode=click () tool=click () element=e18 () link "Using Flash Cards for Self-Propaganda"
[11] mode=tool () tool=extract () element=e3 () heading "Using Flash Cards for Self-Propaganda"
[12] mode=tool () tool=done () element= ()

Steps taken:
  - clicked e13 (link "about")
  - opened start URL https://n.foo
  - opened start URL https://n.foo
  - clicked e13 (link "about")
  - opened start URL https://n.foo
  - clicked e13 (link "about")
  - opened start URL https://n.foo
  - clicked e13 (link "about")
  - clicked e6 (link "about")
  - clicked e7 (link "posts")
  - clicked e18 (link "Using Flash Cards for Self-Propaganda")
  - text captured for user from e3: Using Flash Cards for Self-Propaganda

real    0m8.944s

One difference is the LLM’s “thinking” can be used to help debug (in this case, in the history of the API log): thinking “probably need to click” lol

Comparison - OpenCode

To see how well a more “sophisticated” harness/agent loop could work with the agent-browser skill (as previously discussed in the browser control article), I ran this in OpenCode with DeepSeek Flash v4.1 and it took 1m36s (yikes!). However funny enough, without specifically asking it to use agent-browser, it defaulted to just using WebFetch (which works for this simple example without more complicated interactions) and took 36s: WebFetch via OpenCode

Open Models

I tried 7 open models (from this surprisingly long list), but it kind of felt like some of these were “write-once, run never”. I won’t name names, but for example, one tried to download a file that absolutely does not exist (even though it referenced a repo that does exist and hasn’t changed in months). One was my fault when I didn’t want to deal with upgrading dependencies (good for it to tell me, so props there). And three of them did run and answer questions properly, but they got lost in the agent-browser test loop (one was confident that clicking the home page and calling it done was sufficient, another clicked just about every link).

I did not dig into how the benchmark was calculated so perhaps our question formats were just significantly different. At the risk of sounding like I’m shilling for Jev, though, it’s clear they’ve at least spent some effort to ensure the experience functionally works and feels good.

All that to say, I was hoping to find an open model that would run similarly and build it in as a toggle in rapid.dance, but no luck. It would be handy to be able to route between LLM vs classifier for open-ended vs closed questions (e.g. “extract the node version” vs “is this node 19, is this node 20, is this node 22”). Perhaps the models will evolve a bit with time and I can come back to that; save that for another blog post…