how it works
How fastbrowse runs a task.
An LLM plans the task while the browser opens its start page. Jev chooses from indexed controls and judges whether to read, act or finish. Code checks authorization, copies every quote from the page capture and builds the answer's citations. Model responses use strict structured output.
back to fastbrowse.ai·Browser architecture
The step loop
Observe, choose, then act or read. Each step returns to observation; choosing Done starts a completion check.
01Plan and open
The LLM turns the task into checkable requirements while the start page loads. If no start URL is supplied, an LLM proposes one. Planning can continue alongside the first actions; reading and completion wait for the requirements.
02Observe and choose
The page becomes indexed controls and a text excerpt. A batched Jev request chooses an operation and targets, checks access walls, and judges whether the page holds evidence still needed by the task. When it does, code reads before further interaction can remove it. Read deduplication keys on the document, captured content and unanswered requirements.
03Check an action
Every click, including links, and Enter that submits a form is eligible for Jev's irreversibility judgment. Dialog acceptance is judged separately. Code-driven pagination through a next-page link is explicitly exempt; authorized, confident actions also skip the extra call. A committing action without authorization stops at needs_confirmation; an uncertain choice goes to recovery. A refusal is recorded as a failed step with its reason. The classifier can miss an irreversible action.
Before dispatch, code checks that the target is current and reachable. Secrets are resolved only for an allowed origin at typing time; models receive their names. The loop observes the effect of an action before deciding what to do next.
04Read into notes
Jev can choose a short fact from candidates code cut from the capture. Synthesis, comparisons and uncertain answer shapes go to the LLM reader, which cites capture blocks by id; code copies the quote from the capture, so the model never retypes page text. A card, or a table row with its header, is one block, so a record is read and cited whole.
A count or winner the page does not state is a derived fact with no quote, backed by the records it draws on. Each fact records its requirement, source and reader, jev_choice or llm. Evidence from a continuing list stays partial until the required pages have been read.
05Verify and cite
Jev checks completion against the task, page and notes; an uncertain verdict goes to the LLM verifier. The answer is drafted from the notes or written by the LLM. The claim check judges each claim against its quotes, and a derived fact from the records it draws on. Failed claim checks return unverified rather than complete.
Code adds numbered Markdown links to each claim from verified notes, and links only to quoted page text. RunResult.citations exposes the cited facts, requirement IDs, source URLs, quotes and text-fragment deep links that highlight the quoted words; a derived fact has no quote or link of its own. An answer citing an unknown reference fails the claim check and falls back to one drafted from verified facts. Resolved secrets are redacted before links are built for delivery. Highlighting depends on the browser and the source still containing the text.
- Budgets and recovery
Prompt limits live in Config.tokens and Config.observation; Limits bounds steps, model calls, time and model spend. Cloud browser charges are added after the session closes. Verdict prompts preserve requirement evidence. If it does not fit, the completion check cuts page text first, then stops with observation_limit if the evidence still cannot fit. Cut text carries an omission marker when there is room; otherwise the excerpt is empty. Navigation uses a shorter working memory whose omissions are marked.
An action that makes no progress triggers recovery. Repeated-action and stalled-plan checks log where they would trigger recovery by default, without interrupting the run. Config.stall controls their thresholds and mode. Recovery is bounded; exhausting it ends the run as stuck.
- Step events and live frames
run_task(on_event=...) delivers StepEvent.step with the decider (decided_by), confidence, outcome, reason or detail (note), and facts found by that step. Each fact carries its reader, quote and source link, with resolved secrets redacted.
run_task(on_frame=...) streams JPEG bytes of the active tab over CDP screencast and follows tab switches. Delivered frames are acknowledged after the handler returns, so the rate follows the consumer; superseded frames can be skipped. Delivery runs separately from the task, and handler failures are logged. Capture is off without a handler, and the consumer receives images without access to the CDP endpoint. Live frames and recordings are held back while a resolved secret may show on the page, as PNG step frames are.
- Jev provider failover
Jev uses direct TypeSafe when its key is supplied; otherwise OpenRouter is primary, followed by Vercel AI Gateway when neither other key is available. FASTBROWSE_JEV_SOURCE explicitly selects a primary. OpenRouter and direct TypeSafe use a configured Vercel key as backup; Vercel uses direct TypeSafe when keyed, otherwise OpenRouter. With default endpoints and models, an exhausted retryable HTTP status switches to that backup. The run stays on the backup. Authentication errors, malformed answers and transport failures without a final retryable HTTP status do not switch providers. A custom endpoint or model disables automatic failover. All three routes reach TypeSafe, so failover cannot promise availability.
See task success, time, cost and test settings on benchmarks.