changelog
Release notes and work in progress.
Changes are grouped by version. The Unreleased section lists changes planned for a future release.
back to fastbrowse.ai·CHANGELOG.md
Unreleased
- Abandoned model calls retain charges from responses that arrive within their attempt deadline, including matching provider receipts. Dollar caps still stop runs whose charges cannot be established.
0.5.20 - 2026-10-09
-
Going back from a document that has not been read reads it first when requirements remain, so a pending action can use values from that page after leaving it.
-
Answer checks retain a comparison's collection scope and completeness alongside its quoted records, so a whole-list winner keeps the coverage established during reading.
-
Answer checks retain whether a quote covers a complete downloaded file, so totals can be verified from all its records. Partial excerpts and truncated downloads do not claim complete file coverage.
-
Updated page evidence retires vanished context quotes as well as previous answers, so an earlier subtotal does not conflict with the subtotal after the basket changes. Quotes still shown remain available to checks.
-
A confident request for recovery follows the pending recovery action when it is still valid, so a run can leave a page with no usable controls without exhausting recovery on the same state.
-
Answer checks through the Vercel AI Gateway receive each subject's identifying quote, so supported answers finish as
completethere as they do through OpenRouter. Unsupported answers stay rejected on both. -
Reply controls retain their own comment record as context when only one reply form is open.
-
Provider attempt deadlines bound total response time, including responses that keep sending occasional bytes.
-
Browser commands that remain unanswered for 120 seconds return
unavailableeven if health checks succeed. -
Closing the last owned browser tab returns
unavailablewith retained run evidence instead of a Python error. -
Styled upload labels expose their hidden file inputs as upload targets and receive supplied attachments.
-
Page captures retain visible image DOM metadata, including source and alt text.
-
Missing completion charges are recovered from matching provider generation receipts when available. Dollar caps still stop the run when the charge cannot be established.
-
Optional
--cursorfeedback draws an agent overlay for visible local Chrome on Linux X11 with Cua Driver. It leaves the physical pointer and focus alone, hides when its page or window changes, and skips unverified window mappings. Cloud, headless, Wayland and remote Chrome connections run without it. -
Read captures retain empty form fields, disabled buttons, horizontal overflow and the absence of visible h1 headings across accessible document and shadow roots. These DOM observations cite the source page without text fragments for synthetic labels. Field values supplied as URLs are checked before navigation.
-
Recordings of attached windows leave the caller's page in place. Explicit submit buttons expose the form's destination to the authorization check, while focusing a field does not claim to submit it.
-
Captured hover captions retain their DOM container and sibling position as source metadata. Answer checks can bind positional subjects without treating that context as a quoted value.
-
Answer identity checks emit each literal identifying quote once and reference it across fields. Repeated names no longer consume output tokens for every requested value; citation and subject checks remain.
-
Answer audits consider conflicting current quotes from the same source address and frame, including uncited notes. Conflicting values must be acknowledged rather than presented as one settled value; those other quotes cannot replace a claim's own citation support.
-
Consecutive evidence notes share their exact source URL when it makes the rendered context smaller. Long referral addresses no longer repeat for every fact, leaving more room for required quotes and their basis.
-
Repeated reads keep one copy of each identical prior-evidence packet in the novelty check. Distinct quotes, summaries, frames, headings and changes remain separate, and every stored fact and citation is retained. Duplicate context no longer crowds this check out of its input budget.
-
The agent skill uses uncapped resource defaults, with limits supplied by the caller's task budget. It no longer imposes a 30-step limit or a $0.25 model-spend cap on each delegated task.
-
Revealed hover targets keep their labels and sibling positions while their captions are visible.
-
Final answer audits retain count scope, collection completeness and the quoted records behind each tally.
-
Recovery checks whether a changed page satisfies the task before stopping for lack of further actions. The completion verifier still rejects unfinished work.
-
Requests to expand or load content until a stopping condition keep that page change as an action requirement. A displayed total no longer replaces the requested work in the plan. The feed fixture also checks that every remaining batch was loaded, and advances to task version 6.
-
A ranking shortcut requires a quoted statement of the active list order. Offered sort links no longer close a comparison; rejected shortcuts collect the current records before paging onward. Numbered links to the next page use the same paging path when their address preserves the list filters.
-
Downloaded CSV files retain table structure, so row citations keep the column names needed to interpret their values. Bounded downloads still report omitted file content.
-
Count claim checks retain the requested scope and whether the accumulated tally is complete. An ungrouped count no longer reaches the checker as an unlabeled number.
-
Planning uses the default reading model and separates sign-in identities from requested data filters. A prerequisite account name no longer becomes an extra field in an exported report's answer checks.
-
Answer verification rejects superseded facts from the same page, including calculated facts that remain in browsing history after their requirement evidence is replaced.
-
A slow shortcut can finish reporting its model cost while browsing continues from the start page. Falling back no longer cancels a dispatched request and interrupts a run with a dollar limit. A late cost that exceeds the limit stops the run while retaining its already verified answer and data.
-
Answer checks judge calculated results from their quoted operands. Reader lineage no longer blocks a calculation when its complete source records support it.
-
Scalar reads retain cited headings that identify their values, so answer verification can bind a price or count to its named subject. Heading evidence stays within its frame.
-
Answer verification requires a response for each requested check. Structured model responses cannot omit identity bindings or field judgments through empty maps.
-
Completion checks preserve each requested item's fields even when a draft plan lists generic field names. Quoted identities keep one item's value from answering another item's field, including compound claims. An identity can come from the quoted page body or its captured title.
-
Cancelled Jev requests retain estimated input costs instead of disappearing from the spend ledger. Cancelled LLM requests keep reported earlier charges and mark missing billing as unknown.
-
Older quotes retain their original page title when unchanged text is captured under a new title.
-
Quoted excerpts can include known source-block labels copied by the reader. Those labels are removed before matching real page text, while unknown labels and ambiguous quotes remain rejected.
-
Navigation headings stay within their menu instead of identifying later page facts. Navigation notes show an explicit prefix for long source addresses, leaving room for previously collected facts.
-
Answer repairs reuse quoted notes when selected citations omit available evidence. Each repair must improve source coverage, or reduce assertion failures without worsening source coverage. Identical audit requests are reused within a run; changed evidence, entities, claims or scope are checked again. Component values cannot stand in for totals, and rewritten conclusions do not restore the read budget.
-
Completion checks each requested answer output against its cited sources. Missing or ambiguous outputs return to recovery, so a partially answered requirement cannot finish from its other cited fields.
-
Identical captured spans on different addresses retain separate citations and counted records.
-
Headerless table rows retain the real leading cells identifying their columns and keep separate citations for counting.
-
Removing a doubted answer claim checks whether its requested output is missing, even when other fields still cite the same requirement. Optional details can still be removed.
-
Prose reads use larger source chunks to keep navigation and following details in one read. Quote checks and the smaller choice and counting budgets remain in place.
-
Navigation links keep their stale-attempt history when card labels or context change, so a redraw cannot restore a repeatedly failing destination. Separate in-page actions retain their own retry history.
-
run_taskandconnect_cdptakeallowed_origins, a list of exacthttp(s)://host[:port]origins that limits the documents a run may inspect and control. Navigations, redirects, frames (in or out of process) and popups outside the list are refused before they are sent, and a window or frame already open outside it is never read, captured or screenshotted. It scopes documents, not the network: images, scripts and requests a granted page makes to other hosts still load. Service workers are bypassed so every navigation meets the check, a scoped run delivers no live frames or screenshots, andrecordis refused with it. Opaque child documents are excluded from text and controls. Without it nothing changes. -
run_taskandconnect_cdptakecheck_access, an async callback awaited before browser startup and every observation, capture, screenshot, address, navigation and action. If it raises, the run stops with aBrowserErrorthat keeps only the exception's type. -
Evaluation evidence can carry token-based API-equivalent cost ranges separately from recorded spend.
-
Runs no longer stop at implicit step, Jev-call, LLM-call, dollar or time ceilings. Callers can still set explicit limits when they need them.
-
Complete benchmark suites can be published independently. Existing published suites stay present, and each suite retains its coverage, evidence and regression checks.
-
Navigation notes write each source URL once and reference it from the remaining claims, preserving full addresses while leaving more room for product details and progress.
-
Identical control relevance questions share one Jev score within each pass. Controls remain separate action targets, with fresh selection and authorization.
-
Rewording an already collected source quote no longer counts as read progress, so timer redraws and paraphrased claims cannot keep restoring a barren page's read budget. Newly evidenced requirements still count.
-
Links that repeatedly fail their freshness check leave the choices for that document, so intervening reads cannot keep retrying the same stale link. Recovery can choose another route without weakening input guards.
-
Prose reads can cite a verified excerpt within a large source block, so later reads carry the supporting passage instead of unrelated page text. Code copies the original span and rejects missing or ambiguous matches.
-
Independent answer claim checks run in parallel batches. Each omission check keeps its own quoted-evidence budget instead of sharing it with unrelated claim questions. Missing or failed checks cannot verify an answer.
-
Action decisions keep recent progress claims when a source quote exceeds their notes budget, so collecting a large quote does not hide which items were already checked. Reads and completion checks retain full quotes.
-
Page redraws retain completed Jev relevance scores for unchanged controls in the same document and task context. Unfinished batches still run in parallel, and every action is chosen from the fresh page.
-
Collected evidence includes a verbatim source quote once when the fact text is identical, reducing repeated input in later reads and checks while preserving source quotes and citation links.
-
The internal benchmark catalog includes the opt-in Codex and Cua Driver comparison arm, run by Parallax.
-
Final screenshots allow the fresh page check to finish its loading wait before capture. Loading pages retain their ending frame, with visible secrets still withheld.
-
A directly quoted total can answer a count after an earlier record tally was incomplete. Derived counts still require complete records, and page scope and answer claims remain verified.
-
Runs that stop on a blocked page, an error or a limit retain a final screenshot when step frames are enabled. The capture checks the current page for secrets and preserves the run status.
-
Jev chooses repeated short facts as one value, retaining every source quote and context. Duplicate locations no longer split its confidence and send an otherwise clear lookup to the prose reader. Headerless tables describe values by their row rather than inventing column headers.
-
A run whose form submission the server answered with an error no longer ends
completebecause a later page loaded cleanly. A message form that answered 503 with a page still saying "your message has been sent" was reloaded, the reload answered 200, and the run reported the message delivered. The last submission's status now counts: an error there ends the run the way an error on the final page does, unless a later submission succeeded. -
Icon-only buttons use recognized SVG glyph names when other labels are missing. Button-styled links expose one clickable control. Empty plans retain the requested outcome, and required CI checks reject failed jobs.
-
Evaluation pages report grades, completion, retries, unknown costs and missing coverage separately.
-
Headed sections preserve styled pairs as separate records, so counts can distinguish groups from their items. Covered-target recovery excludes controls belonging to the target's own fixed container.
-
Task-supplied sign-in addresses open directly. Blank opening pages get one reload before any interaction, so a cold app can recover without repeating a submission. Shortcuts leave unspecified filters to observed controls, and recovery can return toward the caller's start page when a guessed address makes no progress.
-
Form controls retain their input name and autocomplete purpose after validation changes their labels. Recovery corrections are bound to the observed field and reuse only supplied, non-secret values.
-
Tasks inspecting access restrictions can read a sign-in wall or HTTP access denial as evidence. Protected retrieval still requires login, and bot challenges still block the run.
-
Record counts use quoted child records instead of a scalar group total. Covered clicks report the visible obstruction, and requested transactions have their outcome read before another interaction. Malformed tally field ranges get one repair attempt on the saved capture before later counts are marked incomplete.
-
External capsule endpoints preserve cookie security attributes while mapping their declared hostnames. Strict state supplements use independently sampled sequences and leave missing evidence ungraded.
-
Every release requires a successful Evals run on its exact commit. A release with no published comparison once skipped the fixture gate and could ship agent code the Evals suites had not passed.
-
The LLM can use Vercel AI Gateway when no OpenRouter key is configured. The CLI, MCP server and fixture evals share that routing, and scheduled evals check their credentials before starting a task.
-
fastbrowse serve, which the JavaScript SDK starts, acceptsAI_GATEWAY_API_KEYalone too. Its startup check still asked forOPENROUTER_API_KEYand refused a run the LLM could have served through the gateway. -
External eval corpora can run pinned tasks with reproducible selection, attempt records and independent grades. Ungraded attempts retain an unknown pass state instead of contributing to a benchmark score.
-
Fixture evals leave a page's scripted cookie banner to the agent. The browser layer's own consent refusal had removed the mock support portal's banner before the agent could accept it, so that task's requirement could not be satisfied.
-
Published eval comparisons require a clean build, complete task coverage, at least three measured repeats and a ledger retaining every attempt and its spend. CI refuses missing evidence and matched correctness, latency or cost regressions. A release needs successful fixture evals and measurements from the same code.
-
Interrupted live comparisons retain in-flight attempts and wait for active browsers to stop before closing the ledger. Spend that cannot be recovered stays unknown instead of disappearing from the report.
-
External corpus comparisons reset each site before every attempt and report completion separately from independent correctness. WindTunnel tasks marked excluded by upstream are left out, and authenticated fixtures use named credentials scoped to their site.
-
External comparison figures require three paired repeats for each included task. Scattered partial repeats cannot combine into a score, and stopped draws retain diagnostics without a comparison headline. A completed draw also refuses a figure when a task every requested arm could attempt has an ungraded or missing repeat, instead of dropping that task after its grades are known.
-
Structured extraction reads scalar values from every page the run quoted, not only the one it ended on. A sorted listing's later page holds the pricier products, so a price quoted on an earlier page had no candidate and the data came back empty while the answer was right. A quote the final page corrected is not offered as a current value, and a candidate pool wider than Jev's option ceiling is grouped rather than dropped.
-
On a previously read page, Jev checks for new evidence alongside short-fact choices. A confident duplicate assessment skips the prose read; changed values and uncertain assessments still reach a reader. Runs that hit a money or time limit return collected facts with citations and keep
budget_exceededstatus. Provider and browser errors also preserve collected evidence with their failure status. -
A portable fastbrowse skill lets Codex and Claude Code delegate website tasks through MCP or the CLI. Coding agents prefer an existing local Chrome connection, then local Chrome, with cloud as a fallback. The skill guide covers installation, configuration, citations, and authorization.
0.5.19 - 2026-10-06
-
run_tasktakescloud_extensions, up to three distinct extension IDs from your Browser Use Cloud account, and loads them into the cloud browser it starts. Asking for extensions with local Chrome or an attached browser is an error, since neither can load them. Runs that ask for none are unchanged. -
fastbrowse can now be used from JavaScript with no Python on the machine.
npm install fastbrowseinstalls a TypeScript SDK and, as an optional dependency, the agent as a native binary for macOS, Linux (glibc) or Windows. No install script runs and nothing is downloaded on first use.Fastbrowse.start()starts the binary andrun(task, options)returns theRunResulta Python caller gets, with the same statuses, citations and cost lines. The secret resolver,until,onEventandonFrameare functions in your own program, so a secret value stays there until the moment it is typed, and anAbortSignalcancels a run and closes its browser.outputtakes a Zod, ArkType or other Standard Schema that can write itself as JSON Schema, or a JSON Schema object; with the first,result.outputis the data as that schema validated it, and has its type. A run fills flat string, number, integer and boolean fields today, and a schema with any other field endsunverified. The README has the install and a first run. -
fastbrowse serve --stdiois the entry point the SDK drives, and any program can speak to it: JSON-RPC 2.0, one message per line, with the methodsinitialize,run,run/cancelandshutdown. It runs one task at a time and takes its keys from the environment, as the command line does.run_task, the command line's existing flags and the MCP server are unchanged. -
One tag now publishes one version to PyPI and to npm. The release builds all five binaries and runs each through a handshake and a task on local Chrome before anything is published. npm publishing uses trusted publishing, as PyPI does, so no token is stored. CI fails when the npm version differs from the Python one. The macOS binaries carry an ad-hoc signature and are not notarized, the Windows binary is not signed, and there is no binary for Alpine or another musl system.
0.5.18 - 2026-10-05
- A run can now take over a window that is already open, in Chrome or in an Electron app such as VS Code or Slack.
--attach(orattach=Trueinrun_task) skips opening a tab and drives an existing page, and--target-match TEXTpicks the first page whose title or URL contains TEXT. The window stays open after the run, and with no--startthe run begins on whatever the window already shows.--cdp-port PORTfinds the browser from its DevTools port, which is how an Electron app started with--remote-debugging-portis reached.connect_cdp()hands a script the same attached page, without the agent. Windows an Electron app opens on its own join the run only with--target-match, because in a browser such a window is as likely a tab someone opened by hand.
0.5.17 - 2026-10-03
-
The mock site's sign-out page now says "You are signed out." instead of sending the browser to the home page, as real sites do. Without that statement a run had no evidence that the sign-out worked, so
mock-sign-outoften ended unverified even when it had signed out correctly. -
A run that only had to act, such as sending a form, is no longer marked unverified because its answer restated what it typed. When the run read the confirmation ("Thanks, we received it.") and its answer also named the sender and topic, the claim check found the quote did not show those details and failed the whole run, though the completion check had already confirmed the submit. Those claims are now dropped instead, as for any doubted claim, and a task with nothing to report still finishes.
-
A drag whose drop target shifts under the pointer, such as a column that nudges or moves when a card hovers over it, now drops on that same control at its new place instead of being cancelled. A drag whose target turns into a different control is cancelled outright, so not even the card's own list receives a drop, where before letting go at the start could still drop the card there. A page that runs its own pointer dragging rather than the browser's has no such cancel; for those, Escape and letting go where the card was picked up remain the fallback.
0.5.16 - 2026-10-02
-
Drag and drop. A run can drag one control onto another, such as a card onto a board column, when the page offers both ends. A drag is checked like a click before it is made: one that would commit something that cannot be undone stops for confirmation unless the run is authorized, and a drop that changes nothing sends the run to recovery instead of being repeated. Both ends are rechecked the moment before pressing, and a drag whose target changed is decided again from the new page instead of being retried. A drop target that changes while the card is in the air is not dropped on: the card is carried back and let go where it was picked up. A drag between a page and a frame hosted separately from it is refused, as is a drag onto its own source.
-
The fixture eval suites (local and mock) run on a schedule and on demand in
.github/workflows/evals.yml, and the job fails on any regression, so a change to browsing behaviour that breaks a task no longer waits for someone to run the suite by hand. The job reports a per-task pass rate over its repeats and refuses a run where any task is below the target. -
An authorized run can submit a form after attaching a file. Unsure whether the submit was the commit the task meant, it asked recovery, and was then refused the same click again whatever recovery said, until it ended
stuckwith nothing sent. When recovery, shown the page and the task, names the same control that was refused, the run may now make that click once. A credential change is still never let through this way, and a run without authorization still stops for confirmation. -
The stateful mock suite grows from 20 tasks to 26: a story opened in a second tab, a native date input, a newsletter form with a hidden field a careful run leaves alone, an export downloaded from behind a sign-in, a document attached to a form and submitted, and a card dragged from one board column to another. A run may now carry attachments, the way a caller supplies a file.
-
Answer grading folds case and typographic punctuation before it matches, so a run that words a right answer differently is not failed for its phrasing, and the runner warns when a task uses more than 60% of its step budget.
-
docs/evals.mdstates how to read a published number, including the caveats that have to travel with one. -
A run asked for something the page does not have, such as the heading of a page with none, now stops after one recovery instead of spending all of them. Recovery used to answer "report that it is absent and finish", the completion check refuses any finish while requested information is unread, and the two repeated until the run ended
stuck. A recovery that can only direct such a finish now ends the run with its diagnosis as the error. -
A
stuckrun returns what it did read. Its answer carries the cited facts gathered before it stopped, marked as partial, with their evidence and citations, so a caller no longer has to browse again for them.
0.5.15 - 2026-10-02
- A request to report the final page's title or to capture a screenshot is now a run report filled in from
browser state, not a requirement the page must evidence. The title is not page text a reader can quote, and
fastbrowse has no screenshot action, so such runs used to end
stuckwith completion rejected. The answer now carriesPage title: ..., and a requested screenshot is returned asfinal_frameeven without step frames, with the answer saying whether it was captured. A run that could not take the screenshot it was asked for endsunverified, notcomplete. The CLI prints where it saved the image, and the MCP result carries it asscreenshot, in the downloads directory or a temporary one.
0.5.14 - 2026-09-29
-
Navigation comparisons enable cloud resizing to match Ultrafast's viewport and retain the actual dimensions for both arms. Publishing rejects scored runs with missing or different dimensions. Ultrafast final evidence is read before its browser driver disconnects, so cleanup cannot change the page being graded.
-
Completion checks read fresh page state after a bounded loading wait, so asynchronous search results cannot be judged from the earlier page before they arrive.
-
Navigation evals retain the final document status and content, so an error page at the requested URL cannot pass. Browser transport timeouts are retried for both arms, and separate site probes no longer discard completed runs. Ultrafast stale decisions do not consume its action budget.
-
Opening a page must leave the requested destination showing. HTTP error pages no longer count as successful visits, and action effects report the failed response to completion checks.
-
Same-named controls can be distinguished by their own visible descriptions when they share an ancestor, instead of falling back to position alone.
-
The Ultrafast eval adapter recognizes the same provider outage statuses as fastbrowse, including Cloudflare 520-524 responses, explicit provider retry markers and proxy transport failures, so outages are retried instead of counted as agent failures.
-
Live evals retain every outage retry in a separate attempt ledger and keep separate retry recordings. Arm launch order rotates between repeats. Slow completed attempts still count with their full time and cost.
-
Eval reports name Browser Use Ultrafast (
jev-ultrafast) separately from the hosted Browser Use agent. The eval guide includes the command for their shared navigation comparison through OpenRouter. -
The eval summary retains prior task versions across releases that publish only other suites, so returning tasks report their version changes instead of appearing new.
0.5.13 - 2026-09-29
- Embedders that enable step frames receive a safe final page PNG in
RunResult.final_frame, so their completed preview can show the destination instead of the page before the last action.
0.5.12 - 2026-09-29
- Requested final URLs and navigation steps come from the browser record, so a run can finish on an image without trying to find page quotes for its own actions. Page facts still require captured evidence.
0.5.11 - 2026-09-28
- The agent can scroll up to revisit content above the viewport. Recovery no longer sends every scroll downward.
0.5.10 - 2026-09-28
-
Required native radio groups stop blocking once an option in the same form is selected.
-
Step frames are PNG images of the page before the action, with a fresh secret-visibility check.
-
Eval commands reject nonpositive concurrency and repeat counts before starting a run.
-
Eval scores retain slow provider calls and recovered request failures, with full time and cost, for every arm. Task versions include the shared grading rule; historical published results keep their original selection bias.
-
Keep HTTP error status in browsing observations and recovery. Failed links are withheld on unchanged source pages, and an unfinished run reports the observed session failure rather than a guessed site restriction.
-
Empty reads no longer excuse repeated navigation or form cycles; reads that add evidence still do.
-
Retire cancelled CDP waits so their late replies are not reported as duplicate responses.
-
A supplied
TYPESAFE_API_KEYselects direct Jev; otherwise OpenRouter is primary, with Vercel AI Gateway supported as a backup or explicit primary.FASTBROWSE_JEV_SOURCEoverrides automatic selection. -
Compare fastbrowse and hosted Browser Use on stateful mock workflows with shared site-state grading, explicit authorization instructions, fresh cloud browsers and recorded per-attempt results.
-
Password changes require a separate
new_passwordsecret and authorization to submit. Replacement values never pass through a model, and current passwords are not reused as replacements. -
Downloaded text is readable with citations to the response URL. Model context is bounded; artifacts retain the full file. Download navigations no longer fail with
ERR_ABORTED. -
Counts accept complete leaf list items and fields at record boundaries, while missing page evidence still blocks completion. Reads retain quoted context needed for later actions, including a code read on another page.
-
An audit suite for the terminal, MCP and embed surfaces.
python -m fastbrowse.auditruns a tiered matrix of contract and behaviour cases, each carrying the status, exit code and per-check evidence it must produce, and writes one JSON report plus a Markdown roll-up. Tier 0 covers the terminal contract without calling a model; the paid tiers are refused unless--spendis passed. -
The CLI assigns a distinct nonzero exit code to each stopped status and documents the mapping in
--help. JSON, embedding and MCP results report the resource and limit behind a budget stop. -
--cloudoverrides local browser defaults from the environment. Secret configuration errors report both missing variables and missing origins, and a login without stored credentials explains the scoped--secretsyntax. -
Read nested gateway errors, recover from input-size refusals, and redact echoed API keys before shortening errors.
0.5.9 - 2026-09-27
- Jev defaults to OpenRouter using the LLM's existing key.
OPENROUTER_API_KEYalone now covers both models. SetAI_GATEWAY_API_KEYfor automatic Jev failover to the Vercel AI Gateway after retries run out. OpenRouter's reported Jev cost is metered, with a list-price estimate when cost is missing or zero while input tokens flow. Explicit model pins and custom endpoints disable failover.
0.5.8 - 2026-09-27
- The date given to every choice names the next two months. Told only a September date, Jev once took "the month after next" as October and clicked a calendar's Next once. Today's date now comes with the month one and two months on.
- A corrected choice replaces its old reading in the answer. A run chose a date, found it wrong, chose again and read the new date, but its answer still quoted the first: both readings evidenced the question. A read that evidences a question now drops earlier readings of the same address whose quote the page no longer shows. They stay as context.
- A page read before is read again when what it showed has changed, even if that read answered nothing yet. A task asked for the name a wizard's Review step showed last and the confirmation, as one question. Reading Review answered nothing alone, so after the run went back and corrected the name, Review was not read again, and the answer named the old name. A return to any page state read before now reads it again once a quote from that read is gone. Superseded readings are compared as whole words ("Priya Sharma" is not shown by "Priya Sharman").
- A pager a loader hid no longer ends the pages. A list read page by page looked for the next page in controls observed while the capture could still be waiting out a loader. With no pager in view, the pages ended early. The page is observed again once the capture has waited.
- A head start is not timed before the run begins. The plan and shortcut start before the browser. A call
that finished after
max_secondsof browser startup ended the run over time the limit does not count. - Published eval rows keep their repeat. Arms are paired attempt by attempt; without the repeat, the report paired each arm's earliest attempts, which can differ when an outage costs one arm a repeat.
- A failed history read no longer ends a run. Every observation asks the browser for its history to decide whether to offer BACK. When that failed as the first page opened, the run ended. BACK is now not offered, and taking BACK still reads the history again.
- Dropdown choices see today's date. The action chooser knew today's date and weekday, but the separate choice of a dropdown option did not. A calendar asked for the month after next could select one month too far ahead. Shared page state now gives the option choice the same date as the action and completion checks.
- A rejected page read can recover before its records enter a tally. An invalid record range or omitted requirement on one page left a cross-page count incomplete, while later reads repeated totals without the missing evidence. The reader now retries that saved capture once before merging its records. Successful reads need no extra call; a second invalid read still leaves the count open.
- A page that fails to open is tried four times. Behind a cloud browser's proxy, a first page failed both of fastbrowse's tries, a second apart, four times in three evals of about 300 runs: a certificate error, two connection timeouts and a document that never became ready, each while the site answered other clients. Navigation now tries four times, waiting 1, 2 and 4 seconds between them.
- Days named relative to today land on the right date. The models were told today's date without its weekday, and on a Sunday the field writer booked "the next Monday strictly after today" a week late in three runs of three. Every check now sees the weekday, and field writers look up a relative day in a list of the next seven dates instead of counting; the same task then passed six runs of six.
- A lookup's explicitly unmet requirements reach the verifier. Completion skipped the verifier for answers read on known pages even when the done check named a missing requirement. Only doubts with no unmet requirements can take that shortcut.
- Completion follows page changes during reading and answering. A redirect or document replacement now refreshes the page used for the caller's completion check and final URL. An unchanged page keeps its observation, avoiding another scan of its controls.
- Published eval times are wall time. Earlier releases subtracted the failed requests and backoff that fastbrowse's client measured inside an attempt, which no other arm's time could have taken out. An attempt with any failed request is now an outage and is run again, and 0.5.6's published times, the only ones that subtracted any, are shown in full: its core median rises from 20.7s to 22.7s.
- Model calls go to the fastest endpoint OpenRouter serves. OpenRouter's default routing sent gemini-3.8-flash reads to Vertex at a 2.4s median, where Google AI Studio answered the same read in 1.3s. Requests now ask OpenRouter to sort endpoints by latency, still requiring structured output and keeping its fallbacks.
- Counts and comparisons read matching list records before leaving the page. The existing read assessment treats a paginated list as partial evidence even before its total is known, avoiding detail-page detours before counting. Pagination observes controls and captures records concurrently after navigation settles, then waits for both before following the next link. Each page still supplies its own captured records, and the final page must confirm the list ended.
- Short answers finish with fewer waits and rewrites. Completion and claim checks now run together on the reader's draft. A labelled value or a comparison with its supporting values can stand as written, without a model call to turn it into a sentence or remove the supporting facts. Failed claims still go to the composer, and its answer keeps the same evidence and receipt checks.
- Numeric rankings compare records from every page in code. The reader identifies matching records and the fields holding their labels and values; code ranks the collected values when the list ends. A winner from an earlier page keeps its quote and page address in the answer's citations. Missing records, ambiguous fields, mixed currencies and ties at the cutoff leave the comparison to the reader. Records used for both counts and value comparisons keep their quoted values visible to the answer checks.
- Calendar days keep their weekday column. Controls inside tables now carry their header names, preserving empty cells and row and column spans, and expanding abbreviations from the page's titles. A request for the first Friday can choose the day in that column instead of reconstructing the calendar from flattened text, which picked a Tuesday and spent extra steps correcting it.
- Filtered counts across pages use recorded tallies. The plan identifies requests for one count, and code counts their matching records even when a reader returns plain continuation records or varies the group's label between pages. Overlapping record ranges count each record once, and reaching the list's end closes the tally without asking the reader to calculate a total. A counted range of adjacent list items counts each item, as a range of records already did, rather than counting it as one.
- Placing an order reliably pauses for confirmation. The check before an irreversible click rated a checkout's final button barely above its threshold, so an unauthorized order could go through. The question now names placing an order among the changes that cannot be undone.
- Six held-out eval tasks moved to the dev sets. Changes in this release were profiled or measured on
new-window,countries-mongolia,table-largest-due,quotes-search,stretch-calendar-first-fridayandstretch-quotes-top-authors, so their scores no longer test unseen work. Fresh tasks for the same skills replace them in the held-out sets. - Browser steps spend less time waiting on remote commands. Clicks check for a popup with their target guard, fills prepare focus with the editor check, and reads wait for loading and capture text in one call. Navigation waits for a readiness event, input starts the settling clock before its reply returns, and new tabs activate while their browser domains are enabled. Stale targets, covered controls, dialogs and secret fields keep their checks, and a press still completes before a release is sent.
- A click that opens a new tab carries on in that tab. The run stayed on the opening page, so the next step was spent choosing the new tab, and a read before it could credit the opening page's heading to the new window. The previous tab is still listed and can be switched back to.
- Forms can fill more than one empty field from a single model call. After a confident field choice, the writer supplies values for the same form and each fill checks that the page and other fields stayed unchanged. A new field, changed value, dialog or navigation returns to the usual decision loop. Secrets, populated fields and submission keep their existing checks, and corrections still follow the task's value order. Fills also avoid activating an already focused tab and let the page's input events drive settling.
- Repeated lists need fewer model reads to count. The reader can identify a field shared by every record, and code groups and counts the original captured blocks. For an unfiltered count across all pages, later pages reuse that field only while their title, section and record structure match; a change goes back to the reader. The final page confirms the list ended, and merged tallies keep every record's quote and page address.
- A date picker's controls say which month they show. A table's header cell no longer names what sits outside its row, so a jQuery UI datepicker's Next and days are named by the month in its title rather than by the "Su" weekday header, and every month reads differently. Choices are also told today's date, so a run asked for "next month" stops there instead of going back and forth past it or paging on for months.
- Text inside a same-origin frame is read where the frame stands. It was read after the whole page, so a subscription form's own heading followed the footer and the page's heading above the frame was reported in its place. Readers are also told which blocks are headings and which lie inside an embedded frame.
- A field the task gives no value for is not offered again. Once a field was found to have no value in
the task, the next choice could pick it again and end the run
needs_input, as a blog's unrelated "Enter Name" did to a date-range booking already made. Recovery can still send a run back to it, and a field the task cannot go on without still ends the run there. - Paginated comparisons spend less time reading and navigating. Pager links open their observed same-origin address directly, with each followed page still recorded. Readers keep matching records from the first page, use a smaller response for intermediate pages, and cite earlier evidence with short references that resolve to the original capture spans. Answered comparisons go straight to the final checks.
- A page read before is read again when what it said is gone. A wizard's Review step, read before the run went back to correct a field, still answered for it afterwards, so the answer could name the value from before the correction. On a return to a page state whose quoted evidence the page no longer shows, the requirement is reopened and the page read as it is now.
- Stepping back and forth through a wizard is quicker. A page the run left unread is not read on the way back when nothing on it changed; an empty fill into an empty field goes to recovery instead of being typed; an unsure pick the run already took from the same page recovers instead of retracing it; and a Next that only shows the values a step already held no longer counts as a setting put back, which tripped the stall recovery on the pass after a correction.
- A site that drops the first page load ends the run
unavailable, noterror. Chrome's empty, reset, refused or proxy-failed connection before the first step says the site was down, as a timeout already did; an address that does not resolve still endserror. - A value the task changes later is typed in order. Asked to enter one value and later correct it, the field writer typed the correction on the first pass, so no step could show the change. It now lists the values a field holds in turn, and the run types the next one in that order, so a field set back to an earlier value is not skipped past it. A box ticked or option chosen early in a long form stays in the record the completion checks read, as typed values already did.
- An answer read on a page the run guessed the address of is always checked when doubted. A doubted answer
read on a built address, such as
github.com/owner/repo, goes to the verifier, and the verifier finding it read off the wrong page sends the run back to work. Only answers read on pages the run clicked to skip it. - Paginated lists are read while the next page loads. After the reader identifies a list needed in full, intermediate pages collect records concurrently. Their quotes merge in page order before the final page answers the comparison, within the run's page, step, call and time limits.
- Cloud clicks, fills and observations use fewer round trips. The pre-press frame wait shares its page evaluation with the target guard and hit test, field hand-off also reads the input type, and page snapshots run beside the history read. Pointer, focus and secret checks still gate input.
- Evals never count a provider's outage against an agent. Browser Use is now scored on its own agent's
answer (its
done, or its final reply when it never callsdone), not on Browser Use's lateris_task_successfulverdict, which failed six correct 0.5.7 answers. It is timed to that answer, not to the session reporting stopped, which came up to two minutes later. Outage waits fastbrowse measures inside a run are left out of every published time. A fastbrowse attempt in which any Jev call took over 2 seconds (healthy calls take about half a second at any page size) is a Jev outage, run again rather than scored. The 0.5.7 results are republished under these rules. - Fewer model hand-offs on the way to an answer. A lookup whose every requirement cites evidence finishes
without the screenshot verifier, which never found an ungrounded claim in 76 0.5.7 verifications. Only an
answer read on a search the run built itself (an address with a query) is still verified. A verifier naming a
requirement as
req_1or1now namesreq-1rather than nothing. A read that answers a lookup finishes without another decision or another look at the page it read, and a page with text but no controls is no longer waited on for 12 seconds. - A run reaches its first page sooner.
run_taskasks for the plan and the shortcut while the browser starts, so the tab opens straight onto the shortcut, and opening the tab sends its independent browser commands at once. - A comparison is read without restating every record. The reader names the records a winner or count rests on as block ranges, and code copies their quotes in, where it wrote a claim for each: a priciest-book read took 4.8 seconds rather than 7.9.
- A page is read once it has finished loading. A capture waits up to 3 seconds for a visible loading indicator to go, where it read "Loading..." with the LLM and then read the page again.
0.5.7 - 2026-09-25
- Every eval comparison sets its arms on the same attempts, and an outage scores no arm. Each suite is split
into comparisons of the arms that ran the same tasks. At each task every arm keeps as many attempts as the arm
with fewest measured, all timed in wall time. A missing hosted verdict, a site answering 5xx and an error sent
with HTTP 200 now count as outages, rerun up to five times. The gateway's $0 Jev metering is priced at list for
both Jev-backed arms. The feed moves to
schema_version2. (#153, #154) - A dense page's step is asked again smaller when Jev sheds it. The gateway answers large Jev requests with
503s far more often than small ones, and a step on a Wikipedia article (160 controls, about 27k tokens) ran out
of retries in six runs of six, ending each one
unavailable. When that happens to a large step request, it is now asked again with the on-screen controls only, then with half of those left; a small request that is still refused ends the rununavailableas before. The error for a request that split into single questions now counts every request the call sent, not only the last question's. - A secret that is also part of a site's hostname no longer breaks the addresses a run cites. 0.5.6 kept such
a host readable in reported addresses, but the model's own view of the page was blanked first, so steps, facts
and citations after signing in as
practicestill linked tohttps://••••••••.expandtesting.com/secure. Masking and redaction now share one rule: a value that sits wholly inside the host of an address on an origin it was typed on is left, wherever the address appears, and every other appearance of it, a password's included, is still blanked or redacted. - An answer states the value, not the requirement ahead of it. A value read straight off the page was answered as the requirement it met followed by the value ("Find the latest released version of httpx." then "httpx 0.28.1"); the answer is now the value alone.
- The page a run began on counts as visited. A task saying "Start at" an address could be planned as a requirement to go there, and a run that opened a deeper page from it was sent back because nothing it was checked against showed it had been there. The done check and verifier now see every address the run has been on, first the one it began on, and the planner no longer makes a start address a requirement of its own.
0.5.6 - 2026-09-25
- Correcting a wizard step is not a loop. Going Back through a multi-step form to fix an earlier step reached pages the run had seen before, and it was stopped as stuck. A return now counts as a loop only when it repeats a move the run already made. (#132)
- What a form holds can be quoted. A filled field, date or select now reads as
Label: valuein the page text, so an answer about the dates or options chosen can cite them. Password and secret values are never read. (#136) - Finishing is judged on what was typed, in order. The completion checks now see today's date and each action's typed value and effect, so "enter one name, then go back and correct it" or "the next Monday" can be confirmed or refused. A field asked to change later now gets its first value first. A control is named by the nearest title before it, so a date picker's Submit is no longer named after the post around it. (#143, #144)
- Live evals require the evidence they grade. A missing final page fails any arm whose page the harness can observe (the hosted agent, which cannot report one, is graded on its answer), unfinished hosted sessions no longer pass, subprocess arms get only allowed environment values, and all arms share one prompt. An opt-in browser-use OSS adapter and pinned external task loaders support preparing independent comparisons.
- Eval results have a site feed. Regenerating docs also writes
docs/results/summary.jsonfrom published rows, with release and suite versions, per-arm time and cost, and changed task versions. Just recipes run, publish and regenerate without editing totals.
0.5.5 - 2026-09-25
- A Jev outage costs a batch, not the run. After two 5xx answers in a row, a batch of questions is asked one at a time within the retries it had left, and the answers that came back are kept. (#131)
- Back stays on the task's site. Back is offered only when the page before is on the same site, so a wizard that shares one address no longer steps back to a blank page. (#132)
- Counts over long lists fit in the run's notes. Records are tallied compactly with their citations, code
computes the totals and rankings, and a run that still fills its notes ends
observation_limitwith what it had grounded rather than an error. (#133) - Star ratings drawn as icons are read. A rating shown only as a class such as
star-rating Three, or an accessible label, now appears in the page text. (#134) - Date pickers and date fields can be used. Links that act as buttons, such as a calendar's days, are offered as controls, and native date, time, month and week fields are filled in ISO form and read back. (#135, #136)
- A browser that never loads the first page reports
unavailable. Such a timeout is retried once; a timeout on the agent's own later navigation remains anerror. (#137) - Evals record what they measured. Every task has a version that a test forces up when its grader changes,
every result row records the build and task version it ran, published results are committed rows that are
never rewritten, and the tables in the docs are generated from them.
--onlysearches every suite and rejects a task it cannot find. (#139) - A comparison missing one of its records no longer names a winner. When a page of a list named a record the reader could not tie to the page's text, or more records than one page holds, the record was dropped and a later page could still settle "the cheapest" without it. That requirement now stays open for the rest of the run, so the run reports it could not settle the comparison instead of answering from part of the list. A page that states its own order still settles it on the leading record. (#129)
- An answer about what the run bought is checked against the page it bought on. An Amazon run bought one pen and reported another, citing the search listing; the checkout page it had read named the right one. In a run authorized to commit, Jev now judges which of its clicks, Enters and accepted dialogs committed something, once the run answers and only for those whose pages it read; the answer is checked against those pages, and the composer is told which notes come from them. Other runs make no extra call. (#117)
- A click that found its element redrawn no longer spends a step. A date picker that redraws under a click
dispatches nothing, but the stale step counted toward
max_steps, and Google Flights runs spent two to four of them. It still counts toward the stall budget, which bounds a page that never stops redrawing. (#101, #94) - Fewer wasted steps on forms and lists. The target question now sets a form's mode (a trip type) before its submit, and knows that a broad "Explore" control is not the search asked for; a calendar's fare per day is an input, not evidence to stop and read; a plan no longer turns "report the total" into "the total once the order is finished"; and a shortcut for a count over records goes to the page listing them, not a page about one of them. (#99, #101)
0.5.4 - 2026-09-25
-
A cheapest or highest from part of a list needs the page to state its order. A reader that cited the page's sort order to settle a superlative on the leading row still settled it when that citation named no block on the page; the row alone no longer answers it.
-
A step with one possible target no longer fails. Jev now refuses a choice of only one option, and the typesafe-ai route reports that refusal as a 503, so a page with one field to fill ended the run, or in the evals retried it for hours as an outage:
wiki-godelnever finished. A choice of one option is answered without asking, and a request left with no open question is not sent. -
A run that has not acted cannot be talked past an action Jev holds undone. Asked to open a project page, a run chose DONE on the start page, Jev's done check held the requirement unmet, and the verifier called it complete, so the run reported
completeon the wrong page. With no action taken, an action requirement Jev holds unmet now keeps the run going. Opening a shortcut address counts as acting, so a run the shortcut already took to the page is not held back. -
The live evals send
GITHUB_TOKENto the GitHub API when it is set. A row's answer key is fetched again on every retry, so a long provider outage spent the anonymous 60 requests an hour and failedgithub-licenseon a 403 that said nothing about the agent. -
A secret that is also part of a site's hostname no longer breaks the reported address. A username of
practicesigned in athttps://practice.expandtesting.com/secure, and redaction rewrote the host as well, so the run reportedhttps://[secret:username].expandtesting.com/secure, which is not an address. Final addresses, trace addresses, fact and citation links now keep their host and port on an origin the secret was typed on, since a value there was published by that site; on any other host, and in the path, query, fragment or sign-in part of an address, it is still redacted, as is every other appearance of it in the run's text. -
Evidence from the wrong page no longer counts as an answer. A proposed address opened a flights summary rather than the search the task described, the reader quoted a price from it, and the requirement counted as evidenced, so the verifier could not hold it open however plainly it was the wrong page. The verifier is now told where each requirement's facts were read and which addresses the run built from the task rather than reached by clicking, and it can name a requirement whose evidence came from the wrong page. A requirement evidenced on such an address goes to the verifier even when Jev's own check accepts. When that evidence was read on a guessed address, the requirement is not excused by having it: the requirement reopens, and the refusal names the page it was read off, so the run goes to find the right one rather than finishing again from the same notes.
-
A reader can settle a superlative the site has already ordered, or name the control that shows the rest. Three reads of a filtered results page returned nothing while the cheapest row was on screen: the reader saw the list go on and never assigned the requirement, so the done check refused and the run stuck. A page that states it is ordered or filtered by the quantity being compared now settles the superlative on its leading record, citing that statement so the claim rests on it. A count or total over a list that goes on is not settled this way. Where the list really does go on, the reader names the control that shows the rest, and the run opens it when the page offers that label.
-
A page that cannot settle a list now has to say what it compared. The reader's prompt asked a continuing page to quote every record it compared, and about half the time it quoted only the leading one, so the winner on a later page could not show the values it beat and the claim check scored it unsupported. The records are now a required part of the reader's answer rather than a request in prose, and code copies each quote from the blocks named, so a record is the page's own text. A read that settles its list pays nothing for this.
-
A run reads what its last interaction changed before it calls itself finished. A run clicked a filter and declared itself done against the results as they were before the filter applied, so the check read a list the run never saw. A run that owes an answer now reads the page its last interaction drew before the done check judges it, asking the reader again for what it had already found. A run that only acts finishes without reading or waiting.
-
A filter put back to a state its page already held is not progress. On a results page, turning a filter on and off redraws the rows underneath it, so every click reached a page state the run had never seen and nothing counted it: runs toggled one control until the step limit. A setting returned to committed values its document has already held no longer counts as progress, however the results redraw, so three of them reach the no-progress check and the run recovers. A setting given a value its page has not held is untouched.
-
A live eval survives a grader that raises. The agent chooses where a run ends, so a grader is handed any address a page can navigate to, and one that could not be parsed took down the whole suite: 59 of 63 runs were discarded after four had finished. A grader that raises now fails that row and nobody else's, and so does an answer key that cannot be fetched for a reason a retry would not cure.
-
A page that rewrites its own text cannot be read for ever. Reads were remembered by the page's exact content, so a ticker, a rotating advert or a live counter minted a key the run had never seen on every observation, and the agent could read one page until its step budget ran out instead of acting. Reads are now also budgeted by the page's address and what it lets you do rather than by its text, and two reads of one page state that add nothing the notes did not already hold make the run act instead. A read that adds a new fact restores the budget. Each page of a list paged in place keeps a budget of its own, and the same records read again off a ticking page are not new facts.
-
A Cloudflare edge error is retried like a 503. A provider behind Cloudflare answered a 503, which was retried, and then a 520 seconds later in the same outage, which ended the run in error and failed its eval row for good, although a re-run of the task passed. Statuses 520 to 524 are now retried with backoff, and once the retries run out the run ends
unavailable, so the live eval retries the row rather than scoring it. -
A value is still the page's when the model retypes its punctuation. A field the page writes with a curly apostrophe, an en dash or an ellipsis was dropped when the model quoted it with a straight apostrophe, a hyphen or three dots, so the fact never landed, the requirement stayed open, and the run read the same page until it stalled. A value now matches across a punctuation family, and across the backslash a capture puts before a table cell's own pipe. The value returned and the quote kept as evidence are both the page's own text, not the model's retyping, and the words, their order and their spacing all still have to be there.
-
A table filtered on the page is read as filtered. Rows a filter hid, with
hidden,display:noneorvisibility, and rows in a hidden header, body or footer still entered the capture, so a reader could answer from a row the page no longer showed. They are left out now, and so are cells a column toggle hid.
0.5.3 - 2026-09-22
- Eval times leave out provider outages. A run's
secondsin the live and local evals no longer counts time spent retrying a provider's 503s and dropped requests, which says nothing about the agent; the row'stransient_secondsholds what was left out. - Relevance passes send small requests in parallel. Packing a page's Noul questions into as few requests as
Jev's 64k limit allows made each one about 26k tokens, and the gateway answered most of those with 503 until
the retries ran out, adding up to 25s a pass or ending the run
unavailable. Questions now go in requests of about 8k tokens, sent together. --proxy-country CCpicks where the cloud browser browses from. It always browsed from the US, so amazon.co.uk opened on "Deliver to United States" and a dispatch-country prompt.--proxy-country ukshows the UK storefront as a UK visitor sees it. A country on a local or attached browser is refused, since that browses from the machine's own IP. The Python API'sproxy_countrygains its command-line surface.- A link's name no longer includes a nested stylesheet. Amazon puts a
<style>block inside a result's link, and the element's name was built from every child's text, so Jev was offered a control named by a page of CSS and tried to click it. Style, script, noscript and template children no longer contribute to a name. - A dense results page is compacted rather than refused. Amazon's signed-in search results offered Jev 112
products with ~200-character titles and ~480-character tracking links, and the run stopped at
observation_limitbefore it could pick one. When a page does not fit even on-screen only, its elements are now sent with labels and links shortened before fastbrowse gives up. - The CLI attaches to a browser already running.
--cdp-url ws://…drives any browser exposing CDP: a container, a VM, a hosted browser. The run opens one tab and closes only the tabs it owns, so the browser is left as it was found. The Python API'scdp_urlgains its command-line surface; flags that shape a browser fastbrowse starts (--local,--headed,--profile,--cloud-profile, and their FASTBROWSE_* environment counterparts) are refused alongside it. --bitwardensigns in past an authenticator-app code. A vault item that holds an authenticator key now also offersone_time_code, the current code computed from the key at the moment it is typed, scoped to the same origin as the username and password. Amazon's two-step sign-in had stopped a run at "Enter OTP" with the key in the vault. Base32 keys andotpauth://totpURIs are read locally, so Bitwarden Premium is not needed; a code with a few seconds left waits for the next one, and a key fastbrowse cannot use (steam://, HOTP) is refused when the item is read.- A stored username is typed into an email or phone sign-in field. Amazon's sign-in field is labelled
"Enter mobile number or email", and Jev, seeing only a secret named
username, chose to write new text, so the run stoppedneeds_inputbefore signing in. The field question now says that stored secrets are the site's sign-in credentials, named by role. - A dense page keeps the controls the task needs, not the first ones in the document. The browser now
indexes up to 320 controls, and when a page offers more than 160, Jev is asked in one batched pass whether each
could serve the task's next actions; the step sees the most relevant 160. Before, a page was cut at its first
160 controls, so a result or filter drawn after a long header and sidebar never reached the choice. Pagers,
blocking fields and controls holding a value or selection are always kept, and a control Jev did not answer
for is kept rather than dropped. A dense page costs one more Jev round trip; a page under the limit costs
nothing more.
A page whose on-screen controls alone outgrow Jev's input no longer ends the run at
observation_limit: the controls that fit are offered, pagers and set fields first, and the rest are reported as omitted so a scroll can reach them. The run stops there only when the page's state does not fit with no controls at all. - Jev reads short facts on long pages too. Jev's quick read of a fact, such as a version or a date, gave up on any page with more than 253 quotable spans or more text than its input allows, which was six of ten real pages measured, from Wikipedia articles to GitHub releases, and the LLM reader then read the page a chunk at a time. Jev now first judges which passages bear on the requirements, one batched pass, and picks the fact from those. A requirement Jev calls absent from a narrowed page still goes to the LLM reader, since the evidence may sit in a passage it set aside. A field Jev picks from a listing record, such as a release date, now quotes the record up to it, so the answer's claim that version 0.1.0 shipped that day still has the version in its citation. A batched Jev pass cut short by the time limit now keeps the cost of the requests that had already answered, and one the call limit cannot cover counts none of its calls.
- A setting switched back to where it was is caught as a loop. A run on a results page could turn a filter on and off until it ran out of steps: each click changed the page, so no two steps looked alike. When every control on a page returns to the values it held before an earlier action, the run now stops and asks the recovery model for a different approach, naming the control and the actions in between.
- Recovery remembers its earlier attempts. The recovery model and Jev now see the last few diagnoses and subgoals of the current stall, and which attempt this is, so a second recovery does not propose the plan the first one already tried.
- A form's Next button no longer ends the run. A click on a "Next" that was not a link to another address was treated as paging to the next part of a list, and the run stopped on a multi-step form's first step. Only a link to another page counts as a pager now.
stretch-devandstretch-heldouteval suites.devandheldoutnow pass almost every run, so these harder tasks (a multi-step form with a correction, a date relative to today, a list aggregated across pages, a filter applied and partly undone) are what show whether an agent change helps.- Every step is credited to Jev or the LLM, never to "code". A step code dispatches carries out a model's
choice, and is now recorded as that model's: the next page of a list is the reader's, since the reader asked for
the rest of the list, and the read taken before an interaction is Jev's, since Jev judged the page to be
evidence. A refused finish is the verifier's when it ran, and Jev's otherwise.
decided_byno longer takes the valuecode, and recording captions name only Jev or the LLM. fastbrowse --versionprints the installed version, which bug reports now ask for.- A field the task has no value for is skipped before it ends the run. Most such fields are optional: Google
Flights opens a "Where else?" box beside the origin, and a run that picked it stopped
needs_inputwith nothing searched. The first time, recovery is told the task gives no value and chooses another step; a field it sends the run back to is a required one, and the run still stopsneeds_inputrather than invent a value. - A finished search is checked on the stronger model. When Jev doubts a run is done, the verifier that has
the last word now runs on
google/gemini-3.8-flashrather than flash-lite. On a Google Flights search with every field right and no nonstop filter, flash-lite took the "Nonstop" rows for the filter and passed it ten times in ten; the default model refused all ten and passed the filtered search every time. A doubted finish costs about two seconds more. - The evals allow 50 steps. A fare search on Google Flights took up to 29 steps, one short of the old limit, so a run that needed a recovery or two could run out of steps rather than fail on its merits.
0.5.2 - 2026-09-22
-
A page is read before it is scrolled. A read takes in the whole page, so scrolling one nobody has read only spends steps: a lookup for Mongolia on a long country list scrolled until recovery ran out and stopped stuck with the answer on the page. Once a page is read, scrolling it goes ahead, as a list that draws more on scroll needs.
-
A reply the provider cut short is an outage, not a truncation. JSON that ends mid-value short of the output cap is asked for again at the same cap; a second one ends the run
unavailable, naming how many tokens came back. Only a reply that used the cap gets more room and, if cut again, is reported as truncated. -
A browser that cannot be reached yet is an outage, not a failed task. A dropped or refused connection when the run first connects, as a cloud browser still starting can give, ends the run
unavailable, so the eval runs it again; any other connection failure is still an error. -
An eval waits out a rate-limited answer key. Fetching a task's expected answer from a public API now retries a 429, 5xx or dropped connection with the same backoff as a run, where it used to end the whole eval.
0.5.1 - 2026-09-21
- A decision the page redraws under is dropped at once. While Jev decides and gates a click, the run watches the controls on offer; if one changes, as Google Flights' date picker does when its prices arrive late, it decides again on the settled page instead of clicking into a stale refusal.
- A guessed shortcut the site does not serve is skipped. A direct address that answers 4xx/5xx sends the run back to its start page instead of reading an error page and spending a BACK to leave it.
- Runs use a Browser Use Cloud browser by default (
BROWSER_USE_API_KEY).--localruns local Chrome;--headed,--profile,FASTBROWSE_HEADEDandFASTBROWSE_PROFILEimply it.--cloudis removed.--cloud-profilewith local Chrome is refused at startup. - A failure says what happened. Each retry logs the call, the provider's status and error, and the backoff; a call that gives up names how many requests failed, over how long, and the last cause. Evals lead every failed run with how it ended, and print which providers (and whether a backup) they use.
unavailableis its own status. A run that ends because a model or browser provider stayed unavailable through every retry reportsunavailablerather thanerror: nothing about the task failed. The live evals run such a run again, so a published result is the agent's own.- An error never quotes a key. Error text built from a provider's response omits the values validation rejected and scrubs the key wherever it appears, including a model's own JSON keys.
- A browser that stops answering is an outage; a slow one is not. A CDP command with no reply after 60s asks
the browser whether it is still there, and keeps waiting if it answers. Only a browser that does not ends the
run
unavailable, where before a lost cloud browser could leave it waiting forever. - Eval arms are named for their agent:
--arms fastbrowse jev-ultrafast browser-use. - Recordings are captioned and kept plain beside them (
<name>.plain.mp4). Captions are timed from the video's first frame; an ffmpeg without libass still writes both videos, uncaptioned, and a failed write leaves an earlier recording at the same path intact. - Clicks are safer on covered controls. A covered control is clicked at an exposed edge only when no other control (a row's own Delete button) sits under that point; a hover-only region inside a button does not count as one. A select the page reverts is a failed step.
- A recording ends on a readable answer. The closing card shows each citation as a number beside the claim and lists the cited site and quote underneath, instead of the raw text-fragment link.
- A run reports where its recording went.
RunResult.recordingslists the videos written; the CLI printsrecorded: <paths>ornot recorded, and prints warnings to stderr. - Local Chrome has no password popups. Every launch switches off Chrome's password manager and leak
detection, in a kept
--profiletoo; a kept profile whose preferences cannot be read is aBrowserError. - Finishing is judged more carefully. A lookup that has its answer ends rather than clicking on; a search the task only asks to run is a requirement to act on; the verifier confirms requirements one by one, sees the controls that are set and the ones the done check doubted, and a narrowing requirement needs its filter applied rather than matching rows.
- Pages are read when they have settled. Navigation waits for the loaded page's DOM to go quiet; a
transparent checkbox styled by its ancestors is offered; an
aria-disabled="false"added during hydration is not a new control.
0.5.0 - 2026-09-21
- Measured on 2026-09-21: 42/42 answer-task runs passed at a $0.0042 median cost, against 41/42 and
$0.0057 for 0.4.1; hosted Browser Use scored 40/42 at $0.4163 the same day. Navigation tasks passed 18/18.
Google Flights is slower (109.3s against 71.7s). See
docs/evals.md. RunResult.would_firelists the shadow tripwires that crossed, one entry per crossing, so a caller or eval can count them without reading logs.- The reader cites page blocks instead of retyping quotes. A claim names the capture's source blocks
and fastbrowse copies the quote from them, so a fact is no longer lost when the model's copy differs from
the page (a table cell's
|, a record quoted as two lines). A count or winner the page never states cites no text: it is kept as a derived fact, and the claim check judges it from the records it counts.StepFact.quote,urlanddeep_linkareNonefor such a fact. - A card or table row is one block. A repeated card (a quote with its author and tags, a product) and each table row with its header are captured as single blocks, so a record is read and cited whole.
- Model responses use strict structured output, and every prompt was audited. Schemas are sent strict, with field docs included. Page text is marked untrusted in the same words everywhere. Context comes before the question. Rules added for one incident became general rules or were removed. Text fields no longer ask the model to retype a quote: the value must appear in the block it names.
python -m fastbrowse.evals.probemeasures the reader and claim check on live pages. It loads the given pages once and runs the agent's own read, draft and claim check over them N times at once.- Counts, totals and superlatives cite their underlying records. The reader preserves every compared
record across pages and records which facts a conclusion draws on. Drafted and composed answers carry
those records through claim checks, citation links and
RunResult.citations. Required evidence keeps its full basis within the notes budget or stops atobservation_limit. - Long lists are read to the end before they are answered. A single block longer than the reader's
input, such as a flight results list, is split at line breaks rather than stopping the run at
observation_limit, and the last chunk of a page decides whether its list goes on. When the list does go on, recovery and the next-step hint say to load the rest or narrow it with the page's own filter or sort, instead of finishing early with the best record seen so far. A load-more button at the foot of a long list ("View more flights", "Show 20 more") stays on offer past the off-screen control limit, as a pager link already did. - The completion check keeps the page state that matters on long pages. When the page and the notes do not both fit, the check drops page text and controls without state before it drops the evidence, so a checked filter such as "Nonstop only" still counts.
- Recovery can direct a read or a finish. A recovery subgoal that names READ or DONE is followed; a finish is still judged by the completion check.
- Page-script output is validated. What the browser scripts return is parsed into typed models, so a
mismatch raises
BrowserErrorat once instead of failing later in the run. - Concurrent live evals keep separate traces. Each run records only its own events, including events from its child tasks. Finishing one run no longer disables trace collection for the others.
- Answer claims link to the words that support them.
RunResult.citationsexposes each cited fact, its requirement, source URL, verbatim quote and text-fragment deep link. An answer citing an unknown reference fails the claim check and falls back to one drafted from verified facts.RunResult.answercontains numbered Markdown links; integrations should render them as Markdown. MCP answers carry the links too, while its citation records retain theirquoteandurlshape. - See what each step learned and why it stopped.
StepResult.factscarries that step's added facts on itsStepEvent, with quotes, source links and the reader (jev_choiceorllm).StepResult.notereports read outcomes, dispatch details, refusal reasons and recovery guidance when available. Resolved secrets are redacted before delivery. - Watch the active tab as the agent works.
run_task(on_frame=...)sends JPEG bytes, follows tab switches and works with browsers reached over CDP. Delivery paces capture by acknowledging each frame after the handler returns, with no fixed frame rate; a later pending frame replaces an earlier one. Slow or failing handlers do not hold up the run. Capture is off unless a handler is supplied. Live images and recordings are held back while a resolved secret may show on the page, as PNG step frames are. - Links pass the same authorization gate as buttons. Jev judges whether a click, an Enter press or dialog acceptance commits an irreversible change. Code-selected pagination is exempt, and authorized actions with sufficient confidence go straight through. Refusals appear as failed steps with reasons before confirmation or recovery.
- Read evidence before a click can hide it. Jev judges whether the page holds information the task needs. Reading preserves that evidence before another interaction, skips unchanged content for the same open requirements, and sends comparisons and partial evidence to the LLM reader. Short facts can be copied from quoted spans by Jev; each fact records which reader supplied it.
- Moving controls are rechecked before input. A click waits for its target to stop moving and checks its meaning and position again. A replaced field must still match and hold focus before it receives text. The browser also gives visible loading indicators time to clear before declaring a page settled.
- Repeated work does not count as progress. Filling or selecting a value already present in the
observed field cannot reset the stall count. Repeated actions and unresolved requirements are also
monitored: by default they log possible stalls without changing the run.
StallRules.tripwirescan enable recovery for them. Productive steps break the unresolved-requirement streak, and recovery resets the evidence used by all three stall checks. Local and live evals record these signals per run; the live summary counts passing runs affected, so repeated signals in one run do not inflate the rate. - Prompt limits preserve required evidence.
ObservationLimitsnames the page-text, working-notes and history limits;TokenBudgetnames input and reader/composer output limits. Shortened page excerpts carry a cut marker when it fits. Verdicts retain requirement evidence or stop atobservation_limit; the Jev completion check makes room by reducing page text before refusing the evidence. - A Jev provider outage can use the other configured provider. With both keys set, a retryable HTTP failure that exhausts retries switches the run to the backup provider and keeps it there. Authentication errors do not switch providers. A custom endpoint or model disables automatic failover.
- CLI secrets can declare their own origin.
--secret NAME=ENV_VAR@ORIGINworks without--start, including wildcard site origins. The shorter form still takes its scope from--start; Bitwarden still needs a start page to match the vault item. - The MCP HTTP token can be read from
.env.FASTBROWSE_MCP_TOKENfollows the same settings rules as model keys, with the process environment taking precedence. - Flights grades use the submitted search even when its fields collapse. The grader decodes the route, date and trip type from the URL, falls back to visible fields when decoding is unavailable, and still requires matching result rows and any requested Nonstop filter. A decoded mismatch fails.
0.4.2 - 2026-09-20
- A form is set up in the order that works. Its mode - which tab of a search, which kind of account or ticket, which category - decides which fields it has and empties what they hold, so it is chosen before any value is typed rather than after, which used to mean typing the values twice. The filters a task asks for are set before submitting where the form offers them, because setting one afterwards submits twice, and a filter the page only reveals once there are results is set there. On a flight search this removed five steps of rework; on a two-package comparison it removed the repeated writes to the search box that had made it the most expensive lookup in the suite.
0.4.1 - 2026-09-20
- A secret can be declared for a site rather than for one of its hosts.
https://*.example.comcoverswww.example.com,accounts.example.comandexample.comitself, which is how one sign-in spans a site: the login typed on the account host is the login the shop host asks for. The wildcard stands for whole labels only, so it does not coverexample.com.evil.test, and neither the scheme nor the port is ever wildcarded. An exact origin behaves exactly as before. This reaches the MCP server too (--secret NAME=ENV_VAR@https://*.example.com), where each secret keeps the scope it was declared with rather than the start page's. ScopedSecrets.per_secret({name: (value, origins)})holds a person's credentials each scoped to the sites it belongs to, for an application that stores them that way. The single-origin constructor is unchanged.- An IPv6 origin survives being read back.
origin_ofreturnedhttps://::1, which is not a URL any parser reads again, so a check against an IPv6 origin could raise rather than answer. The literal keeps its brackets.
0.4.0 - 2026-09-20
Everything an application needs to run fastbrowse as its browser engine rather than as a command someone types. Each of these came from wiring it into a product that already had one.
- Drive a browser you already have.
cdp_urlattaches to any browser over the DevTools protocol, wherever it runs: a container, a VM, a machine you own. The run opens one tab and closes that tab, so a browser handed over is left exactly as it was found, and nothing is billed to a cloud account. This is the option to reach for when the browser should live next to the user rather than in someone else's cloud. - Start from the task alone.
startis optional now. A caller whose own interface takes a goal and no URL had nowhere to get one; the first address is worked out from the task, as a person would.--startis optional in the CLI andstartis optional on the MCP server'sbrowsetool, wheretaskis now the only thing a call must carry, and it holds for an attached browser too, which the run opens its own tab on. A secret is only ever typed on the start origin, so asking for one without a start page is refused rather than quietly dropped. - Stop a cloud browser you did not start. The browser event carries the cloud browser's id, so an application that has to end a run out of band (a user pressing cancel, a subscription ending) can.
proxy_countryandviewportreach a cloud browser the run starts, instead of being fixed at what the library guessed.
Fixed in the same release, from tasks that failed in the field:
- A list longer than one page is answered from the whole of it. A task over a paginated catalogue read the first page, answered from it and called that done. The run now follows the pager until what was asked for is evidenced or the page cap is reached, the reader is told when the page it is reading continues, and a claim about a whole list is not accepted from one page of it.
- A bot check is reported as one.
Status.BLOCKEDis new: a CAPTCHA is not a sign-in and no credential passes it, so a run that meets one says so rather than ending asstuck. A challenge that clears itself once its script runs is still waited out first, and the check is made whether or not a secret is held for the site. - A reply cut short is asked for again. A read whose answer hit the output limit was parsed as though it were whole, so facts after the cut were lost without a word.
--jsonkeeps its contract on a bad limit.--max-steps 0printed a traceback and nothing parseable; it is now refused like any other bad flag, with the error on stdout as JSON.- A limit reads as what it is in the message that reports it: a dollar limit as money, a duration as a duration.
0.3.4 - 2026-09-20
- A step frame can no longer carry a secret the step itself revealed.
Config(step_frames=True)checked whether a resolved secret was on screen using the reading of the page the step was decided from, which is the page before the action ran. A fill that a page mirrors into ordinary text put the secret on the page after that check, so the frame sent to the caller could contain it as pixels. The check now reads the page as it is when the image is taken. Affects 0.3.2 and 0.3.3 with step frames enabled; no other surface sent an image.
0.3.3 - 2026-09-20
- Runs on Python 3.13. The floor was 3.14, which an application pinned below that could not work around:
uv add fastbrowsesimply would not resolve. Nothing in the package needed 3.14. CI now runs the whole gate on 3.13 and 3.14, so the floor is exercised rather than claimed.
0.3.2 - 2026-09-20
- A picture of each step, for an interface that shows a run as it happens.
Config(step_frames=True)puts a PNG of the page a step acted on onto every step event. It is off by default, because it costs a screenshot round trip per step. A step whose page is showing a resolved secret sends no frame: pixels cannot be masked the way text is.
0.3.1 - 2026-09-20
- A cloud browser can run as a profile someone already signed in.
--cloud-profile ID,run_task(cloud_profile=...)and the MCP server's--cloud-profilestart a Browser Use Cloud browser from one of that account's profiles, so a run acts as whoever set the profile up. No credential is shown to a model, and a local run's--profile DIRkeeps working as before. Passing a cloud profile id to local Chrome is refused rather than ignored. - Install from PyPI. The README opened with
git clone, which was the only way to run fastbrowse before it was published and is now the contributor path. It opens withuvx fastbrowse.
0.3.0 - 2026-09-20
- A list split across pages is read whole. Counting or ranking over a paginated list had no path to the answer: a run either stopped at the step limit or finished early on page one. The reader can now say a list continues past the page it read, a claim from part of a list cannot close the question, and when a page has a single next-page link the agent follows it and reads what it opened without asking the choice model each time.
- A click that changed nothing is not repeated. An action that left the page as it was goes to recovery instead of being taken again from the same page, and fields a form will not submit without (required and empty, or marked invalid) are named in what the action reports.
- An MCP server.
fastbrowse-mcpserves onebrowsetool over MCP, so Claude Code, Claude Desktop, Cursor or any other client can hand it a task. What a calling model may do is fixed by the operator's flags: a call can ask for less, never more. - Runs on Windows. A checkout failed before any test ran: files were read in the locale's code page, a date used a flag only glibc has, and Chrome was never found where its Windows installer puts it.
- Fixes found by the live suite: a run outlasts a brief provider outage, a redrawn control's twin is acted on rather than decided again, an empty page is drawn before it is read, a start page that never loads is tried again, and a finish stands when the verifier doubts only what the notes already cite.
0.2.0 - 2026-09-18
--record FILEsaves an MP4 of the tab, ending on the answer, its time and its cost.- Sign in with a stored login the task never mentions.
--secret NAME=ENV_VARand--bitwarden ITEMtype a credential on the start origin without the value entering a model's context.
0.1.0 - 2026-09-18
- First release: a browser agent that picks its next action from the controls the page actually has, with an LLM to plan and read, and code owning verification, safety and secrets.