tablaze.
Try Tablaze
Developer previewBROWSER MCP / DEVELOPMENT

The web.
At your
agent’s pace.

A compact browser MCP that keeps sessions warm, batches the busywork, and checks what actually happened.

No extra model API key for MCP/MIT licensed

One session. A clear path.MCP → CHROMIUM
•••⌕   localhost / stays01
STAYS /LOCAL DEMO

Somewhere worth staying.

e1DestinationChoose a city
e2Nights1
Free cancellatione3
tab_act4 actions / 1 call
01fill e1 "Lisbon"
02select e2 "3"
03check e3 true
04click e4

Illustrative replay of the included local fixture. Not a live browser connection.

A standard protocol.
Your preferred agent.
CodexClaude CodeCursor{ MCP }
Local stdio configuration

LESS CEREMONY. MORE CONTROL.

One loop.
Nothing hand-wavy.

Your agent stays in charge. Tablaze handles the browser mechanics and returns evidence it can use.

01 / OBSERVE

e1 textbox Destinatione2 combobox Nightse4 button Find a staysnapshot_id: …

Know what’s on the page.

Compact observations expose useful controls with references and a snapshot revision. Ask for changes instead of another full dump.

tab_snapshot

02 / ACT

fillselectclickone ordered batch

Move through the busywork.

Keep the browser running and submit multiple actions together. If a target change is detected or a step fails, stop and report what already ran.

tab_act

03 / VERIFY

✓ URL✓ visible text✓ field value✓ element count

Check the result, explicitly.

A successful click is only a successful click. Check URLs, text, values, and visibility against the intended outcome; add post-checks to an action batch when the expected result is known.

tab_verify / tab_act.post_checks

Batches run in order and stop on error. They are not transactions; completed actions are not rolled back.

SMALL SURFACE. USEFUL PRIMITIVES.

Sixteen tools. Room to build.

The current development branch supports forms, virtual-list search, files, tabs, structured extraction, and PDF export. MCP needs no additional model; the optional agent loop is configured separately.

tab_opentab_snapshottab_findtab_acttab_extracttab_verifytab_capturetab_listtab_closetab_navigatetab_tabstab_downloadstab_dialogtab_statetab_extract_structuredtab_pdf

OPTIONAL / OWNED PERSISTENT PROFILE

Keep durable site state
across restarts.

Use a dedicated private Chrome directory. Reopening requires the expected profile ID, and concurrent ownership is refused. A real-Chrome test verified durable cookies, localStorage and IndexedDB across processes; check the active site account before sensitive writes.

Configuration, tested behavior and limits ↗

REAL MODEL RUNS / 2026-09-22—23

What finished.
What did not.

The five-task comparison and later independent follow-ups are shown separately. Business outcomes and complete agent runs are checked separately. These few visible development tasks do not establish overall superiority.

Round 3: business checks passed 5/5 for both

Complete successful runs: Tablaze 4/5, Browser Use 5/5. The order was recorded, but Tablaze’s old inference bridge then terminated a call before verification and reporting; its total tokens are unknown.

TaskTablazeBrowser Use
Form45.664 s58.220 s
Popup approval62.133 s70.208 s
Iframe form60.227 s57.984 s
CSV download93.255 s80.975 s
Order after interrupted response46.226 s · failed return84.718 s

These are end-to-end times, including startup, cleanup, and Browser Use’s default judge. Browser Use’s agent reported completion earlier on all four mutually successful tasks. A failed order return is not a faster success.

Round 3 report, raw data, and conditions ↗

Separate order follow-up: both completed successfully

After the bridge fix, each engine ran once with fresh state. Each recorded one order with zero duplicates, passing independent checks, agent success, and return within the deadline.

MeasurementTablazeBrowser Use
Agent reported completion59.247 s75.354 s
End-to-end time59.415 s93.706 s
Model calls (including judging)44 (1 judge)
Input / output tokens59,123 / 66470,868 / 1,002

The two clocks start at different boundaries; Browser Use totals already include judging. This single follow-up does not replace round 3’s 4/5, combine success rates, or establish a stable advantage. Monetary cost was not estimated.

Follow-up report and actual terminal-event evidence ↗

New virtual-list task: one complete success each

The page renders only seven rows at once. Both engines found VIRTUAL-130 beyond the initial DOM and made exactly one reservation, with zero duplicates; both agents finished successfully before the deadline.

MeasurementTablazeBrowser Use
End-to-end time86.830 s141.597 s
Model calls (including judging)65 (1 judge)
Input / output tokens93,209 / 94491,929 / 2,746

2026-09-23, gpt-6-astra / ultra, shared 120,000-token budget. Browser Use's default judge is included in time and usage. One visible development task cannot establish overall or stable superiority. Two 50,000-token attempts hit the comparison gateway budget; see the report.

Virtual-list report, measurements, and conditions ↗

Shadow DOM and dynamic menu: both agents passed

In two new matched tasks, both agents made one correct write per task with no duplicates and completed before the deadline. Tablaze later added optional action post-checks; a separate Codex Shadow DOM run finished in two model calls, without a concurrent Browser Use arm.

Initial matched task / whole runTablazeBrowser Use
Shadow DOM form87.539 s ✓69.014 s ✓
Delayed menu90.417 s ✓81.933 s ✓

2026-09-23, gpt-6-astra / ultra, seed 31, one attempt per task and arm; Browser Use's default judge is included. Tablaze took longer on both matched tasks. A separate Tablaze-only post-check attempt took 42.012 s and two model calls; cross-run differences do not prove a stable speedup or overall superiority.

Full report, source states, and raw traces ↗

A delayed target, one action batch

Click an observed menu opener, then wait for one uniquely named new button or menu item. In the final visible run, Codex combined both steps into one Tablaze action call. Both agents passed the independent business judge with one write and no duplicates.

Seed 37 / whole runTablazeBrowser Use
Delayed menu80.715 s ✓77.720 s ✓
Model calls44

2026-09-24, gpt-6-astra / ultra, one visible development task; Browser Use's default judge is included. An earlier role-bearing version led the model to guess menuitem incorrectly, then recover without a duplicate write. The final API matches a unique name across buttons and menu items. Cross-version observations do not prove stable speed or overall superiority.

Full report and all failure/success traces ↗

Cross-origin authorization return: one success each

A local provider popup authorizes the original app tab. Both engines made one authorization and one correct submission with zero duplicates, passing business, Agent, and deadline checks.

MeasurementTablazeBrowser Use
End-to-end time123.719 s119.515 s
Model calls (including judging)86 (1 judge)
Input / output tokens127,014 / 1,245111,839 / 1,517

2026-09-23, gpt-6-astra / ultra, shared 250,000-token budget and 240-second deadline. Browser Use's default judge is included in time and usage. It was slightly faster here; a synthetic popup task does not establish real OAuth capability or an overall ranking.

Authorization-return report, measurements, and conditions ↗

One Codex agent, four browser MCP entries

An independent server checks for exactly one correct write. All four entries passed form and canvas. On the virtual list, three direct-tool entries passed; structured direct tools did not complete.

Task / Codex process timeTablazeHarness MCPBU CLI-MCPBU MCP
Form69.331 s ✓111.232 s ✓93.260 s ✓74.824 s ✓
Visual canvas65.037 s ✓77.263 s ✓71.074 s ✓71.326 s ✓
Virtual list90.791 s ✓175.487 s ✓193.335 s ✓195.185 s, no write

Scroll the table sideways to view all entries.

2026-09-23, gpt-6-astra / ultra, one direct-tool attempt per arm and task. Times cover the Codex process, model, MCP and auto-review, excluding Chrome startup and independent judging. The structured MCP's nested Agent separately passed after a longer tool timeout (318.247 s, six additional model calls); its first nested attempt timed out. Both programmable Browser Use entries passed the virtual list. A direct-tool failure is not a product-wide failure, and these samples do not establish overall superiority.

Four-entry report, full fallback and raw traces ↗ · Initial two-entry record ↗

Authenticated response: three direct entries passed

After a local provider popup returns, the receipt reference exists only in the network JSON response. The independent server requires one authorization and exactly one correct submission. Tablaze's response tool is opt-in; Browser Use's programmable CLI-MCP also passed.

Task / Codex process timeTablazeHarness MCPBU CLI-MCPBU MCP
Read response and submit once149.560 s ✓190.380 s ✓291.881 s ✓300.012 s · deadline, no write

Scroll the table sideways to view all entries.

2026-09-23, gpt-6-astra / ultra, one visible development attempt per direct entry. Timing includes Codex, model, MCP and auto-review, excluding Chrome startup, server judging and cleanup. Two separate structured-MCP nested-Agent attempts also failed the server judge, but used different deadlines and a second Agent, so are not pooled with this table. Browser Use CLI-MCP passed; this cannot establish overall superiority or a stable speed advantage.

Network-response report, failed attempts and raw data ↗ · Opt-in network-tool contract ↗

Page scripts: both entries pass once

The same Codex model handled an authenticated-response task. Tablaze enabled page scripts explicitly; Browser Use used CLI-MCP. The independent server recorded one correct submission and zero duplicate writes per entry.

Direct entryBusiness judgeCodex processInput / output tokens
Tablaze --page-scriptPass, 1 write198.359 s317,115 / 1,557
Browser Use CLI-MCPPass, 1 write239.400 s340,202 / 2,461

2026-09-23, gpt-6-astra / ultra, one visible development attempt per arm, 300 s deadline. Tablaze's Chrome launch fell inside its process time; Browser Use's isolated Chrome started before its timer. Times are not a controlled speed ranking, and one synthetic task does not establish overall superiority.

Page-script report and raw traces ↗ · Page-script authority and limits ↗

Four new matched tasks: both finished 4/4

The same Codex model completed an iframe form, interrupted order, virtual list and visual canvas. Independent server checks passed in all eight runs, with one attempt per task and engine.

Task / whole runTablazeBrowser Use
Iframe form83.799 s ✓52.515 s ✓
Interrupted order65.939 s ✓78.520 s ✓
Virtual list74.301 s ✓118.083 s ✓
Visual canvas54.428 s ✓66.514 s ✓

2026-09-24, gpt-6-astra / ultra, isolated Chrome sessions; Browser Use's default judge is included in whole-run time. Tablaze was slower on the iframe sample. One run per task does not establish stable speed or capability superiority. The later real-session recording and initial iframe wait were not part of these runs.

Full report and eight raw traces ↗ · Real-session recording guide ↗ · Browser Use capability and issue audit ↗ · Viewport and permission settings ↗

Iframe follow-up: both pass, speed gap remains

In three paired same-model tasks, both engines made three correct writes with no duplicates. Tablaze's first observation found the child iframe, but two runs spent extra model rounds on an unnecessary text check of a form value.

SeedTablazeBrowser Use
4352.308 s ✓49.108 s ✓
4479.699 s ✓57.496 s ✓
4585.689 s ✓56.040 s ✓

2026-09-24, gpt-6-astra / ultra, one attempt per side and seed on the same source commit; Browser Use's default judge is included. Three-run median whole time was 79.699 versus 56.040 s, missing the provisional 1.25× efficiency target. This sample is still too small to estimate stable production behavior.

Follow-up report and six raw traces ↗

These are separate visible development tasks with budgets and deadlines stated in their own reports. Business outcome, Agent completion and whole-run return are recorded separately; the samples do not establish overall superiority.

For the current SDK extension, see typed custom tools: input/output validation, trusted application context, and browser bindings. The real-model records above cover their original source versions, not success rates for the new extension or production provider adapters.

NEW PROVIDER / SEPARATE VALIDATION

Run a real task.
With Codex directly.

Two visible tasks through the new production entry point, one attempt each: independent acceptance and complete Agent success 2/2, zero duplicate writes. This is not a new matched Browser Use comparison.

TaskWhole runInput / output tokens
Form64.030 s43,391 / 462
Avoid duplicate order78.943 s59,401 / 710

gpt-6-astra / ultra, seed 23, including startup and cleanup; usage includes Codex CLI context overhead. Anthropic / Ollama have local protocol tests only. Overall superiority remains unproven.

Production provider report and raw data ↗ · Four provider options ↗

HISTORICAL LOCAL SAMPLE / 0.1.0

Speed deserves
a measurement.

Preserved 0.1.0 fixture timings, not current development-branch performance. Real-model comparisons are recorded separately above.

LOCAL CHROMIUM / HOTEL SEARCH FIXTUREPreparing measurement report
Cold browser + first pagemedian wall-clock time
Warm snapshotmedian wall-clock time
Action batchmedian wall-clock time
Verified runsthis fixture, this environment

Five historical fixture runs on Apple M3 Max, 2026-09-22 04:40 UTC, with source hashes retained in the raw JSON. Direct engine calls exclude MCP, models, and internet websites. The command below tests your current checkout; it does not recreate the old version.

$npm run bench

START WITH THE FIRST CALL.

Give your agent
the keys to a tab.

Node.js 20+ and Chromium. Start in an isolated browser, or explicitly connect a Chrome session you choose.

The downloadable 0.1.0 package and source ZIP are the historical eight-tool preview. For sixteen tools, the agent loop, and typed custom tools, build the current GitHub source ↗.

Follow the source build instructions on this page. A public npm release is not available yet.

Build from source; npm publication is pending

git clone https://github.com/SweetDianDian/tablaze.git
cd tablaze
npm ci
npm run build
node dist/cli.js setup
node dist/cli.js doctor

CLEAR EXPECTATIONS.

Before you build.

What do MCP and the agent loop require?

The core browser tools make no model API calls; your MCP client supplies reasoning. The development agent accepts Codex, Anthropic, Ollama, or compatible endpoints. Codex uses its existing CLI login; other providers use their own credentials. Choose the model explicitly. See provider setup.

Can I use my existing Chrome login?

Yes, through an explicit CDP connection. The default launches an isolated browser. Attached mode shares that Chrome profile’s identity, and Tablaze closes only the tabs it owns.

Can I restrict navigation destinations?

The current source supports exact-origin allow/deny rules through CLI/SDK configuration for HTTP(S) document requests, including redirects, frames and popups, in owned isolated browsers. External CDP is excluded. Fetch, images and other non-document traffic remain unrestricted; interception connection loss can release requests. This is not a network firewall. See navigation policy and limits.

What happens if the page changes?

Actions include a snapshot revision and element reference. Detected stale snapshots or changed targets are rejected; re-observe before retrying. Validation and input are not atomic, so a short race remains. These checks are not a security sandbox.

What does this preview not promise?

It does not promise success on every website, overall superiority over Browser Use, or rollback. The current branch supports explicit file workflows and owned popups; closed shadow roots and arbitrary complex sites remain limitations. Checks and input are not atomic; reconcile uncertain submissions before retrying.

BROWSER ACTIONS. LESS WAITING.

Make the next move.

Try Tablaze