Batch the steps
Fill, select, and click in order, with fewer calls between steps.
A compact browser MCP that keeps sessions warm, batches the busywork, and checks what actually happened.
No extra model API key for MCP/MIT licensed
MCP → CHROMIUMtab_act4 actions / 1 callfill e1 "Lisbon"—select e2 "3"—check e3 true—click e4—Illustrative replay of the included local fixture. Not a live browser connection.
{ MCP }REAL BROWSER · THE COMPLETE RECORDING
Watch Tablaze batch a form, verify the result, stop at a changed target, then follow a popup through approval and file download. Real MCP calls, with a trace you can inspect.
Fill, select, and click in order, with fewer calls between steps.
Check field values and page status after the actions finish.
Reject a detected stale target; observe again before continuing.
Follow an owned popup with an explicit policy, approve, and check the download.
A local MCP fixture with authentic responses and screenshots. English male narration and optional captions accompany presentation pauses; this is not a continuous feed from the controlled tab. No model inference.
LESS CEREMONY. MORE CONTROL.
Your agent stays in charge. Tablaze handles the browser mechanics and returns evidence it can use.
01 / OBSERVE
e1 textbox Destinatione2 combobox Nightse4 button Find a staysnapshot_id: …Compact observations expose useful controls with references and a snapshot revision. Ask for changes instead of another full dump.
tab_snapshot02 / ACT
Keep the browser running and submit multiple actions together. If a target change is detected or a step fails, stop and report what already ran.
tab_act03 / VERIFY
A successful click is only a successful click. Check URLs, text, values, and visibility against the intended outcome; add post-checks to an action batch when the expected result is known.
tab_verify / tab_act.post_checks↳Batches run in order and stop on error. They are not transactions; completed actions are not rolled back.
SMALL SURFACE. USEFUL PRIMITIVES.
The current development branch supports forms, virtual-list search, files, tabs, structured extraction, and PDF export. MCP needs no additional model; the optional agent loop is configured separately.
tab_opentab_snapshottab_findtab_acttab_extracttab_verifytab_capturetab_listtab_closetab_navigatetab_tabstab_downloadstab_dialogtab_statetab_extract_structuredtab_pdfOPTIONAL / OWNED PERSISTENT PROFILE
Use a dedicated private Chrome directory. Reopening requires the expected profile ID, and concurrent ownership is refused. A real-Chrome test verified durable cookies, localStorage and IndexedDB across processes; check the active site account before sensitive writes.
REAL MODEL RUNS / 2026-09-22—23
The five-task comparison and later independent follow-ups are shown separately. Business outcomes and complete agent runs are checked separately. These few visible development tasks do not establish overall superiority.
Complete successful runs: Tablaze 4/5, Browser Use 5/5. The order was recorded, but Tablaze’s old inference bridge then terminated a call before verification and reporting; its total tokens are unknown.
| Task | Tablaze | Browser Use |
|---|---|---|
| Form | 45.664 s | 58.220 s |
| Popup approval | 62.133 s | 70.208 s |
| Iframe form | 60.227 s | 57.984 s |
| CSV download | 93.255 s | 80.975 s |
| Order after interrupted response | 46.226 s · failed return | 84.718 s |
These are end-to-end times, including startup, cleanup, and Browser Use’s default judge. Browser Use’s agent reported completion earlier on all four mutually successful tasks. A failed order return is not a faster success.
Round 3 report, raw data, and conditions ↗After the bridge fix, each engine ran once with fresh state. Each recorded one order with zero duplicates, passing independent checks, agent success, and return within the deadline.
| Measurement | Tablaze | Browser Use |
|---|---|---|
| Agent reported completion | 59.247 s | 75.354 s |
| End-to-end time | 59.415 s | 93.706 s |
| Model calls (including judging) | 4 | 4 (1 judge) |
| Input / output tokens | 59,123 / 664 | 70,868 / 1,002 |
The two clocks start at different boundaries; Browser Use totals already include judging. This single follow-up does not replace round 3’s 4/5, combine success rates, or establish a stable advantage. Monetary cost was not estimated.
Follow-up report and actual terminal-event evidence ↗The page renders only seven rows at once. Both engines found VIRTUAL-130 beyond the initial DOM and made exactly one reservation, with zero duplicates; both agents finished successfully before the deadline.
| Measurement | Tablaze | Browser Use |
|---|---|---|
| End-to-end time | 86.830 s | 141.597 s |
| Model calls (including judging) | 6 | 5 (1 judge) |
| Input / output tokens | 93,209 / 944 | 91,929 / 2,746 |
2026-09-23, gpt-6-astra / ultra, shared 120,000-token budget. Browser Use's default judge is included in time and usage. One visible development task cannot establish overall or stable superiority. Two 50,000-token attempts hit the comparison gateway budget; see the report.
Virtual-list report, measurements, and conditions ↗In two new matched tasks, both agents made one correct write per task with no duplicates and completed before the deadline. Tablaze later added optional action post-checks; a separate Codex Shadow DOM run finished in two model calls, without a concurrent Browser Use arm.
| Initial matched task / whole run | Tablaze | Browser Use |
|---|---|---|
| Shadow DOM form | 87.539 s ✓ | 69.014 s ✓ |
| Delayed menu | 90.417 s ✓ | 81.933 s ✓ |
2026-09-23, gpt-6-astra / ultra, seed 31, one attempt per task and arm; Browser Use's default judge is included. Tablaze took longer on both matched tasks. A separate Tablaze-only post-check attempt took 42.012 s and two model calls; cross-run differences do not prove a stable speedup or overall superiority.
Full report, source states, and raw traces ↗Click an observed menu opener, then wait for one uniquely named new button or menu item. In the final visible run, Codex combined both steps into one Tablaze action call. Both agents passed the independent business judge with one write and no duplicates.
| Seed 37 / whole run | Tablaze | Browser Use |
|---|---|---|
| Delayed menu | 80.715 s ✓ | 77.720 s ✓ |
| Model calls | 4 | 4 |
2026-09-24, gpt-6-astra / ultra, one visible development task; Browser Use's default judge is included. An earlier role-bearing version led the model to guess menuitem incorrectly, then recover without a duplicate write. The final API matches a unique name across buttons and menu items. Cross-version observations do not prove stable speed or overall superiority.
Full report and all failure/success traces ↗A local provider popup authorizes the original app tab. Both engines made one authorization and one correct submission with zero duplicates, passing business, Agent, and deadline checks.
| Measurement | Tablaze | Browser Use |
|---|---|---|
| End-to-end time | 123.719 s | 119.515 s |
| Model calls (including judging) | 8 | 6 (1 judge) |
| Input / output tokens | 127,014 / 1,245 | 111,839 / 1,517 |
2026-09-23, gpt-6-astra / ultra, shared 250,000-token budget and 240-second deadline. Browser Use's default judge is included in time and usage. It was slightly faster here; a synthetic popup task does not establish real OAuth capability or an overall ranking.
Authorization-return report, measurements, and conditions ↗An independent server checks for exactly one correct write. All four entries passed form and canvas. On the virtual list, three direct-tool entries passed; structured direct tools did not complete.
| Task / Codex process time | Tablaze | Harness MCP | BU CLI-MCP | BU MCP |
|---|---|---|---|---|
| Form | 69.331 s ✓ | 111.232 s ✓ | 93.260 s ✓ | 74.824 s ✓ |
| Visual canvas | 65.037 s ✓ | 77.263 s ✓ | 71.074 s ✓ | 71.326 s ✓ |
| Virtual list | 90.791 s ✓ | 175.487 s ✓ | 193.335 s ✓ | 195.185 s, no write |
Scroll the table sideways to view all entries.
2026-09-23, gpt-6-astra / ultra, one direct-tool attempt per arm and task. Times cover the Codex process, model, MCP and auto-review, excluding Chrome startup and independent judging. The structured MCP's nested Agent separately passed after a longer tool timeout (318.247 s, six additional model calls); its first nested attempt timed out. Both programmable Browser Use entries passed the virtual list. A direct-tool failure is not a product-wide failure, and these samples do not establish overall superiority.
Four-entry report, full fallback and raw traces ↗ · Initial two-entry record ↗After a local provider popup returns, the receipt reference exists only in the network JSON response. The independent server requires one authorization and exactly one correct submission. Tablaze's response tool is opt-in; Browser Use's programmable CLI-MCP also passed.
| Task / Codex process time | Tablaze | Harness MCP | BU CLI-MCP | BU MCP |
|---|---|---|---|---|
| Read response and submit once | 149.560 s ✓ | 190.380 s ✓ | 291.881 s ✓ | 300.012 s · deadline, no write |
Scroll the table sideways to view all entries.
2026-09-23, gpt-6-astra / ultra, one visible development attempt per direct entry. Timing includes Codex, model, MCP and auto-review, excluding Chrome startup, server judging and cleanup. Two separate structured-MCP nested-Agent attempts also failed the server judge, but used different deadlines and a second Agent, so are not pooled with this table. Browser Use CLI-MCP passed; this cannot establish overall superiority or a stable speed advantage.
Network-response report, failed attempts and raw data ↗ · Opt-in network-tool contract ↗The same Codex model handled an authenticated-response task. Tablaze enabled page scripts explicitly; Browser Use used CLI-MCP. The independent server recorded one correct submission and zero duplicate writes per entry.
| Direct entry | Business judge | Codex process | Input / output tokens |
|---|---|---|---|
Tablaze --page-script | Pass, 1 write | 198.359 s | 317,115 / 1,557 |
| Browser Use CLI-MCP | Pass, 1 write | 239.400 s | 340,202 / 2,461 |
2026-09-23, gpt-6-astra / ultra, one visible development attempt per arm, 300 s deadline. Tablaze's Chrome launch fell inside its process time; Browser Use's isolated Chrome started before its timer. Times are not a controlled speed ranking, and one synthetic task does not establish overall superiority.
Page-script report and raw traces ↗ · Page-script authority and limits ↗↳These are separate visible development tasks with budgets and deadlines stated in their own reports. Business outcome, Agent completion, and whole-run return are recorded separately; the samples do not establish overall superiority.
For the current SDK extension, see typed custom tools: input/output validation, trusted application context, and browser bindings. The real-model records above cover their original source versions, not success rates for the new extension or production provider adapters.
NEW PROVIDER / SEPARATE VALIDATION
Two visible tasks through the new production entry point, one attempt each: independent acceptance and complete Agent success 2/2, zero duplicate writes. This is not a new matched Browser Use comparison.
| Task | Whole run | Input / output tokens |
|---|---|---|
| Form | 64.030 s | 43,391 / 462 |
| Avoid duplicate order | 78.943 s | 59,401 / 710 |
gpt-6-astra / ultra, seed 23, including startup and cleanup; usage includes Codex CLI context overhead. Anthropic / Ollama have local protocol tests only. Overall superiority remains unproven.
Production provider report and raw data ↗ · Four provider options ↗HISTORICAL LOCAL SAMPLE / 0.1.0
Preserved 0.1.0 fixture timings, not current development-branch performance. Real-model comparisons are recorded separately above.
LOCAL CHROMIUM / HOTEL SEARCH FIXTUREPreparing measurement reportFive historical fixture runs on Apple M3 Max, 2026-09-22 04:40 UTC, with source hashes retained in the raw JSON. Direct engine calls exclude MCP, models, and internet websites. The command below tests your current checkout; it does not recreate the old version.
Raw measurement data ↗npm run benchSTART WITH THE FIRST CALL.
Node.js 20+ and Chromium. Start in an isolated browser, or explicitly connect a Chrome session you choose.
Download 0.1.0 package ↓The downloadable 0.1.0 package and source ZIP are the historical eight-tool preview. For sixteen tools, the agent loop, and typed custom tools, build the current GitHub source ↗.
Follow the source build instructions on this page. A public npm release is not available yet.Build from source; npm publication is pending
git clone https://github.com/SweetDianDian/tablaze.git
cd tablaze
npm ci
npm run build
node dist/cli.js setup
node dist/cli.js doctorUse the built server’s absolute path
codex mcp add tablaze -- \
node /absolute/path/to/tablaze/dist/cli.jsFor clients that use mcpServers configuration
{
"mcpServers": {
"tablaze": {
"command": "node",
"args": ["/absolute/path/to/tablaze/dist/cli.js"]
}
}
}CLEAR EXPECTATIONS.
The core browser tools make no model API calls; your MCP client supplies reasoning. The development agent accepts Codex, Anthropic, Ollama, or compatible endpoints. Codex uses its existing CLI login; other providers use their own credentials. Choose the model explicitly. See provider setup.
Yes, through an explicit CDP connection. The default launches an isolated browser. Attached mode shares that Chrome profile’s identity, and Tablaze closes only the tabs it owns.
The current source supports exact-origin allow/deny rules through CLI/SDK configuration for HTTP(S) document requests, including redirects, frames and popups, in owned isolated browsers. External CDP is excluded. Fetch, images and other non-document traffic remain unrestricted; interception connection loss can release requests. This is not a network firewall. See navigation policy and limits.
Actions include a snapshot revision and element reference. Detected stale snapshots or changed targets are rejected; re-observe before retrying. Validation and input are not atomic, so a short race remains. These checks are not a security sandbox.
It does not promise success on every website, overall superiority over Browser Use, or rollback. The current branch supports explicit file workflows and owned popups; closed shadow roots and arbitrary complex sites remain limitations. Checks and input are not atomic; reconcile uncertain submissions before retrying.
BROWSER ACTIONS. LESS WAITING.