# Real MCP demonstration / 真实 MCP 演示

The current website recording is **154.104 seconds at 2880×1800**, with eleven segments of conversational English male neural narration, optional captions (off by default), seven chapter shortcuts and a visible action-guide cursor. The guide points to measured targets over authentic browser captures. Its movement is paced for narration; it is not a recording of the physical mouse. The MCP trace and tool timings are unchanged.

官网当前演示为 **154.104 秒、2880×1800**，包含十一段英语男声旁白、默认关闭的中英字幕、七个章节和可视操作指针。指针按本地页面实测控件位置标注，配合讲解节奏移动；它是讲解辅助，并非真实物理鼠标录屏。原有 MCP 调用记录和工具耗时没有改变。真实使用时，Tablaze 也会在通过引用验证的点击、输入等动作上短暂显示指针，可关闭。

## Reproduce / 复现

Install the source dependencies and build the runtime. The current voice is Microsoft Edge's `en-US-AndrewMultilingualNeural` (male), synthesized from the checked-in [English script and Chinese subtitles](https://github.com/SweetDianDian/tablaze/blob/main/demo/narration.en-US.json) with [`edge-tts`](https://github.com/rany2/edge-tts). This recipe sends the public narration text to the online speech service. Use an FFmpeg executable with `libx264` and AAC:

```sh
npm ci
npm run build
python3 -m venv /absolute/path/voice-venv
/absolute/path/voice-venv/bin/pip install -r demo/requirements-narration.txt
/absolute/path/voice-venv/bin/python demo/narrate-neural.py \
  --script demo/narration.en-US.json \
  --output /absolute/path/narration --ffmpeg /absolute/path/ffmpeg
TABLAZE_BROWSER_CHANNEL=chrome \
TABLAZE_DEMO_FFMPEG=/absolute/path/ffmpeg \
TABLAZE_DEMO_NARRATION=/absolute/path/narration/narration.json \
node demo/record.mjs
```

The generator measures the decoded duration, peak/RMS and SHA-256 of all eleven clips. It uses the male voice at its native pitch, with ordinary punctuation and a relaxed speaking rate. It does not pitch-shift an existing voice. Online synthesis may change over time; the published manifest records the actual generated bytes. `TABLAZE_DEMO_OUTPUT` optionally selects a new recording directory.

生成器通过在线语音服务合成英语男声，逐段解码，测量时长、音量和哈希。使用男声本身的音色，不对旧女声降调；句子缩短，保留自然停顿。在线服务的合成结果可能变化，已发布版本保留实际音频的哈希记录。

To update narration on an existing complete recording without changing its pictures or tool trace:

```sh
TABLAZE_DEMO_FFMPEG=/absolute/path/ffmpeg \
node demo/revoice.mjs /absolute/path/original-recording \
  /absolute/path/narration/narration.json /absolute/path/revoiced-recording
```

Each clip must fit its original presentation hold. `revoice.mjs` copies the H.264 packets, verifies identical picture-stream hashes before/after, regenerates captions and records a separate `media_revision`. It keeps the original events, assertions, tool timings and chapters. This is an audio edit, not a new benchmark run. The legacy `narrate.mjs` recipe remains available for offline macOS synthesis of the historical v3 script.

To annotate the narrated recording with the measured action guide:

```sh
TABLAZE_DEMO_FFMPEG=/absolute/path/ffmpeg \
node demo/annotate-cursor.mjs /absolute/path/male-narrated.mp4 /absolute/path/annotated.mp4
```

[`cursor-track.json`](https://github.com/SweetDianDian/tablaze/blob/main/demo/cursor-track.json) records native screenshot control centers and narration-paced intervals. `annotate-cursor.mjs` verifies that the AAC stream is copied unchanged and emits an evidence JSON beside the new MP4. These overlays are editorial guides over actual still captures, not new browser actions or a continuous mouse recording. In the live browser, the pointer appears only briefly after an action target passes validation and never receives pointer input. The CLI switch `--no-visual-pointer` turns it off.

The recorder starts the actual compiled stdio server, connects a real MCP SDK client and completes one continuous travel workflow in a localhost fixture. It verifies browser results, server counters and downloaded bytes. An isolated presentation browser renders actual returned JSON and `tab_capture` images. At each display update it saves a lossless 2× PNG; those frames retain the observed timeline. `render-video.mjs` produces 2880×1800 H.264 at CRF 16/12 fps and AAC narration at 48 kHz, targeting −16 LUFS. The embedded browser captures remain the engine's original 1280×800 JPEG screenshots; the larger form typography improves their legibility, not their native resolution. Frame scheduling is quantized to 1/12 second and is not a continuous browser video feed.

录制器通过真实 MCP SDK 完成整条任务，再用实际响应和截图呈现过程。每次展示更新保存 2 倍像素的无损 PNG，并按原时间轴编码；嵌入的网页截图仍为引擎返回的原始 1280×800 JPEG，表单字体放大提高了可读性，没有冒称网页截图本身是 2880 像素。视频帧时间对齐到十二分之一秒，不是被控页面的连续视频流。

Output defaults to `demo/output/`:

- `tablaze-demo-hd.mp4`: the current high-resolution recording with English narration.
- `narration.en-US.vtt` and `narration.zh-CN.vtt`: eleven English caption cues and their Chinese translations, synchronized to the actual audio.
- `demo-report.json`: all requests/results, tool timing, source/frame/audio/video hashes, chapter timing and independent assertions.
- Five original JPEG captures, `poster.png`, and the verified `wayfar-itinerary.csv`.
- `frames/` and `frames.ffconcat`: local lossless presentation inputs, retained for inspection.

The current website loads the [identical MP4 from a fixed GitHub commit](https://raw.githubusercontent.com/SweetDianDian/tablaze/821c441eea6433dcc6c1bbadc041c2e7030107cb/demo/media/tablaze-demo-hd.mp4?v=34223f1e). The fixed commit identifies the exact published bytes; external media hosting avoids recurring large-archive upload timeouts. The VTT, report, CSV, five captures and `poster-hd.png` remain hosted with the website. For self-hosting, copy all output assets into the website's `demo/` directory and change the video URLs back to `demo/tablaze-demo-hd.mp4`; copy `poster.png` as `poster-hd.png`. Set `release.json` to `"demo": true`. Generated videos, audio and frames are excluded from source packages. The current reproduction scripts are in this GitHub repository; downloadable website preview packages retain their labeled original build.

Without narration configuration, the original silent WebM recipe remains available. `TABLAZE_DEMO_QUICK=1` is only for a short harness check and cannot be used with narrated publication. Do not present quick-mode output as the normal-speed website recording.

## What is shown / 展示内容

1. The SDK discovers fifteen tools and opens the travel form in one session.
2. One six-action batch sets Lisbon, three nights, two travelers, a €900 budget and free cancellation, then searches. Server counters confirm one search with those filters.
3. Five assertions check the resulting URL, title, visible text, city value and matching hotel count.
4. The fixture replaces the observed approval button. Its stale reference is rejected with zero completed actions; server counters confirm zero approval popup visits, submissions and orders at that point.
5. A fresh snapshot supplies a new reference. Explicit `follow-single` policy follows the owned approval popup, then four actions complete its form.
6. Three receipt assertions pass. The server independently confirms one submission and one order.
7. The saved CSV's bytes and SHA-256 match the expected contents. Both owned tabs close and zero sessions remain.

同一条差旅流程展示：六步批量填表、五项结果验收、目标变化后拒绝旧引用且没有审批写入、重新观察恢复、跟随自有弹窗、四步审批、三项回执检查、CSV 内容和哈希核对，以及关闭两个所属标签页。

## Recorded result / 本次结果

The underlying browser workflow passed on **2026-09-23 (Asia/Shanghai)**: 22 MCP calls, eight browser assertions, one approval order, one matching 145-byte CSV and zero remaining sessions. The 154.104-second video retains the original 138.063 seconds of presentation holds. The v4 edit contains 96.984 seconds of English male narration and preserves the original picture stream and chapter positions. Recorded tool calls total 2,161.036 ms; the display, narration and holds are excluded. This is one local scripted sample, not a model or Browser Use speed comparison.

本次正式录制通过：22 次 MCP 调用、八项浏览器检查、一份审批订单、一个内容匹配的 145 字节 CSV，零遗留会话。视频 154.104 秒，保留原来约 138.063 秒的讲解停留；v4 含约 96.984 秒的英语男声旁白，只替换声音与字幕，画面和章节不变。实际工具耗时合计 2,161.036 毫秒，不能作为模型执行能力或 Browser Use 的速度对比。

v5 在原有讲解视频上叠加操作位置指引，AAC 语音数据保持逐包相同；视频画面因此重新编码，原始 MCP 记录、章节和工具耗时均未改动。旧引用被拒绝的一段只指向候选控件，没有模拟成功点击。

[Complete trace](https://github.com/SweetDianDian/tablaze/blob/main/docs/evidence/demo-run.json) · [34 website checks](https://github.com/SweetDianDian/tablaze/blob/main/docs/evidence/demo-site-qa-v4.json) · [Decoded narration evidence](https://github.com/SweetDianDian/tablaze/blob/main/docs/evidence/demo-narration-v4.json)

[Cursor annotation evidence](https://github.com/SweetDianDian/tablaze/blob/main/docs/evidence/demo-cursor-v5.json)

The v4 media/player release passed 34 browser checks, including 2880×1800 video metadata, an actual decoded audio track, initially unmuted playback, keyboard sound control, eleven English cues and eleven Chinese subtitle cues, complete video decode and all seven chapter jumps. All eleven source clips and the final AAC are decoded and checked separately; the narration evidence includes the actual measurements and exact voice metadata. These checks establish audio in the file and player; local device volume is controlled by the viewer.

This deterministic SDK demonstration does not invoke a reasoning model, make external bookings or payments, or prove general website reliability or superiority over Browser Use. Source hashes bind the trace to the recorded runtime. The original [female-narrated v3](https://github.com/SweetDianDian/tablaze/blob/main/docs/evidence/demo-run-v3.json), [short v1](https://github.com/SweetDianDian/tablaze/blob/main/docs/evidence/demo-run-v1.json) and [silent 126-second v2](https://github.com/SweetDianDian/tablaze/blob/main/docs/evidence/demo-run-v2.json) traces remain historical records.

The later [caption-default check](https://github.com/SweetDianDian/tablaze/blob/main/docs/evidence/demo-caption-default-v1.json) verifies that both tracks start disabled, the approval scene is unobstructed, English audio decodes, and an optional caption choice survives language and chapter changes. Captions remain available through native video controls.

字幕显示调整已单独验证：首次播放默认不显示字幕，审批画面不再被大字幕遮挡，英语声音正常；手动开启字幕后，切换语言或章节会保留选择。
