Import AITURK IDE 1.0.0-beta.1 from Hermes 63279301; preserve MIT license

This commit is contained in:
2026-09-05 13:26:46 +03:00
commit 03634b1ca3
11340 changed files with 3442369 additions and 0 deletions
+114
View File
@@ -0,0 +1,114 @@
# Browser Use Mode Benchmark
The A/B battery behind PR [#81958](https://github.com/NousResearch/hermes-agent/pull/81958)
(Browser Use CLI 3.0 mode, salvage of #66476 by @laithrw): built-in
`browser_*` toolset vs the single `browser_exec` driver, measured as total
task tokens / tool calls / wall clock at accuracy parity on live multi-step
web tasks.
## Design
- **Arms differ only by tree + config.** `base` runs the built-in twelve
`browser_*` tools from a merge-base checkout; `pr` runs `browser_exec`
(`browser.backend: browser-use`) from the branch checkout; `prns` is `pr`
with the schema's helpers digest stripped to the header (isolates the
digest's value). Each cell gets a throwaway `HERMES_HOME`; web-fetch
credentials are stripped so every arm must actually drive the browser.
- **Tasks are oracle-checked.** toscrape-family sites (stable content, no
anti-bot), regex oracles over the final answer. `tasks/easy.json` (5 tasks:
price lookup, category extract, count/aggregate, login, pagination) and
`tasks/hard.json` (6 tasks: full-category multi-page crawls, five-star
rating aggregation, JS/delayed render, login chain, cross-category
compare).
- **Resume-safe.** Completed cells in `results/*.jsonl` are skipped on rerun
(same pattern as `scripts/toolperf_abeval`).
- **Backend matrix.** `orchestrate.py` drives a local headless-Chrome CDP;
`orchestrate_cloud.py --backend nous-cloud|browserbase` provisions a real
cloud browser per cell through the same provider plumbing the product uses.
## Run
```bash
# arms are pinned checkouts — e.g. merge-base worktree vs your branch
export BUBENCH_BASE_TREE=/path/to/merge-base-tree
export BUBENCH_PR_TREE=/path/to/branch-tree
```
Note: since #81958 merged (and #85170 made Browser Use the default driver),
a current-main checkout resolves to `browser_exec` in BOTH arms. The `base`
arm only measures the built-in `browser_*` toolset when `BUBENCH_BASE_TREE`
is pinned to a pre-#81958 tree (the original run used the PR's merge-base
worktree). For future A/Bs of new browser changes, pin `base` to the
merge-base of the change under test — the arms are generic.
```bash
google-chrome --headless=new --remote-debugging-port=9333 \
--user-data-dir=/tmp/bubench-chrome --no-first-run --disable-gpu about:blank &
python3 orchestrate.py --tasks tasks/hard.json --reps 3 # 108 cells @ 2 models x 3 arms
python3 report.py results/results.jsonl
```
## Baseline scorecard (Aug 8-10 2026, the #81958 run — 204 cells total)
**Hard-task battery, local Chrome CDP** (6 tasks x 3 reps per cell; final
corrected-oracle readout, nothing excluded):
```
model arm ok tok_mean tok_med calls wall_s vs base tok
opus4.8 base 18/18 64594 63776 4.1 25.2 —
opus4.8 pr 18/18 25934 25030 2.0 17.5 -60%
opus4.8 prns 18/18 25578 27934 3.2 23.7 -60%
kimi-k3 base 18/18 56464 53276 5.3 50.0 —
kimi-k3 pr 18/18 19230 16710 2.4 33.3 -66%
kimi-k3 prns 18/18 23099 21160 4.1 50.5 -59%
```
Digest ablation: pr (with helpers digest) 36/36 ok, mean 22,582 tok; prns
(header-only) 36/36 ok, mean 24,339 tok — the pinned 3.4KB digest costs
nothing and saves a little; the full 11KB live skill dump adds nothing.
**Backend matrix** (pr arm, same tasks):
```
model backend ok tok_mean calls wall
opus4.8 local-cdp 17/18 25934 2.0 17.5
opus4.8 nous-cloud 12/12 33330 2.8 33.8
opus4.8 browserbase 6/6 26712 2.2 23.2
kimi-k3 local-cdp 18/18 19230 2.4 33.3
kimi-k3 nous-cloud 12/12 22050 2.9 41.4
kimi-k3 browserbase 6/6 22121 2.8 35.2
```
**Easy battery, round 1** (5 tasks x 3 reps, sonnet-5 + qwen3-coder-30b;
after excluding provider-noise runs — raw chat-template XML, 0 tool calls):
```
model arm ok prompt compl total calls wall_s
claude-sonnet-5 base 15/15 39771 324 40095 2.7 16.5
claude-sonnet-5 pr 15/15 27482 509 27991 2.4 14.3
qwen3-coder-30b base 13/14 59509 559 60068 5.7 21.5
qwen3-coder-30b pr 10/11 57146 1616 58763 6.8 26.3
```
sonnet-5: 30% tokens at parity. qwen3-30b: a wash — weak coders burn the
savings retrying exec code. The token win concentrates on multi-step tasks
and grows with task hardness; strong models also finish in fewer tool calls.
Compatibility probes from the same run: Firecrawl cloud browsers attach fine
(CDP websocket); Camofox has no CDP surface — structurally incompatible,
hence the automatic fallback to the built-in toolset in #81958.
Caveats: toscrape-family sites (no anti-bot, no heavy SPA); n<=3 per cell;
success-rate deltas at this n are noise — audit sub-100% cells run-by-run
before calling a regression.
## Provenance
The original per-run `results*.jsonl` files lived in `/tmp/bu-bench/` (tmpfs)
and were lost in a host reboot on Aug 12 2026. The harness, task definitions,
and aggregate readouts in this directory were recovered verbatim from the
session transcripts of the benchmark run (session `20260808_050008_5f615e`
tool-call history); `single_run.py`/`orchestrate*.py` are the recovered
scripts with the hardcoded `/tmp/bu-bench` paths parameterized. Rerunning the
battery reproduces fresh per-run data.