Skip to content

Correct the AIMultiple benchmark date, scope, and tier framing - #1363

Open
hmishra2250 wants to merge 3 commits into
mainfrom
docs/search-measured-performance
Open

Correct the AIMultiple benchmark date, scope, and tier framing#1363
hmishra2250 wants to merge 3 commits into
mainfrom
docs/search-measured-performance

Conversation

@hmishra2250

Copy link
Copy Markdown
Contributor

Summary

The measured-performance table added to features/search.mdx dated the AIMultiple study by the
article's update date rather than the measurement, omitted that its query set is AI/LLM-domain only,
and carried a Note whose "ties, not a ranking" sentence contradicted the table's own "rank 2 of 8".
This branch corrects all three against the source, and links the first "15 tokens" callout on
features/extract.mdx to the billing and per-page credit statements.

Why (evidence)

Source report: Weekly Deep Insights DI-2026-09-03-WEEKLY, report date 2026-09-03.

Findings: F6-SCALAR-ASYMMETRY, grade C minus, verdict file verdicts/deep-F6-benchmarks.md
(docs.firecrawl.dev is the owned host agents actually reach per the exposure ledger, yet carried no
measured-performance information), and F1-PRICING-UNIT, grade B minus, verdict file
verdicts/deep-F1-pricing.md (disconnected credit statements).

Verified against the archived source at review-grounding/aimultiple.md:

  • Line 536: "This reflects December 2025 snapshot only." Lines 175 and 667: "updated on May 25,
    2026" and "Retrieved May 25, 2026".
  • Line 535: "All queries are AI/LLM-related. Results don't generalize to medical, legal, e-commerce,
    or general domains." Line 494: "10,000 new datasets by randomly sampling 100 queries with
    replacement."
  • Line 248: "Ranked in the top tier, with no statistically significant difference compared to
    Firecrawl, Exa, or Parallel Search Pro." Line 500 confirms paired bootstrap.
  • Line 521: Agent Score 14.58, 95% CI 13.12 to 15.98, Mean Relevant 4.30 of 5, Quality 3.39 of 5,
    rank 2 of 8. All match the table exactly.
  • Line 422: "Source: Top 500 queries from AIMultiple.com organic search traffic (Dec 2024 to Jan
    2025)."

The removed phrase "the ~0.3-point gaps between them fall inside the reported confidence intervals"
also understated the actual tier span, which is Brave 14.89 to Parallel Pro 14.21, or 0.68. The
phrase is deleted entirely rather than corrected.

Final verdict: APPROVE per the resolution addendum to
review-grounding/FINAL-REVIEW-docs-mcp.md (the review body records APPROVE-WITH-NITS; the single
nit was resolved in commit 3c66fc32).

Changes

  • features/search.mdx, Measured Performance table. The Date cell reads
    Dec 2025 (source updated May 25, 2026). The n / ± cell reads
    95% CI 13.12–15.98, n=100 AI/LLM-domain queries drawn from AIMultiple's own organic search traffic (10,000 bootstrap resamples); the source notes results don't generalize to other domains.
  • features/search.mdx, the Note below the table. The self-contradicting "ties, not a ranking"
    sentence is replaced with "Point-estimate rank 2 of 8; paired-bootstrap testing found no
    statistically significant gap versus Brave, Exa, or Parallel Search Pro." The link to
    firecrawl.dev/benchmarks is retained.
  • features/extract.mdx, line 41, the top-of-page <Info> and the first "15 tokens" statement a
    reader hits. It now carries "See Billing for how credits map to plan pricing; scrape
    and crawl are billed at 1 credit per page
    (Scrape cost)." The second occurrence at roughly
    line 221, under "Billing and Usage Tracking", was already linked by the prior commit and is
    unchanged, so both occurrences now carry matching text.

Diffstat: 2 files changed, 14 insertions, 2 deletions.

Verification

Anchors resolve: billing.mdx exists, and ## Scraping a URL with Firecrawl is at
features/scrape.mdx:76.

mint broken-links under Node 22.23.2: 208 broken links in 95 files, identical to the pre-existing
baseline and byte-identical to the output on the other two docs worktrees after stripping spinner
frames. Neither features/search.mdx nor features/extract.mdx appears.

mint validate: fails with 17 warnings, all pre-existing "Could not find file" warnings for missing
locale snippet files under /snippets/v2/extract/short/ and /snippets/*/v1/scrape/agent-f1/. No
touched file appears.

Every figure in the table was re-checked line by line against the archived source during the
independent final review, not taken from the fix report.

Not in this PR

  • The same "15 tokens" statement appears in v1/features/extract.mdx,
    api-reference/endpoint/token-usage*.mdx and
    developer-guides/usage-guides/choosing-the-data-extractor.mdx, all unlinked. Out of scope for
    this wave.
  • All localized copies (es/, fr/, pt-BR/, zh/) are untouched; the repo CLAUDE.md prohibits
    modifying them.
  • The F6 discrepancy "openbenchmarks 70.3 ± 1.5, 2nd of 11" is byte-verified from a snapshot but was
    not locatable live on 2026-09-04 (the nearest source, Artificial Analysis, shows Firecrawl roughly
    8th of 12). It is deliberately not quoted anywhere on this page and should be located or retired
    before it is quoted again.
  • DevDex 63.1% Recall@10 is first-party, not third-party, and is deliberately excluded from a table
    of third-party measurements.

Links

Files changed:

  • features/search.mdx
  • features/extract.mdx

Evidence packet:
agent-experience-deepinsights-cleanroom/artifacts/deep-insights-sep3-verification-20260904/

…tics link

Deep Insights Sep-3 verification (F6, F1): docs.firecrawl.dev is the owned
host agents actually reach, yet /features/search carried no measured
performance while competitor pages cite numbers about Firecrawl search.
Add a dated "Measured Performance" table with one independently-verified
third-party figure (AIMultiple Agentic Search Benchmark, Agent Score
14.58 / 4.30 of 5), framed as ties within CI and pointing to
firecrawl.dev/benchmarks for first-party numbers. Two other figures from
the evidence packet (an "openbenchmarks" 70.3 ± 1.5 / 2nd-of-11 claim and
DevDex 63.1% Recall@10) could not be attributed to a matching public
source or are first-party, not independent, and were excluded rather than
guessed — see report for detail.

Also link the extract page's "each credit is worth 15 tokens" statement
to /billing and to the in-docs 1-credit-per-page statement on
/features/scrape, so credit semantics aren't a dead-end statement (F1).
…d credit link (F6/F1 Deep Insights Sep-3)

Grounding review (Deep Insights Sep-3, D1-D3 report) found three defects
in the prior commit:

- The Date column read "2026", but AIMultiple's own Limitations section
  states the benchmark "reflects December 2025 snapshot only"; May 25,
  2026 is only the article's update/retrieval date. Change to
  "Dec 2025 (source updated May 25, 2026)" and disclose that the 100
  queries are AI/LLM-domain queries drawn from the source's own organic
  search traffic, carrying the source's own non-generalizability caveat.
- The Note said the top-tier gaps "are ties, not a ranking", directly
  contradicting the table's own "rank 2 of 8" two lines above, and
  understated the tier span as "~0.3-point gaps" (actual span Brave
  14.89 to Parallel Pro 14.21 = 0.68). Replaced with the source's own
  framing: point-estimate rank 2 of 8, no statistically significant gap
  under paired-bootstrap testing versus Brave, Exa, or Parallel Search
  Pro.
- The prior commit linked "each credit is worth 15 tokens" to /billing
  and the scrape 1-credit-per-page statement only on the second
  (Billing and Usage Tracking) occurrence in features/extract.mdx. The
  first occurrence a reader hits, the top-of-page <Info> callout, was
  left unlinked. Added the same link there.

mintlify broken-links and mintlify validate (Node 22) both pass with only
pre-existing, unrelated warnings (missing zh/es/fr/ja/pt-BR snippet
files); neither touches features/search.mdx or features/extract.mdx.

Local only, no push.
…own traffic

The 100 queries aren't an independent sample; AIMultiple drew them from
its own site's organic search traffic in the AI/LLM domain. Per the
archived source (review-grounding/aimultiple.md:422): "Source: Top 500
queries from AIMultiple.com organic search traffic (Dec 2024 to Jan
2025)" -- narrowed to the 100 used in the benchmark. Making that
self-selection legible on the page, not just in the commit message,
per the D1 nit in FINAL-REVIEW-docs-mcp.md.
@mintlify

mintlify Bot commented Sep 4, 2026

Copy link
Copy Markdown

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
firecrawl 🟢 Ready View Preview Sep 4, 2026, 4:55 PM

💡 Tip: Enable Automations to automatically generate PRs for you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant