Skip to content

Weekly benchmark review: 2026-08-17 #178

Description

@github-actions

Weekly benchmark review (2026-08-17)

Automated check from scripts/weekly-benchmarks-check.mjs. Triage and either:

  • Update data/benchmarks.json if a new flagship model dropped this week, then close this issue, OR
  • Comment noop and close if nothing actionable surfaced.

Current state of data/benchmarks.json

  • lastUpdated: 2026-08-02 (15 days ago)

  • Models tracked: 24

  • Benchmarks tracked: 9

  • Models released within last 60 days: 2

    • 2026-07 | Moonshot AI | Kimi K3
    • 2026-07 | Anthropic | Claude Opus 5

Model-release-flavored news, last 7 days

Matched 13 articles (keyword scan; not all will be real releases).

Date Source Title
2026-08-17 ZDNet AI Google Workspace lets Gemini access your company data by default - how to shut it down
2026-08-17 The Verge AI Anthropic explains how Claude’s invisible text watermarks will work
2026-08-17 Google AI Blog Get closer to the game with Gemini and Pixel
2026-08-15 WIRED AI Amazon Can Use Your Twitch Content to Train Its AI—Unless You Opt Out
2026-08-14 ZDNet AI Google Meet can take notes for your in-person meetings now - here's how it works
2026-08-14 The Verge AI You can now turn off Google Gemini’s visible watermarks
2026-08-14 ZDNet AI This free Android assistant fixes my biggest Gemini frustration - and keeps my data private
2026-08-13 ZDNet AI Gemini voice calling on Android Auto keeps failing me - and Google has until September to fix it
2026-08-12 Hugging Face Blog Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
2026-08-12 NVIDIA AI Blog NVIDIA CEO Tops Glassdoor’s 2026 List of Best CEOs
2026-08-12 NVIDIA AI Blog NVIDIA AI Factory Compute Is Becoming an Investable Asset Class
2026-08-11 Google AI Blog AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.
2026-08-11 NVIDIA AI Blog NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI

Sources: ZDNet AI (4), NVIDIA AI Blog (3), The Verge AI (2), Google AI Blog (2), WIRED AI (1), Hugging Face Blog (1)

HF Open LLM Leaderboard top 10

Captured: 2026-08-17

Rank Model
1 MaziyarPanahi_calme-3.2-instruct-78b_bfloat16 (avg 52.08 · 78B)
2 MaziyarPanahi_calme-3.1-instruct-78b_bfloat16 (avg 51.29 · 78B)
3 dfurman_CalmeRys-78B-Orpo-v0.1_bfloat16 (avg 51.23 · 78B)
4 MaziyarPanahi_calme-2.4-rys-78b_bfloat16 (avg 50.77 · 78B)
5 huihui-ai_Qwen2.5-72B-Instruct-abliterated_bfloat16 (avg 48.11 · 73B)
6 Qwen_Qwen2.5-72B-Instruct_bfloat16 (avg 47.98 · 73B)
7 MaziyarPanahi_calme-2.1-qwen2.5-72b_bfloat16 (avg 47.86 · 73B)
8 newsbang_Homer-v1.0-Qwen2.5-72B_bfloat16 (avg 47.46 · 73B)
9 ehristoforu_qwen2.5-test-32b-it_bfloat16 (avg 47.37 · 33B)
10 Saxo_Linkbricks-Horizon-AI-Avengers-V1-32B_bfloat16 (avg 47.34 · 33B)

What "needs update" usually means

  1. A flagship from Anthropic / OpenAI / Google / Meta / Mistral / DeepSeek / xAI launched this week → add a row to data/benchmarks.json.
  2. A tracked model has materially-shifted benchmark scores (re-running, methodology change) → update the row.
  3. A new benchmark itself (e.g. a successor to MMLU-Pro) is becoming canonical → add it.

What it usually does NOT mean

  • Research papers about benchmarks (those land on /research, not /benchmarks).
  • HN opinion threads about a model.
  • Pricing-only changes (those go in data/pricing.json).

Bump lastUpdated in data/benchmarks.json whenever you change anything else in the file.

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarksWeekly benchmark score reviewmaintenanceScheduled review issues for ongoing maintenance

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions