Files
hermes-skills/model-capability-benchmark/SKILL.md
T

7.0 KiB

name, description, version, author, license, platforms, metadata
name description version author license platforms metadata
model-capability-benchmark Use when picking a model for a capability. Benchmark. 1.0.0 Hermes Agent MIT
linux
hermes
tags related_skills
benchmark
evaluation
model-selection
ollama
extraction
agent-routing
cognee-brain
save
save-q-memory

Model Capability Benchmark

Use when the user needs to pick which model to use for a specific capability (memory extraction, entity/relation extraction, classification, summarization quality, etc.). The goal is a data-backed model choice, not a vibe check.

When to Use

Load this skill when the user needs to pick which model to use for a specific capability (memory extraction, entity/relation extraction, classification, summarization quality, etc.). The goal is a data-backed model choice, not a vibe check.

Triggers:

  • "which model should I use for X"
  • "run a test of the models we have"
  • "compare models on "
  • Selecting a model for a pipeline (memory extraction, Cognee-style work, etc.)

Method

  1. Run the SAME test on every candidate — no variations. When the user says "test these models" or "do the same test", send the identical prompt, identical temperature, identical system message to every model. Do NOT add toggles (thinking on/off, different prompts, extra modes) unless the user explicitly asks for a comparison of those modes. The user corrected this directly: "WHY DIDN'T you just do the same test??? DO not turn off thinking. Just DO THE SAME TEST." A benchmark's value is comparability; a variation you introduce silently breaks it.

  2. Enumerate the candidate models. For Ollama cloud models, ollama list shows them; curl -s http://localhost:11434/api/tags reveals remote_model/remote_host (cloud models have a remote_host; local models have a real byte size). Filter to the requested scope (e.g. "cloud only" = has remote_host). For a non-Ollama endpoint (e.g. a llama.cpp server on another box), hit its OpenAI-compatible /v1/chat/completions directly with urllib — same system prompt, same temperature: 0, same test cases. The harness just needs a chat(messages) function; the scoring is identical regardless of backend.

  3. Write a test-case set. Each case = input turns + expected extraction + explicit fail conditions. Cover the hard cases for the capability, not just happy paths. For memory extraction the canonical hard cases are: identity/stable facts, preference vs one-off, update/contradiction, relative time, negation, specificity, multi-entity relation, pronoun resolution (needs prior turn), plan vs completed fact, abstention (no facts), assistant contamination, quantity/attribute, soft preference + intensity, two-people-easy-to-merge.

  4. Build a harness that sends each case to each model with a fixed system prompt demanding a clean structured list (JSON array), temperature=0, and captures raw output. Save raw outputs to disk (JSON) — never score in the same pass that runs.

  5. Score kept / missed / invented per case, then total. A good extractor is high recall (kept) with near-zero inventions. For Cognee-style work also reject answers that aren't clean entity/relation lists.

  6. Pick the winner on recall + zero inventions, and report per-case detail for the interesting failures (not just totals).

Pitfalls

  • Normalize BOTH sides before substring scoring. Expected facts and model output must go through the same re.sub(r'[^a-z0-9 ]', ' ', s.lower()) — otherwise hyphenated facts ("sci-fi", "15-gauge", "fine-tune", "gpt-oss-120b") are falsely marked missed. This bit the first scoring pass.
  • Retired cloud models return HTTP 410. A model in ollama list can still be retired upstream; every call returns ERROR: ... was retired at ... (status code: 410). Detect this and exclude the model rather than scoring it as zero-kept.
  • Verify relative-time anchors yourself. "last Tuesday" relative to a reference date must be computed with date -d <ref> +%A etc. The user's expected answer may itself be wrong (e.g. "Aug 19" when Aug 19 is a Wednesday) — compute the correct date and flag the discrepancy rather than silently scoring against a wrong target.
  • Contamination test needs an assistant-role turn. To test that the model ignores assistant-suggested facts, the middle turn must be role: "assistant", not user.
  • Substring scoring can give false credit. A model that answers instead of extracting (e.g. recommends vector DBs) may contain the expected keywords in prose and score as "kept" when it actually failed the format. Inspect raw output for the cases that matter before trusting the score.
  • Score from raw output, not from a live re-run. Models are non-deterministic even at temperature 0; re-running changes results. Persist raw outputs and score the file.
  • Large cloud models are slow — run sequentially in the background. A 397b/675b/120b cloud model can take ~20s per test case, so 14 cases ≈ 5 min per model. A single foreground run of several models will hit the terminal timeout. Run one model per background process (or a for loop over models in ONE background process), and poll the error log for DONE <model> markers rather than blocking on wait. Background processes that get killed mid-run leave a partial JSON — check which models actually completed before scoring.
  • Strip fences/whitespace in the abstention check. A model that correctly returns an empty list may wrap it as json\n[]\n or [ ]. The abstention check must strip ``` fences and whitespace and accept [], [ ], "", null — otherwise a correct abstention is falsely scored as a failure.
  • Negation FORBID substrings false-positive. A FORBID fact like "allergic to shellfish" will match inside a correct negation ("not allergic to shellfish") and be falsely scored as invented. For negation cases, either check for the negated form explicitly (e.g. forbid "allergic to shellfish" only when NOT preceded by "not"), or inspect the raw output manually before trusting an "invented" flag on a negation test.
  • Non-Ollama endpoints need their own auth. A llama.cpp server (e.g. e1's qwen38-27b at http://10.0.0.26:8099/v1) may require a Bearer key while /health and /v1/models stay public — only a real completion proves auth. Reuse the same harness by swapping the transport; don't assume the ollama python client works for it.

Support Files

  • references/memory-extraction-results.md — 2026-08-26 benchmark of 11 Ollama cloud models on memory extraction, with per-model findings and the winner.
  • references/recommended-sampling-params.md — official temperature/top_p guidance per model (DeepSeek V4 Pro = 1.0/1.0, MiniMax M3 = 1.0/0.95, API use-case table), plus how temperature is actually set for Hermes custom providers (extra_body.temperature, not a top-level key) and how to research user-experience threads (HN Algolia API). Load when the user asks about optimal temperature for a model.