7.9 KiB
name, description, version, author, license, platforms, metadata
| name | description | version | author | license | platforms | metadata | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| agent-routing | Routing guide for DeepSeek-V4-Pro orchestrator — which sub-agent (GLM-5.2, Kimi K2.7 Code, MiniMax-M3) to assign for each task type. Eliminates guesswork in peer selection. | 1.0.0 | Hermes Agent | MIT |
|
|
Agent Routing — Sub-Agent Selection Guide
You are DeepSeek-V4-Pro. You have three specialized sub-agents available. Use them deliberately. Do not default to yourself for everything — routing correctly reduces token burn and improves quality.
When to Load This Skill
Load this skill whenever you need to decide which peer agent to dispatch a task to. If you find yourself thinking "which agent should handle this?" — load this skill first.
Sub-Agent Profiles
1. GLM-5.2 (Primary coding & long-horizon engineering specialist)
Core identity: Flagship long-horizon model with a truly usable ~1M context window. Highest-ranked open-source model on long-horizon coding benchmarks (FrontierSWE, PostTrainBench, SWE-Marathon). Supports High / Max effort levels.
Strengths:
- Project-level and multi-file software engineering
- Full development workflows (requirements → design → implementation → testing → multi-platform deployment)
- Complex debugging, performance optimization, large-scale refactors
- Consistent adherence to engineering standards over long sessions
- Strong Terminal-Bench and SWE-bench Pro performance
- Reliable multi-step tool use in coding environments
Use GLM-5.2 when the task involves:
- Building, extending, or refactoring real codebases
- Multi-step coding agents that need to stay coherent across many files or long sessions
- End-to-end feature development or system design + implementation
- Tasks that benefit from High or Max effort modes
- Any serious software engineering work where correctness and engineering discipline matter more than raw speed
Avoid for: Pure research/synthesis over many external documents, heavy multimodal analysis, or very short simple questions.
Dispatch via: ask-glm skill (hermes -p glm chat -q "..." -Q --max-turns 600 --yolo)
2. Kimi K2.7 Code (Long-horizon coding + efficient agentic specialist)
Core identity: Coding-focused agentic model optimized for real-world long-horizon software engineering. ~30% lower thinking-token usage than K2.6 while improving task completion. Strong multi-step tool calling and MCP support. Native multimodal (image + video). 256K context.
Strengths:
- End-to-end software engineering across backend, frontend, infrastructure, systems, security, ML/data, and performance work
- Efficient multi-step tool calling and agent loops (lower token overhead)
- Coherent long coding sessions with interleaved thinking preserved
- Strong real-world coding agent performance (Claude Code / OpenClaw / Hermes style workflows)
- Multimodal input when screenshots, diagrams, or short videos are involved
Use Kimi K2.7 Code when the task involves:
- Long coding sessions that need good token efficiency
- Multi-tool / MCP-heavy development workflows
- Coding agents that must stay coherent and productive over many steps
- Tasks that mix code with visual inputs (UI screenshots, architecture diagrams, etc.)
- Situations where you want strong coding performance without burning as many thinking tokens as GLM-5.2 Max
Avoid for: Tasks that truly need the full 1M context window or the absolute highest long-horizon engineering scores (prefer GLM-5.2).
Dispatch via: ask-kimi-c skill (hermes -p kimi-c chat -q "..." -Q --max-turns 600 --yolo)
3. MiniMax-M3 (Frontier coding + agentic + multimodal + long-context specialist)
Core identity: Coding & agentic frontier model with native multimodality and a true 1M context window (guaranteed ≥512K). Strong autonomous task decomposition, tool use, and browsing/retrieval. First open model that simultaneously delivers frontier coding, million-token context, and deep multimodal alignment.
Strengths:
- Long-range agent tasks and long-range coding
- Native multimodal understanding (text + image) with deep alignment
- Autonomous task decomposition and multi-step reasoning
- Excellent browsing / information retrieval (very high BrowseComp)
- Long-video understanding capability
- Strong general agentic workflows that are not purely code-centric
Use MiniMax-M3 when the task involves:
- Very long context requirements (large documents, many files, long histories)
- Multimodal inputs (images, screenshots, diagrams, video)
- Autonomous research + coding combinations
- Complex agent workflows that need strong decomposition and tool use outside pure coding
- Browsing-heavy or information-gathering + synthesis tasks
- Situations where you want the combination of 1M context + multimodality + coding strength
Avoid for: Pure high-discipline software engineering where GLM-5.2's long-horizon coding scores are superior, or when maximum token efficiency on pure coding is the priority (Kimi K2.7 Code).
Dispatch via: ask-minimax skill (hermes -p minimax chat -q "..." -Q --max-turns 600 --yolo)
Routing Rules (apply these in order)
- Pure or heavy software engineering / multi-file coding / full feature development → GLM-5.2 (default to Max effort for hard tasks).
- Long coding sessions that need better token efficiency or strong multi-tool/MCP loops → Kimi K2.7 Code.
- Tasks needing 500K–1M context, multimodal input, or strong autonomous research + coding → MiniMax-M3.
- Mixed or ambiguous tasks → Prefer the specialist that matches the dominant requirement. You may spawn parallel sub-agents when different aspects benefit from different strengths.
- Simple questions, quick reasoning, or light planning → Handle yourself with the appropriate reasoning mode (No thinking / Thinking / Max thinking). Do not spawn sub-agents unnecessarily.
- Always give the sub-agent a clean, self-contained brief. Do not dump your entire conversation history unless the task genuinely requires it.
Quick Reference Table
| Task Type | Agent | Why |
|---|---|---|
| Multi-file refactor, feature build, system design | GLM-5.2 | Best long-horizon engineering scores |
| Long coding session, MCP-heavy, token-efficient | Kimi K2.7 Code | 30% fewer thinking tokens, strong agentic |
| 500K+ context, multimodal, research+coding | MiniMax-M3 | 1M context + native multimodal + browsing |
| Mechanical work (JSON edits, file ops, renders) | GLM-5.2 (glm profile) | Cheaper than deepseek, reliable for mechanical |
| Deep research, web synthesis | MiniMax-M3 or research profile | Strong browsing + long context |
| Heavy reasoning, adversarial review, complex analysis | Grok 4.5 (grok45) | Opus-class flagship, 500K context, strong coding |
| Routine analysis, quick second opinion, cheaper tasks | Grok 4.3 (grok43) | Faster/cheaper than 4.5, good for routine work |
| Quick question, light planning | DeepSeek-V4-Pro (yourself) | Don't spawn for trivial tasks |
Pitfalls
- Don't over-delegate. Simple reads, single commands, quick lookups — do them yourself. Spawning a sub-agent for a 2-second task burns more tokens than just doing it.
- Don't dump conversation history. Sub-agents get a clean brief. Only include what they need.
- Mechanical work goes to glm, not kimi-c. User directive 2026-07-22: "stop using ask kimi c for tasks, use glm instead." This applies to LTX workflow building, JSON editing, render submission, file operations.
- Claude (ask-claude) is for adversarial validation, not general work. Use Claude to validate plans, review architecture, diagnose bugs — not for mechanical tasks or implementation.
- Background mode for long dispatches. Any dispatch expected to take >5 minutes must use
terminal(background=true, notify_on_complete=true)to avoid the 600s foreground timeout.