Files
hermes-skills/research-dispatch-pitfalls/SKILL.md
T

263 lines
14 KiB
Markdown

---
name: research-dispatch-pitfalls
description: Dispatch pitfalls for deep-research and better-search — foreground pipe kills, missing result files, session_id format variance, -Q session DB gaps. Load alongside either dispatcher skill.
version: 1.1.0
author: Hermes Agent
metadata:
hermes:
tags: [research, dispatch, pitfalls, deep-research, better-search]
related_skills: [deep-research, better-search]
---
# research-dispatch-pitfalls — Dispatch Hard-Learned Lessons
Companion to `deep-research` and `better-search`. Load this alongside either
dispatcher when dispatching research. These are the things that went wrong in
real dispatches and how to avoid them.
## Pitfall 1: Foreground pipe through `head` kills the process
**Status:** Fixed in `deep-research` v2.5.0 and `better-search` v1.0.1. Both dispatchers now use `terminal(background=true, notify_on_complete=true)` as the primary dispatch pattern. No pipe, no SIGPIPE.
**If it still happens** (e.g., you ran foreground manually): the pipe closes as soon as `head` exits after N lines. SIGPIPE fires, hermes dies. Re-dispatch as background. The truncated ledger from the killed run is irrelevant — start fresh.
## Pitfall 2: Result file not written despite completed research
**Status:** Fixed in `deep-web-research` v2.2.0. Move 5 now has an explicit `mkdir -p` + write to `/home/n8n/workspace/research/results/<date>-<slug>.md` with YAML frontmatter. The /tmp confinement rule has an explicit exception for the results directory.
**If it still happens:** after `notify_on_complete` fires, verify the file exists. If missing, extract the report from the process log and save it manually. Do not leave the research stranded in process output only.
## Pitfall 3: Session_id format varies
**Symptom:** One dispatch emits `session_id: research-1784565747`, another
emits `session_id: 20260720_115450_cdb929`. A regex hardcoded for one format
misses the other.
**Root cause:** The research agent uses different session_id formats depending
on the run context. Both are valid.
**Fix:** Capture whatever follows `session_id: ` — don't hardcode a format.
Use `grep -oE 'session_id: \S+'` or just read the line after the marker.
Both formats work for `--resume`.
## Pitfall 4: -Q sessions are NOT in the session DB
**Symptom:** `session_search` on the research profile returns zero results for
a `-Q` (quiet mode) dispatch, even though the session completed successfully.
**Root cause:** Quiet mode (`-Q`) suppresses session storage in the SQLite DB.
The session ran but left no DB record.
**Fix:** The session_id from a `-Q` dispatch is for the operator's reference
and for `--resume` only. Do not expect `session_search` to find it. To check
whether a `-Q` session completed, check for the result file or the process
exit status instead.
## Pitfall 5: First dispatch killed → re-dispatch, don't give up
**Symptom:** First dispatch died (pipe kill, timeout, etc.). Ledger has only
the strategy section. No result file.
**Fix:** Re-dispatch the same question as a background process. The research
agent starts fresh — the truncated ledger from the killed run is irrelevant.
Don't try to resume a killed session; start a new one. The second dispatch
will run the full flow independently.
## Pitfall 6: Shell metacharacters in `-q` break the dispatch
**Symptom:** `hermes -p <profile> chat -q "<long prompt with special chars>"`
fails with `hermes: error: unrecognized arguments`. The shell interprets
backticks, quotes, parentheses, and dollar signs before hermes sees them.
Prompt fragments get parsed as CLI flags. Exit code 2.
**Root cause:** The `-q` argument is passed through the shell. Any
shell-special character in the prompt body breaks the argument boundary.
**Fix:** Write the prompt to a temp file first, then dispatch with a short
`-q` that tells the peer to read the file:
```bash
write_file("/tmp/peer-task.txt", content=full_prompt)
terminal("hermes -p dev chat -q 'Read /tmp/peer-task.txt and execute the task.' -Q --max-turns 30 --yolo",
background=true, notify_on_complete=true)
```
This avoids all shell escaping issues. Use for any prompt longer than ~3
lines or containing backticks, quotes, parentheses, or dollar signs.
Real failures (July 2026): ask-dev round 5 crashed when a validation prompt
containing backtick-quoted filenames and parenthetical notes broke the `-q`
argument. Same pattern recurred with ask-kimi-c dispatch.
## Pitfall 7: Don't bail on software without due diligence
**Symptom:** A tool, model, or package appears unavailable from one source.
You pivot to an alternative without exhausting other sources.
**Root cause:** Taking a single "not found" signal as definitive. Software
distributes through multiple channels (HuggingFace, ModelScope, GitHub,
community mirrors). One being empty doesn't mean the software doesn't exist.
**Fix:** Before declaring something unavailable, check ALL distribution
channels:
- HuggingFace API: `curl -s "https://huggingface.co/api/models?search=..."`
- GitHub: direct URL probe (`curl -sI https://github.com/org/repo`)
- ModelScope: direct URL probe
- Community mirrors: search for GGUF quants, forks, mirrors
- Direct download URLs from official docs
Only after exhausting all channels should you pivot. Document each channel
checked and its result.
Real failure (July 2026): Wan 2.7 was declared unavailable after checking
ModelScope (empty placeholder). The user called this out as insufficient
due diligence. Subsequent exhaustive search confirmed: zero results on
HuggingFace API, GitHub repo 404, ModelScope placeholder — the conclusion
was correct, but the process was sloppy. The user's correction stands:
exhaust the channels before pivoting.
## Pitfall 8: Don't ask a peer to "build everything" — coordinate piece by piece
**Symptom:** You dispatch a peer with "build the whole pipeline" and it either
times out, hits the turn ceiling, or produces a half-finished result you
can't verify.
**Root cause:** Long build tasks exceed foreground timeouts, exhaust turn
budgets, and produce unverifiable self-reports. The peer has no checkpoints
and you have no visibility into intermediate state.
**Fix:** Break the build into numbered pieces. Dispatch one piece at a time.
Verify each piece's output before dispatching the next. The coordinator
(you) owns the sequence; the peer owns each piece's execution.
Pattern:
```
Piece 1: System deps (apt-get, ffmpeg, git-lfs) — verify each binary
Piece 2: Runtime install (ComfyUI, venv, PyTorch) — verify launch
Piece 3: Custom nodes (git clones) — verify imports
Piece 4: Small model smoke test — verify generation
Piece 5: Full model download — verify file sizes
Piece 6: Full render test — verify output MP4
Piece 7: Cleanup and report
```
Each piece gets its own temp file with exact commands and verify steps.
The peer reads the file, executes ONLY that piece, reports results, and
stops. You verify, then dispatch the next piece.
**Why this works:**
- Each piece fits in a foreground timeout
- Each piece has a verifiable output (binary version, file size, exit code)
- A failed piece doesn't waste the work of prior pieces
- You can resume from the last successful piece
- The peer can't drift into unrelated work
Real application (July 2026): AI video pipeline build on 10.0.0.175.
Piece 1 (system deps) dispatched first because it fits in the current
15GB free disk space. Pieces 2-3 also fit. Piece 4+ need the disk
increase. The coordinator tracks which pieces are done and which are
blocked.
**When to use:** Any build task that spans multiple install steps, model\ndownloads, or verification stages. Especially when disk space, timeouts,\nor turn budgets are constraints.\n\n## Pitfall 9: Don't make the peer monitor long-running tasks — poll directly from the coordinator\n\n**Symptom:** You dispatch a peer to run a long task (render, download, build).\nThe peer's polling loop consumes its turn budget. It hits the turn limit\nmid-task and returns incomplete results. You have no output and no prompt ID.\n\n**Root cause:** The peer's `--max-turns` budget (even at 30) is consumed by\nthe polling loop. Each `sleep 10; curl status` is a turn. A 40-minute render\nat 10-second polls burns 240 turns — far beyond any reasonable budget.\n\n**Fix:** The peer submits the job and returns the job ID. The coordinator\nthen runs a background shell script (`terminal(background=true,\nnotify_on_complete=true)`) that polls status and VRAM directly on the\ntarget host. The peer's job is to get the job submitted and validated —\nthe coordinator owns the wait.\n\nPattern:\n```\n# Peer submits the job, returns prompt_id, then STOPS\n# Coordinator polls directly:\nterminal(\"ssh target 'while true; do curl -s http://localhost:8188/history/$ID | ...; sleep 10; done'\",\n background=true, notify_on_complete=true)\n```\n\n**Why this works:**\n- The peer stays within its turn budget (submit + validate = ~5 turns)\n- The coordinator's background script has no turn limit\n- The coordinator gets notified on completion\n- If the task fails, the coordinator has the full log\n\nReal failure (July 2026): kimi-c hit 30-turn limit mid-render on a Wan 2.2\nTI2V-5B 60-second generation. The coordinator polled directly and got the\nresult 40 minutes later. Same pattern recurred on a 5-minute render attempt.
## Pitfall 10: Render submitted but never started — diagnose the hang, don't just wait
**Symptom:** The peer submits a render job, gets a prompt_id, starts polling.
VRAM stays at idle (~1490 MiB) for 9+ minutes. The polling loop burns turns
waiting for a render that will never complete.
**Root cause:** The job errored immediately on submission (missing package,
wrong node input, model not found) but the error is only visible in the
history endpoint, not in the submit response. The submit returned
`{"prompt_id": "...", "node_errors": {}}` — empty node_errors, so it looked
like success. The actual error was in the execution traceback inside the
history object.
**Fix — three-step diagnosis when VRAM stays at idle:**
1. `curl -s http://localhost:8188/queue` — if `queue_running` is empty, the
job errored immediately
2. `curl -s http://localhost:8188/history/$PROMPT_ID` — check
`status.status_str` for `error` and inspect `messages` for the traceback
3. `tail -50 /tmp/comfyui.log | grep -i error` — the ImportError or node
error will be there
**Common causes:**
- Missing Python package (SageAttention installed but ComfyUI not restarted)
- Wrong node input field name (e.g., `lora_name` instead of `lora`)
- Model file not found at the specified path
- Channel mismatch (wrong VAE, wrong model type)
**Pattern:** When the peer reports "VRAM stuck at idle, render not starting,"
do NOT wait longer. Diagnose immediately with the three checks above. The
render will never start — it already failed.
Real failure (July 2026): kimi-c submitted an optimized Lightning LoRA
workflow. SageAttention was installed but ComfyUI hadn't been restarted.
The submit returned success, but the render errored with `ImportError:
Selected attention mode not available`. VRAM stayed at 1490 MiB for 9+
minutes while the peer polled. The coordinator diagnosed the hang, restarted
ComfyUI, and re-submitted.
## Pitfall 11: Peer hits turn limit on long render — coordinator polls directly
**Symptom:** The peer submits a render job, starts a polling loop, and hits
the turn limit before the render completes. You get a partial report with
no output file and no final timing.
**Root cause:** The peer's `--max-turns` budget is consumed by the polling
loop. Each `sleep 10; curl status` is a turn. A 40-minute render at 10-second
polls burns 240 turns.
**Fix:** The peer submits the job and returns the prompt_id. The coordinator
then runs a background shell script that polls status and VRAM directly on
the target host. The peer's job is to get the job submitted — the coordinator
owns the wait.
```bash
# Peer submits, returns prompt_id, then STOPS
# Coordinator polls directly:
terminal("ssh target 'while true; do curl -s http://localhost:8188/history/$ID | ...; sleep 10; done'",
background=true, notify_on_complete=true)
```
Real failure (July 2026): kimi-c hit 30-turn limit mid-render on a Wan 2.2
TI2V-5B 60-second generation. The coordinator polled directly and got the
result 40 minutes later. Same pattern recurred on a 5-minute render attempt
and on the Lightning LoRA optimized render (Piece 8c fixes).
## Pitfall 12: Research agent exits after Move 0 — no result file, no findings
**Symptom:** The background process completes (exit 0, notify_on_complete
fires), but the result file at `/home/n8n/workspace/research/results/<date>-<slug>.md`
does not exist. The process log shows the agent got through Move 0 (strategy
phase) and then stopped. The ledger at `/tmp/research-<date>-<slug>.md` exists
but contains only the strategy section — no findings, no condensation.
**Root cause:** The research agent hit an error during Move 1 (landscape pass)
or Move 2 (deep-dive) — likely a tool failure (SearXNG returning empty, web
extract timeout, or a Python error in the research script). The agent exited
without writing the result file. The `-Q` flag means no session DB record.
**Fix:** When the result file is missing after completion:
1. Check the process log for the session_id (last ~20 lines)
2. Check if the ledger exists: `ls -la /tmp/research-<date>-<slug>.md`
3. If the ledger has only the strategy section, the research didn't progress
past Move 0 — re-dispatch with a fresh session
4. If the ledger has findings but no result file, the condensation step
(Move 5) failed — resume the session with `--resume <session_id>` and
ask it to condense from the existing ledger
5. Do NOT treat a Move-0-only exit as a completed research — it produced
nothing useful
**Prevention:** When dispatching deep research, always note the expected
result path and the stem (`/tmp/research-<date>-<slug>`) so you can
diagnose failures quickly.
Real failure (July 2026): Deep research on LTX-2.3 artifacts dispatched
with 10 specific research questions. Process completed but result file
never created. Process log showed the agent got through Move 0 (strategy)
and stopped. Ledger at `/tmp/research-2026-07-22-ltx-2.3-artifacts-deep`
had only the strategy section. Session `20260722_092743_bee18c` on the
research profile had no DB record (quiet mode). Root cause not determined
— likely a tool failure during Move 1 landscape pass.