Files
hermes-skills/media/ltx-video-pipeline/references/artifact-review-2026-07-21.md
T

47 lines
2.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# LTX-2.3 Artifact Root Causes — 6-Agent Consensus (2026-07-21)
Full research result from 6 independent agents diagnosing Story 1 "Ghost in the Wire" artifacts.
## Methodology
6 agents reviewed the exact technical setup independently:
- Claude Opus 4.8 (SSH print mode, 9 turns)
- Kimi K2.7 Code (kimi-c profile, 30 turns)
- Kimi K2.6 (kimi profile, 30 turns)
- MiniMax M3 (minimax profile, 30 turns)
- GLM-5.2 (glm profile, 30 turns)
- Deep research (research profile, 600 turns, 18 sources including Reddit, GitHub, HuggingFace)
## Consensus Findings
### CRITICAL (6/6 agents agree)
1. **LoRA stacking at 1.0+1.0** — Distilled + ID LoRA both at full strength cause interference on faces/hands. Fix: Distilled → 0.7, ID → 0.6 (or drop ID if no audio).
### HIGH (4-5/6 agents agree)
2. **10s exceeds temporal coherence** — LTX-2.3 degrades after 5-6s. Fix: test at 5s (121 frames).
3. **linear_quadratic scheduler** — Distilled model is fragile with non-standard schedules. Fix: switch to `simple`.
4. **Inconsistent guide_strength** — 0.5 on scene 1, 0.7 on others breaks flow. Fix: 0.5 uniform.
### MEDIUM (2-3/6 agents agree)
5. **Audio track enabled with no audio** — AudioVAE NaN risk. Fix: disable audioTrackEnabled.
6. **Text in prompts** — LTX can't render readable text. Fix: replace with abstract descriptions.
7. **512×512 square aspect** — LTX prefers widescreen. Fix: 768×512 (requires new frames).
8. **fp8_scaled v1.0 quality** — Kijai's v1.1 is better. Fix: model upgrade.
## Deep Research Sources (18 total)
Key community sources:
- GitHub #148: End-of-video logo/watermark artifacts (upscaler v1.0 contamination)
- HF Discussion #13: Upscaler sigma fix table by frame count
- Reddit r/StableDiffusion: Skin compression fix (Nearest Exact interpolation)
- Reddit r/StableDiffusion: Deformed bodies / identity drift in I2V
- Reddit r/comfyui: Official workflow vs ComfyUI built-in (prompt ignoring bug)
- Reddit r/StableDiffusion: RL LoRA for coherence (OmniNFT)
- HuggingFace RuneXX: Dev model sampler comparison
- HuggingFace LiconStudio: MSR V2 LoRA for identity preservation
- GitHub #255: Image conditioning does not preserve identity
- GitHub #244: Noise in distilled 1.1 model
- LTX.io blog: Official artifact reduction guide
Full research: `~/workspace/research/results/2026-07-21-ltx-2.3-artifacts.md`