Judgment · rulings · merges
Claude
Best at careful judgment and clean merges, so it held the rulings and the only merge key. It also ruled from memory once, and the ruling had to be withdrawn.
Rule nowCite the source before ruling. One model merges; the others propose.
Planning · long runs
ChatGPT + Codex
A tireless planner and runner. Its status reports understated problems three times, and a five-agent run burned 98% of a 5-hour window in one night.
Rule nowEvidence over status. One session per lane, and every run stops on the meter.
Testing · research · documents
Gemini + Antigravity
A strong second opinion. It ran the smoke tests on finished pieces, and handled research and longer documents in long solo runs.
Rule nowDon't let the builder test its own work. Give testing to a different model.
Wide research, low cost
DeepSeek
Fast and cheap for breadth: market scans, source checks and research briefs. It ran a full shift until it hit the meter line.
Rule nowSpend expensive models on judgment and cheap ones on breadth. Research arrives as a brief; another agent commits it.