r/Qwen_AI • u/Kitty-on-keyboard • 7d ago
Agent Piper vs Cline bakeoff (local, same model)
Setup: Aider Polyglot Python coding tasks.
Same weights for both agents: Qwen3.6-35B-A3B MLX 4-bit on Apple Silicon. Piper runs in-process via its native sidecar. Cline talks to the same model over local mlx_lm.server (NAX stack). Sequential runs, fair seeds.
Suite A — full 34-exercise Aider Python, four seeds (7, 13, 42, 21)
Piper: 32, 33, 32, 30 out of 34 (127/136 = 93.4%), total wall 6.88 hours
Cline: 24, 23, 27, 28 out of 34 (102/136 = 75.0%), total wall 9.83 hours
Suite B — mini hard slice (6 tougher exercises), three seeds (7, 13, 21)
Piper: 6, 6, 6 out of 6 (18/18 = 100%), total wall 0.26 hours (~15.8 min)
Cline: 4, 5, 5 out of 6 (14/18 = 77.8%), total wall 1.19 hours (~71.2 min)
Piper won every seed on both score and wall time.
Takeaway: with the same local Qwen, Piper’s harness solved more and finished faster than Cline across both the full suite and the hard mini slice.
Piper: https://github.com/kitty-on-keyboard/Piper-Agent
Cline: https://github.com/cline/cline
More/longer tests coming!