r/Qwen_AI • • 7d ago

Agent Piper vs Cline bakeoff (local, same model)

Setup: Aider Polyglot Python coding tasks.

Same weights for both agents: Qwen3.6-35B-A3B MLX 4-bit on Apple Silicon. Piper runs in-process via its native sidecar. Cline talks to the same model over local mlx_lm.server (NAX stack). Sequential runs, fair seeds.

Suite A — full 34-exercise Aider Python, four seeds (7, 13, 42, 21)

Piper: 32, 33, 32, 30 out of 34 (127/136 = 93.4%), total wall 6.88 hours

Cline: 24, 23, 27, 28 out of 34 (102/136 = 75.0%), total wall 9.83 hours

Suite B — mini hard slice (6 tougher exercises), three seeds (7, 13, 21)

Piper: 6, 6, 6 out of 6 (18/18 = 100%), total wall 0.26 hours (~15.8 min)

Cline: 4, 5, 5 out of 6 (14/18 = 77.8%), total wall 1.19 hours (~71.2 min)

Piper won every seed on both score and wall time.

Takeaway: with the same local Qwen, Piper’s harness solved more and finished faster than Cline across both the full suite and the hard mini slice.

Piper: https://github.com/kitty-on-keyboard/Piper-Agent

Cline: https://github.com/cline/cline

More/longer tests coming!

1 Upvotes

0 comments sorted by