Join the conversation
Join the community of Machine Learners and AI enthusiasts.
Sign UpThe per-harness split holds up. The step-1,000 sizes mostly don't.
I took the six complete checkpoints from 500 to 1,000 in the Space's training-results.json and compared the two LFM runs under each harness.
Multi-harness is ahead under Claude Code and Codex at all 6. OpenCode-only is ahead under OpenCode at 5 of 6. Mini-SWE-Agent is a tie (mean gap +0.2).
But the plateau gaps are about 5 points each: OpenCode 4.9, Claude Code 5.7, Codex 5.3.
Step 1,000 shows 8.4, 6.8 and 10.4. The Codex gap there is twice its plateau mean.
Checkpoint-to-checkpoint spread per harness is 1.1 to 3.4 points.
One attempt on 250 cells gives about 3.0 to 3.2 from sampling alone.
So after step 500 each per-harness curve is flat plus noise, and those six checkpoints are six free repeats.
Same in the savings heatmap. OpenCode-only under Claude Code uses 9.6% more calls at step 1,000, but 5.9 to 8.6% fewer at steps 600 through 900. Generated tokens are the consistent part: above base at 5 of 6.
Multi-harness tool savings is the one number still moving: 23.7% at step 500, 31.1% at 1,000, calls 6.91 to 5.91, while overall accuracy sits flat inside noise.
That fits your all-correct-groups argument. Once most groups are all correct, the bonus is the only gradient left.
Do you still have the per-cell outcomes at each checkpoint? Pooling 500 to 1,000 gives every task six draws per harness, enough for a paired test of the crossover.