NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Abstract
NeoHorse-1 uses agentic post-training with intelligent routing, structured feedback loops, and curriculum-based distillation to improve model capabilities across agent benchmarks.
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model (2026)
- RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement (2026)
- Agentic Routing: The Harness-Native Data Flywheel (2026)
- AREX: Towards a Recursively Self-Improving Agent for Deep Research (2026)
- UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations (2026)
- LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (2026)
- Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.08183 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 10
TokenRhythm/NeoHorse-1-9B
Datasets citing this paper 0
No dataset linking this paper