KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill Paper • 2607.12625 • Published 8 days ago • 59
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 7 days ago • 96
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 14 days ago • 74
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 9 days ago • 213
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training Paper • 2607.05804 • Published 16 days ago • 18
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation Paper • 2607.05147 • Published 17 days ago • 37
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published 20 days ago • 80
Agentic Abstention: Do Agents Know When to Stop Instead of Act? Paper • 2606.28733 • Published 26 days ago • 148
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published 27 days ago • 111
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents Paper • 2606.28480 • Published 27 days ago • 48
Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It Paper • 2606.26027 • Published 29 days ago • 18
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents Paper • 2606.24551 • Published Jun 22 • 28
The Verification Horizon: No Silver Bullet for Coding Agent Rewards Paper • 2606.26300 • Published 29 days ago • 51
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published 28 days ago • 56
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 95
OpenThoughts-Agent: Data Recipes for Agentic Models Paper • 2606.24855 • Published about 1 month ago • 47
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published about 1 month ago • 151