hudiege.cn
Preprint · Under review at TMLR (#12935)预印本 · TMLR 在审(#12935)

Stopping as a Database Property: Reason-Bearing Termination for Long-Running LLM Agents

Baofeng Zhao · Independent Researcher · baofeng@hudiege.cn
DOI: 10.5281/zenodo.23144013 2026-10-05 CC BY 4.0

论文正文(摘要、数字、图注)为英文原文;中文版 PDF 见上方「中文版」链接。

Abstract摘要

Loop termination should be a property of the data layer: when the substrate owns resource and intent state, stopping becomes a reason-bearing, auditable predicate rather than an arbitrary step cap or an unreliable model decision. We document this third option, observed rather than invented, in a self-built agent research testbed. The substrate's signal ledger holds 48,061 self-monitoring events over 45 days; loop-related entries reference round counts up to 212 (208 distinct values), with controlled convergence at 29 rounds and substrate fidelity of 30/40 across a fixed 60-round probe. On the official ARC-AGI-3 benchmark: 28.57 (3/6 levels, 31 steps) and 0.15 on Kaggle, operating under a self-imposed 82-step budget. A self-narration layer grew from zero to 12,102 entries (10,393 commit-verified at the September audit; the remainder from a live read eight days later). We crystallize the mechanisms into four data-computable termination predicates (one implemented, three specified), and document the boundary the same record forces: convergence control and converged-to content are separable, and a relay loop faithfully amplifies upstream wiring faults. No cross-system superiority is claimed; thousand-round stress tests have not been run.

Key numbers关键数字

48,061
self-monitoring events in the substrate signal ledger
29 rounds
natural convergence of the relay loop, before any predicate
30/40
substrate fidelity on the frozen drift probe
28.57
ARC-AGI-3 official score (3/6 levels, 82-step budget)
Termination control has three homes; only one makes the stop
Termination control has three homes; only one makes the stop reason a queryable property of system state.

The honesty bound诚实边界

Boundary边界

Three case files show a healthy relay loop—correct convergence behavior included—amplifying upstream wiring faults into systematic self-misunderstanding. Convergence control and content fidelity are separable; a relay loop's honesty is capped by the honesty of its wiring.

Post-run substrate fidelity by question type, against per-ty
Post-run substrate fidelity by question type, against per-type chance.