When Should the Agent Speak? A Survey of Intervention Timing for Always-On AI Assistants

Tao An

Video

Paper PDF

Thumbnail of paper pages

Abstract

Agents now complete multi-step tasks against verifiable goals, and egocentric perception resolves user intent with high reported accuracy. What remains unmodeled is the decision of when to act unprompted. We survey intervention timing for always-on assistants (smart glasses, MR headsets, ambient copilots) around a single decision rule, intervene iff the expected benefit of acting exceeds the expected cost of interrupting, and organize the literature into five layers: signals, decision, action, memory, and evaluation. We reconnect two lineages that have proceeded almost without citing each other: the 1999-2017 interruptibility literature, which formalized interruption cost rigorously but had no capable actor, and the 2024-2026 proactive-agent wave, which has actors but rediscovers the cost term only in fragments (false-alarm pricing, cognitive load, social violation, compute) that no work unifies. We argue that evaluation is the gating layer. Mobile health already learns intervention timing online by randomizing decision points against a validated proximal outcome; open-world assistance has no analogous outcome, so a reinforcement-learning objective for proactivity has no shared, validated reward to train against. We therefore propose the design of a benchmark for open-world intervention timing with an explicit cost term, machine-score its metric suite over 131.5 hours of seeded labels to show it separates the reference policies, and assemble a reference architecture for the priced decision from surveyed components. Six open problems close the survey.