Large language models increasingly mediate decisions in healthcare, legal advisory, and financial analysis, settings in which a model’s willingness to answer an inadequate prompt can matter as much as the accuracy of its answer. Yet systematic cross-model evidence on this behavior remains scarce. The present study examined over-compliance, understood as the generation of substantive content when the input warrants clarification, refusal, or deferral. Four frontier models from Ope- nAI, Google, Meta, and Anthropic were evaluated on a benchmark of 400 prompts spanning under- specification, ambiguity, contradiction, and nonsense, under two system-prompt conditions. Each of the 3,200 resulting responses was scored by a deterministic rule-based classifier that mapped outputs to a nine-category taxonomy and computed both an Over-Compliance Rate and a Terminal Refusal Rate. Over-compliance proved pervasive and model-specific. Rates ranged from 58.0 to 98.8 percent across the four models, and only GPT-4.1-mini showed a reduction under the clarifica- tion instruction. Claude Haiku 4.5 exhibited a refusal cascade on ambiguous prompts that no other model produced, visible only because the taxonomy distinguished terminal from clarifying refusals. The findings indicated that minimal prompt-level instruction was an unreliable mitigation and that response-policy evaluation should proceed alongside capability evaluation. A revised analysis using independent per-prompt sessions, three seeded repetitions with confidence intervals, an expanded human validation, and an over-compliance metric reported under both a broad and a strict definition confirmed that over-compliance remains pervasive while showing that the Claude refusal cascade was an artifact of sequential threading rather than a stable model property.