Selective Refusal as Protective Calibration

Mohammed Faisal Parvez · Hyderabad, IN

Abstract
Which VLM should a safety-critical robot trust when it's uncertain?
Method
Six-pipeline benchmark — Qwen2.5-VL 3B / 7B / 72B + InternVL2-2B — on Modal A100s over robot perception frames.
Key Result
Failure modes are model-family-specific: one family refuses under uncertainty, another confabulates on identical inputs. The refusing family reaches ~2× precision and 5–10× recall on engaged frames.
Implication
Refusal is calibration, not weakness — pick model families by failure mode, not leaderboard rank.
Fig. 1 — Precision / recall vs confabulating family
precision ~2× recall 5–10× engaged frames · refusing family vs baseline