Prompt Priming Masquerading as Visual Attribution

Mohammed Faisal Parvez · Hyderabad, IN

Abstract
An apparent visual-misattribution effect in detector-grounded VLM benchmarks.
Method
Marker × prompt ablation isolating the visual channel from the text channel.
Key Result
The effect was prompt-text priming — wording shifted results 9.4pp; visual markers ≤0.2pp.
Implication
Benchmark conclusions can flip when input channels aren't isolated.
Fig. 2 — Effect size by input channel
prompt text 9.4pp visual marker ≤0.2pp shift in benchmark outcome (percentage points)