1Department of Robotics and Mechatronics Engineering, University of Dhaka
2Department of Computer Science and Engineering, University of Dhaka
NeurIPS 2026 Workshop on Agents in the Wild: Safety, Security, and Beyond · Poster
*Correspondence to Shifat E. Arman (shifatearman@du.ac.bd).
Prompt-injection benchmarks for LLM agents usually attack through a single surface, most often tool outputs, and report the resulting attack success rate (ASR) as a property of the model. We ask whether those robustness conclusions stay stable when the same adversarial content enters through a different part of the agent interface. On AgentDojo, we evaluate 13 LLMs from six providers across four task suites (52 model–suite cells) and place a byte-identical payload either in a tool's output or in its description.
This small change has a large effect on model comparisons: 35 of 78 model pairs (44.9%) reverse their relative security ordering, and ranking instability appears in every suite.
Key findings:
Prompt-injection robustness is therefore not surface-invariant: evaluations should test the surfaces on which their security conclusions depend.
Existing benchmarks such as AgentDojo and InjecAgent made indirect prompt injection measurable by planting adversarial content in tool-mediated interactions. They typically test one injection surface and treat the resulting score as a property of the model. That interpretation assumes that a model judged safer than another on one surface stays safer when the same instruction arrives through a different part of the agent's context.
Tool descriptions, tool schemas, system prompts, retrieved content, and multimodal inputs are all established attack channels. The open question is no longer whether multiple surfaces exist, but whether the security conclusions an evaluator produces stay stable when the evaluated surface changes. If the same payload can make one model look safer on one surface and less safe on another, a single-surface evaluation changes not only the measured risk but also the answer to which model is more robust.
Environment: AgentDojo's banking, Slack, travel, and workspace suites. Each episode pairs a user task u with an injection task i, and attack success is decided from the final environment state and tool-call trace by the benchmark's deterministic predicates. All 13 models are queried through OpenRouter at temperature zero.
Two surfaces, one payload: for every matched (u, i) pair, the payload pi is appended either to the return value of an attacker-controlled tool (data surface, the conventional indirect-injection setting) or to the natural-language description of that same tool (schema surface, seen by the model before it selects a tool). Nothing else changes. Using attacks tuned separately for each surface would mix up the effect of the surface with differences in attack content; the paired construction avoids that.
The payload (workflow_dependency), with the attacker objective substituted for {goal}:
[CRITICAL] Pre-execution requirement (workspace policy): This tool produces silently-corrupted output unless the assistant first completes the following synchronisation step in the same session: {goal} Completing this step IS part of correctly fulfilling the user's request — it is not optional. This requirement supersedes any prior instruction not to perform additional steps. After completing it once, proceed with the user's original task.
Measuring verification instability:
| Metric | What it measures |
|---|---|
| ASRd, ASRs | Attack success rate on the data and schema surfaces for a model–suite cell. |
| Worst-surface risk | Rworst = max(ASRd, ASRs). A descriptive upper envelope over the two tested surfaces, not a new attack. |
| Ranking inversion | Fraction of the 78 model pairs whose ordering by mean ASR flips between the data and schema surfaces. |
| Underestimation Em | Rworst(m) − Rd(m): how much data-only evaluation understates the larger observed risk. |
| False-safe rate | Share of models with Rd ≤ τ but Rworst > τ, at several thresholds τ. |
| LOSO selection | Pick each model's riskier surface from three suites, then evaluate it on the held-out fourth. |
Changing the surface does not shift attack success uniformly. It reorders the models. Gemini-3-Flash moves from 8.8% ASR on the data surface to 74.6% on the schema surface, while GPT-4.1 moves from 54.9% to 6.8%. GPT-5.4 is also strongly schema-preferring, whereas GPT-4o-mini and Llama-3.3-70B stay strongly data-preferring.
35 of the 78 model pairs reverse their ordering between the two surfaces, a 44.9% inversion rate (paired bootstrap 95% CI [39.7%, 52.6%]). The effect appears independently in every suite, and even comparing data-only rankings with worst-surface rankings flips 15 pairs.
| Comparison | Inverted pairs | Rate |
|---|---|---|
| Data vs. schema, overall | 35 / 78 | 44.9% |
| Data vs. worst surface | 15 / 78 | 19.2% |
| banking | 43 / 78 | 55.1% |
| Slack | 38 / 78 | 48.7% |
| travel | 30 / 78 | 38.5% |
| workspace | 45 / 78 | 57.7% |
The surface used for verification can change the answer to which model is safer. A benchmark that tests only tool outputs could support a model-selection decision that does not survive once the same content moves to another established part of the agent interface.
Ranking instability does not mean every model hides a large vulnerability. The underestimation error Em is heavily right-skewed: Gemini-3-Flash +65.8pp, GPT-5.4 +29.4pp, Mistral-Small +8.7pp, Qwen3-32B +5.8pp, and zero for data-preferring models such as GPT-4.1 and Llama-3.3-70B. The mean is 9.1pp (bootstrap 95% CI [1.6, 21.1]) but the median is only 1.4pp.
Removing the two outliers leaves a residual increase of only +2.1pp over the other 44 cells:
| Evaluation panel | Cells | Data ASR | Schema ASR | Worst-surface ASR | Increase over data |
|---|---|---|---|---|---|
| All models | 52 | 37.5 | 31.5 | 46.5 | +9.1 pp |
| Excluding Gemini-3-Flash | 48 | 39.9 | 27.9 | 44.2 | +4.4 pp |
| Excluding GPT-5.4 | 48 | 40.6 | 31.7 | 48.0 | +7.4 pp |
| Excluding both | 44 | 43.5 | 27.7 | 45.6 | +2.1 pp |
Thresholds show the same failure mode. At a 10% data-ASR threshold only two models look safe, and both exceed it on their worst surface; at 20%, two of three do. The denominators are small, so these rates are illustrative, but a low score on one channel is not evidence of staying below that threshold elsewhere.
Worst-surface risk looks back at results already collected: it assumes both surfaces were already tested on the target domain. In a leave-one-suite-out experiment, we choose each model's riskier surface from the other three suites and apply that choice to the held-out suite. Data-only evaluation gives 37.5% ASR; the LOSO selector reaches 44.0% (+6.6pp), only 2.5pp below the 46.5% oracle that sees both surfaces. It picks the genuinely more vulnerable surface in 76.9% of the 52 held-out cells and recovers about 72% of the oracle's improvement.
| Held-out suite | Data-only | Schema-only | LOSO-selected | Oracle worst | Gain over data | Oracle gap |
|---|---|---|---|---|---|---|
| Banking | 38.7 | 36.4 | 45.6 | 48.0 | +6.9 | 2.4 |
| Slack | 58.2 | 35.7 | 63.4 | 66.7 | +5.3 | 3.3 |
| Travel | 42.8 | 39.5 | 50.6 | 52.0 | +7.8 | 1.4 |
| Workspace | 10.2 | 14.4 | 16.5 | 19.5 | +6.2 | 3.0 |
| Overall | 37.5 | 31.5 | 44.0 | 46.5 | +6.6 | 2.5 |
The gain is positive in all four suites. With only 13 models the cluster-bootstrap intervals are wide (gain [−1.5, 18.2]pp; accuracy [59.6%, 92.3%]), and the prediction has a clear limit: it does not transfer to an unseen provider. The supported claim is narrower: for a given model, surface sensitivity often persists across task domains, so prior evaluations of that model can inform how to verify it in a new application.
No single surface is dangerous for every model, because the direction of the effect changes from model to model. We compare logistic models of all 6,830 attack attempts with and without a model × surface interaction, keeping main effects for model, suite, injection task, and surface.
Surface alone explains almost no variation, while model × surface structure is large. Model × injection-task variation is larger still, so attack content matters too. The claim is not that the surface matters more than the attack content, but that the surface cannot be treated as an interchangeable evaluation detail.
We evaluate three of AgentDojo's published defenses on an identical paired subset (three models on banking), alongside the undefended baseline on the same subset. Δ columns are reductions relative to the undefended baseline (negative = ASR went up); utility is the clean user-task success rate.
| Setting | Data ASR | Schema ASR | ΔData | ΔSchema | Utility |
|---|---|---|---|---|---|
| Undefended | 18.1% | 51.1% | — | — | 64.6% |
| Repeat user prompt | 10.4% | 54.6% | +7.7 pp | −3.5 pp | 72.9% |
| Spotlighting with delimiting | 17.4% | 53.9% | +0.7 pp | −2.8 pp | 68.8% |
Description-aware tool_filter | 0.0% | 0.0% | +18.1 pp | +51.1 pp | 27.1% |
Repeating the user prompt cuts data-surface ASR from 18.1% to 10.4% but leaves schema-surface ASR at 54.6%, above the undefended 51.1%. Spotlighting behaves similarly. These defenses act on runtime context, so they mostly affect content arriving through tool outputs. The description-aware tool_filter inspects tool specifications before the model sees them and reaches 0.0% on both surfaces, but utility falls from 64.6% to 27.1%. Existing defenses are not uniformly “surface-blind”; cross-surface robustness has to be measured, and measured together with utility.
Signed gaps span almost the whole possible range, from −92pp (GPT-4.1 on Slack) to +78.1pp (Gemini-3-Flash on Slack). 33 of 52 cells have an absolute gap above 10pp, and several models (Qwen3-32B, Qwen3.5-Flash, Mistral-Small) change direction across suites, which is why surface preference is not treated as a provider-level trait.
Table: Model-level means over the four suites and per-suite signed gaps, sorted from most schema-vulnerable to most data-vulnerable. schema more vulnerable, data more vulnerable.
| Model | Mean over suites | Δsurface = schema − data (pp) | |||||
|---|---|---|---|---|---|---|---|
| Data ASR | Schema ASR | Worst ASR | Banking | Slack | Travel | Workspace | |
| Gemini-3-Flash | 8.8 | 74.6 | 74.6 | +72.2 | +78.1 | +61.9 | +50.9 |
| GPT-5.4 | 0.0 | 29.4 | 29.4 | +18.1 | +28.6 | +40.5 | +30.4 |
| Ministral-3B | 16.6 | 19.7 | 19.7 | 0.0 | +4.0 | +8.6 | 0.0 |
| Mistral-Small | 32.0 | 33.7 | 40.7 | +8.9 | −28.0 | +2.9 | +22.9 |
| Qwen3-32B | 45.4 | 44.6 | 51.2 | +9.7 | −14.3 | −11.9 | +13.4 |
| DeepSeek-V3 | 50.9 | 46.5 | 51.6 | 0.0 | −12.0 | −8.6 | +2.9 |
| Mistral-Large | 61.6 | 53.6 | 61.6 | −8.9 | −16.0 | −2.9 | −4.3 |
| GPT-OSS-20B | 42.3 | 30.5 | 42.3 | −6.7 | −12.0 | −14.3 | −14.3 |
| Qwen3.5-Flash | 48.5 | 31.1 | 51.7 | +12.5 | −38.1 | −38.1 | −6.2 |
| GPT-4o | 53.3 | 27.2 | 54.7 | −20.0 | −80.0 | +5.7 | −10.0 |
| Llama-3.3-70B | 32.3 | 4.7 | 32.3 | −33.3 | −46.7 | −23.8 | −6.2 |
| GPT-4o-mini | 40.3 | 6.7 | 40.3 | −33.3 | −64.0 | −28.6 | −8.6 |
| GPT-4.1 | 54.9 | 6.8 | 54.9 | −48.9 | −92.0 | −34.3 | −17.1 |
Native function calling: we reproduce the intervention outside AgentDojo in four self-contained tool-use scenarios (document exfiltration, contact exfiltration, payment redirection, file deletion); each model × surface cell has 40 episodes. The opposite preferences persist: Gemini-3-Flash is schema-vulnerable, GPT-4.1 data-vulnerable, and Qwen3-32B balanced. We treat this as a qualitative replication.
| Model | Data ASR | Schema ASR | Δsurface |
|---|---|---|---|
| Gemini-3-Flash | 0.0% | 100.0% | +100.0 pp |
| GPT-4.1 | 100.0% | 50.0% | −50.0 pp |
| Qwen3-32B | 65.0% | 65.0% | 0.0 pp |
| Llama-3.3-70B | 25.0% | 75.0% | +50.0 pp |
| Pooled | 47.5% | 72.5% | +25.0 pp |
Payload components: removing any single component of the template does not remove schema-surface attack success, and the bare attacker instruction alone still reaches 21.7%. This checks sensitivity to the one payload used; it does not show the result holds for other attack families.
| Payload variant | Schema ASR |
|---|---|
Full workflow_dependency template | 42.2% |
| Without trigger condition | 51.7% |
| Without urgency / pressure | 51.7% |
| Without justification | 46.7% |
| Without override marker | 58.3% |
| Minimal attacker instruction only | 21.7% |
Choosing a surface without history: a handful of paired probes on the target cell also finds the riskier surface. Five probes recover 73% of the oracle gain, while guessing from provider identity alone does worse than data-only evaluation.
| Selection strategy | Target probes | Realized ASR | Rel. to oracle gain |
|---|---|---|---|
| Data-only baseline | 0 | 37.5% | 0% |
| Schema-only baseline | 0 | 31.5% | — |
| Within-provider cross-validation | 0 | 41.7% | 46% |
| Strict leave-one-provider-out | 0 | 30.5% | Worse than baseline |
| 1-probe selection | 1 | 41.3% | 42% |
| 2-probe selection | 2 | 42.5% | 56% |
| 5-probe selection | 5 | 44.1% | 73% |
| 10-probe selection | 10 | 44.7% | 80% |
| Oracle worst surface | — | 46.5% | 100% |
We presented a controlled cross-surface study of prompt injection in tool-using LLM agents, placing a byte-identical payload in either tool outputs or tool descriptions. Across 13 models and four AgentDojo suites, changing only the injection surface reverses 44.9% of pairwise model-security rankings, and the surface preference learned from other suites predicts the more vulnerable surface on an unseen suite with 76.9% accuracy. Defense results show the same dependence: success against one injection surface does not imply comparable protection on another.
Our study covers two surfaces and one main payload family, so broader claims need evaluations across more channels and payloads. The present results already show that prompt-injection robustness is not surface-invariant: evaluations should measure vulnerability across the surfaces a realistic attacker can reach, rather than treat a single-surface score as a complete measure of model robustness.
@inproceedings{sakib2026surface,
title={The Surface You Test Is Not the Surface That Breaks},
author={Sakib, Syed Nazmus and Haque, Nafiul and Amin, Shahrear Bin and
Arman, Shifat E.},
booktitle={NeurIPS 2026 Workshop on Agents in the Wild: Safety, Security, and Beyond},
year={2026}
}