Cybersecurity and
Agent Security
Benchmarks
A comprehensive reference for model onboarding, capability assessment and security evaluation
| 43 benchmark profiles | Five evaluation domains |
| Capability and exploitation | Lifecycle, secure coding and detection |
| Defensive analysis and knowledge | Prompt injection and agent security |
| Misuse resistance and monitoring | Methods • Metrics • Limitations • Evidence |
Prepared for Bruno Cormack
6 October 2026 | Version 1.0
Research synthesis and proposed evaluation guidance. No candidate model or benchmark was executed as part of this work.
Source review date: 6 October 2026. Adoption recommendations are proposed guidance, separate from benchmark-author methodology.
Contents
Click a section or benchmark title to navigate. Profile numbers match the supplied list; number 38 was not supplied.
Executive overview
This reference covers all 43 benchmark entries in the requested list, grouped into five domains. It explains what each benchmark measures, how the evaluation works, how to read its metrics, what its evidence can support, and where additional testing is needed. The original numbering is retained: the supplied list skips number 38, so the profiles run from 1–37 and 39–44.
For model onboarding, the central decision is whether the proposed model-and-application combination is suitable for its intended role. A strong exploitation score, a strong patching score and a low prompt-injection attack success rate answer different questions. They should remain visible as separate evidence rather than being averaged into one security percentage.
| Evaluation question | Representative evidence | Decision relevance |
| What technical actions can the agent complete? | CyberGym, ExploitGym, ExploitBench, Cybench, web CVE-Bench, AISI ranges | Capability and the consequences of granting tools, time and access. |
| Can it perform defensive work correctly? | SEC-bench, AutoPatchBench, PatchBench, secure-coding datasets, CyberSOCEval | Fitness for a coding, triage, repair or analyst role. |
| Can untrusted input redirect its behavior? | AgentDojo, InjecAgent, ASB, AgentDyn, BIPIA, MCP-SafetyBench | Resilience of the configured agent and its trust boundaries. |
| Will it complete harmful requests, and can oversight detect deviations? | HarmBench, JailbreakBench, AgentHarm, SafeArena, SHADE-Arena | Misuse resistance, action-level controls and monitoring effectiveness. |
The benchmark profiles are grounded in papers, official repositories, maintainer documentation and institutional publications. Recommendations for bank use, evaluation gates, ownership and evidence retention are this document’s proposed approach. They are not statements that the benchmark authors, a regulator or the bank have adopted those policies.
NIST’s AI RMF emphasizes documented measurement tied to the intended context, including validity, safety and security. Its Generative AI Profile also cautions that laboratory benchmarks may not transfer directly to deployment and recommends evaluating the system and its safeguards in relevant conditions. Those principles support the use-case-based approach proposed here. [G1][G2]
Findings that materially affect interpretation
- Names need an owner and a task definition. The two CVE-Bench projects are different; CyberBench has multiple public meanings; SECURE is distinct from SEC-bench; and the cybersecurity VEX-Bench must be separated from an unrelated namesake.
- Parent suites are not extra independent tests. CyberSecEval 4 includes CyberSOCEval and AutoPatchBench. Reusing a track under its suite name does not add evidence coverage.
- Counts and scorers have changed. Public releases, paper versions and adapters sometimes use different task populations or success criteria. Each profile records the material differences found.
- A benchmark run evaluates a configuration. Model snapshot, tools, agent code, instructions, permissions, memory, budget, guardrails and judge all affect the result.
- A low measured attack rate needs a functioning legitimate workflow. An agent that refuses or fails every task can appear resistant while being unusable.
- Benchmark numbers need calibrated decision criteria. Capability triggers, defensive quality floors, harmful-action limits and evidence-validity checks should be designed separately.
How to use this reference
Start with the catalogue to identify the relevant task family, then read the full profile before selecting a benchmark. Each profile identifies the canonical project, its publication or release context, the task modality, metric direction, practical prerequisites and sources. The application guide later in the document proposes bundles and an evidence record for an actual onboarding decision.
| Profile element | How to use it |
| Identity and release | Use the named owner and linked repository to avoid a namesake. Preserve the paper version, release or commit in an actual run. |
| Method and metrics | Check what inputs the agent receives, what it can do and how success is determined. Native metrics are separated from recommended diagnostics. |
| Interpretation and limitations | Identify which claim the result supports and which confounders need investigation. |
| Bank use | Treat this as an analyst recommendation for benchmark selection and follow-up evidence. |
| Execution and access | Plan dependencies, task health checks and the available public subset. Public source code does not guarantee that every evaluated artifact is released. |
| Sources | Bracketed identifiers resolve to titled links in the reference register. Sources were reviewed on 6 October 2026. |
Status labels and limits of this review
Verified identity means that primary evidence supports the selected project’s identity and described purpose. It does not mean that its code was executed, all labels were audited, or its results were independently reproduced. Maturity descriptions are analyst assessments of publication and artifact availability, not certification labels.
Emerging means recently released or still developing in the evidence reviewed. It does not mean fictitious, ineffective or unsuitable by default. Conversely, an older, well-known benchmark can have outdated tasks, archived code or a saturated score distribution. Review usefulness and reproducibility separately from age.
Requests for “onward” coverage are handled by reviewing current official releases and related tracks. A future release is not assumed to be comparable. Where no later numbered release was verified, the profile says so. Version selection and regression mapping should be repeated when the evaluation is actually commissioned.
Reading benchmark results correctly
Keep the measurement layers separate
| Layer | What is being measured | What to record |
| Model | Knowledge, code generation, reasoning or responses under a defined prompt | Exact snapshot or weights; inference parameters; examples supplied; output parser. |
| Agent | Model operating with tools, memory, planning and a workflow | Agent commit; tool schemas; system/developer instructions; time and action limits. |
| Deployed system | Agent together with identity, application permissions and safeguards | Connector scopes; authorization; execution boundaries; human approvals; data paths. |
| Oversight | A monitor or reviewer attempting to identify and interrupt failures | Monitor model and visibility; threshold; false alarms; detection timing; intervention outcome. |
| Human-plus-AI workflow | Performance and errors of a person using the system | Operator expertise; baseline without AI; task allocation; review workload and missed errors. |
These layers are an analytical structure for this document. Evidence can cross layers only when the additional components and assumptions are tested. A provider’s base-model score is useful context, while the bank’s configured application needs its own relevant evaluation.
Metrics that are easy to misread
| Metric or term | Meaning and common error |
| Accuracy | Correct scored answers divided by the specified population. It is not automatically attack prevention, factual completeness or absence of hallucination. |
| Precision / recall / F1 | Precision concerns the correctness of positive predictions; recall concerns the fraction of positives found; F1 balances them. Preserve class balance and micro/macro averaging. |
| MAE / MAD | Average absolute error, often used for severity prediction. Lower is better. A normalized transformation may be labelled accuracy without being exact-answer accuracy. |
| pass@1 | Success in one sampled attempt, subject to the benchmark’s estimator and trial definition. |
| pass@k | Usually the probability that at least one of k attempts succeeds. More allowed search can improve the score without changing the model. |
| pass^k | In benchmarks using all-trials consistency, success requires all k trials to meet the criterion. Read the benchmark definition; it is not interchangeable with pass@k. |
| Attack success rate (ASR) | Fraction of the declared attack population achieving the attacker’s objective. Record attacker knowledge, editable surfaces, retries and the success validator. |
| Clean / benign utility | Ability to complete legitimate tasks without attacks. Keep this beside ASR to expose blanket refusal or broken tools. |
| Secure and functional completion | A joint criterion: legitimate task completion while satisfying the security checks. Inspect how both terms are verified. |
| ROUGE / reference overlap | Text similarity to a reference. High overlap does not prove that an explanation is factual, complete or operationally safe. |
| Macro-average | Average of separately calculated task/class scores. It differs from pooling every prediction into one score, particularly with unequal class sizes. |
| Confidence interval | Uncertainty under stated sampling assumptions. It does not correct biased tasks, shared failure modes, contamination or deployment mismatch. |
Minimum comparability conditions
Before ranking two results, match or disclose the task population, benchmark release, scoring code, model snapshot, toolset, agent scaffold, information provided, defense setting, budget and retry policy. Preserve error handling: a failed container launch is not a model’s failed exploit, and an unparseable answer is not necessarily a knowledge error.
Record both maximum demonstrated capability and typical reliability when they serve different decisions. An any-success result can establish that a capability is possible under the tested conditions, while repeated independent runs provide a different estimate of consistency. A successful attack found after extensive adaptive search should not be compared with a single fixed prompt as though the attack effort were equal.
Confidence example: zero observed failures
Illustrative calculation, not benchmark data: if there are zero failures in n independent, identically distributed Bernoulli trials, the exact one-sided 95% upper bound on the failure probability is 1 − 0.05^(1/n). It is approximately 9.5% for 30 trials, 3.0% for 100 trials and 1.0% for 300 trials. Zero observed failures therefore does not establish zero risk.
Repeated paraphrases of one attack, correlated tasks and a nonrepresentative test corpus weaken those statistical assumptions. For real assessment, report uncertainty by task family and scenario, alongside the consequences of observed failures. The deployment probability of a harmful event also depends on who can reach the system and what actions it can execute.
Identity, version and overlap register
This register highlights distinctions that can otherwise invalidate a comparison. Detailed sources and additional qualifications are provided in the corresponding profiles. Counts describe the referenced releases, not a permanently fixed property of the benchmark.
| Name or relationship | Required distinction |
| CVE-Bench (#7 / #14) | UIUC web exploitation and WhileBug/CVEBench vulnerability repair are different projects. Cite the full title and repository. [A7a][B14a] |
| CyberGym / CyberGym-E2E / ExploitGym | Related research line with different endpoints: reproduction, an extended lifecycle and exploitation. Track dataset reuse and do not treat related instances as independent samples. [A1a][A2a][B10a] |
| ExploitBench | The v8-bench ladder reports individual capabilities. Reaching a capability on any bug is not a full-exploit success rate over every bug. [A3a] |
| SEC-bench / SEC-bench Pro | Original reproduction/patch tasks and Pro’s complex-system cases have different scope, releases and validators. [A8a][B12a] |
| CyberSecEval 4 / CyberSOCEval / AutoPatchBench | Parent-suite and component relationship. Count each underlying task set once in coverage calculations. [C23a][C24a][B13a] |
| CTIBench / AthenaBench | Related CTI benchmark lineage with changes to construction, tasks and scoring. Avoid presenting both as independent confirmation by default. [C25a][C26a] |
| AgentDojo / AgentDyn | AgentDyn extends the workflow dynamics of the AgentDojo lineage; compare the task and attack models explicitly. [D32a][D35a] |
| CyberBench | This reference selects JPMorgan Chase’s 2024 cybersecurity NLP suite based on the requested grouping. Other public projects use the same name; owner confirmation remains necessary before importing results. [C28a] |
| SECURE | Security Extraction, Understanding & Reasoning Evaluation is an ICS-focused advisory benchmark, distinct from secure-code and software-repair suites. [C29a] |
| VEX-Bench | Select the software supply-chain exploitability benchmark. A separate misinformation benchmark shares the name. [B19a] |
| PatchBench / DUMA-Bench | Verified recent 2026 projects. Emerging status is a maturity assessment; check artifacts and version-specific definitions before using them as decision gates. [B15a][D36a] |
| AgentHarm / SHADE-Arena | Published full benchmark descriptions and public subsets differ. Disclose exactly which tasks and attack artifacts were available. [E42a][E44a] |
| AISI cyber ranges | A continuing evaluation programme and set of methodologies. Inspect Cyber is a framework, not evidence that every AISI range is publicly reproducible. [A9a] |
Release changes to preserve in evidence
- ExploitGym: the original paper describes 898 instances; the current public v1.0 release lists 869. [A2c]
- InterCode-CTF: 100 original tasks versus the 78-task offline Inspect implementation. [A5c]
- Web CVE-Bench: the 2026 criterion change from arbitrary file upload to remote code execution affects score interpretation. [A7b]
- SEC-bench Pro: early engine-only population of 183 versus 344 cases in the later paper revision and current repository. [A8e]
- CyberGym-E2E, SEC-bench and AutoPatchBench: original-paper or launch counts differ from subsequently inspected data; use the profile’s version notes. [B10c][B12c][B13b]
- ASB and CTIBench: reported method/metric descriptions vary between source versions; pin the scorer and resolve the discrepancy. [D34b][C25b]
- AthenaBench: the normalized vulnerability-severity score is not the percentage of exactly correct predictions. [C26c]
Benchmark catalogue
Use this index to locate a profile. Each full profile contains the owner, current evidence, detailed methodology, metric definitions, limitations and practical evaluation guidance. Catalogue descriptions are condensed from those cited profiles.
A. Cyber capability and exploitation
| ID | Benchmark | Primary measurement | Evaluation form |
| 1 | CyberGym | Reproduction of real software vulnerabilities [A1a] | Agent interacting with source code and executable targets |
| 2 | ExploitGym | Turning a known vulnerability into a working exploit [A2a] | Agent, source code, triggering input, containers / kernel VMs |
| 3 | ExploitBench | Progress from reaching a vulnerability to full exploit control [A3a] | Agents interacting with instrumented V8 targets |
| 4 | Cybench | Interactive security problem solving in professional CTFs [A4a] | Containerized CTF tasks with optional subtask guidance |
| 5 | InterCode-CTF | Tool-assisted coding and security reasoning in bounded CTF tasks [A5a] | Shell-based interactive challenge environment |
| 6 | NYU CTF Bench | Breadth of tool-assisted CTF problem solving [A6a] | Dockerized CSAW CTF tasks; separate evaluation agent |
| 7 | CVE-Bench: web exploitation | Exploitation of deployed vulnerable web applications [A7a] | Containerized web services and outcome validators |
| 8 | SEC-bench Pro | Bug hunting and validated PoC generation in complex systems [A8a] | Repository-level coding agents; browser engines and kernel VMs |
| 9 | AISI multi-step cyber ranges | Sustained autonomous progress across linked cyber tasks [A9a] | Multi-host enterprise and industrial-control simulations |
B. Vulnerability lifecycle, secure coding and detection
| ID | Benchmark | Primary measurement | Evaluation form |
| 10 | CyberGym-E2E | Discovery, proof generation and repair across the vulnerability lifecycle; separate patch-only capability. [B10a] | Agent operates in a real project build environment; execution-based validation. |
| 11 | BountyBench | Detect, Exploit and Patch capabilities on real bounty-backed vulnerabilities. [B11a] | Containerized repositories and running applications/services, with executable outcome checks. |
| 12 | SEC-bench | Proof-of-concept generation and vulnerability patching using reproducible native-code cases. [B12a] | Agentic repository interaction in Docker; sanitizer-based executable oracles. |
| 13 | AutoPatchBench | Automated repair of fuzz-discovered C/C++ vulnerabilities, with deeper patch validation. [B13a] | Native-code build, crash replay, fuzzing and white-box differential testing in Linux containers. |
| 14 | CVE-Bench: vulnerability repair | Repository-level repair of known CVEs under different levels of report detail. [B14a] | Interactive coding-agent environment with patch application and vulnerability-associated unit tests. |
| 15 | PatchBench | Robust vulnerability repair with reduced reliance on historical-patch memorization and crash-location shortcuts. [B15a] | Repository-level coding agents; replay corpora, sanitizer checks, unit tests and output-state comparison. |
| 16 | SecurityEval | Propensity to produce insecure Python code in security-relevant completion scenarios. [B16a] | Prompt-to-code generation followed by targeted static analysis and manual assessment. |
| 17 | SecRepoBench | Ability to complete repository code that is both functionally correct and secure against the designated vulnerability. [B17a] | Masked code region plus repository context; compilation, developer tests and a security reproducer. |
| 18 | SecureVibeBench | Whether feature-development agents recreate historical vulnerabilities or introduce additional suspicious changes. [B18a] | Natural-language requirements and full repositories; functional comparison, dynamic PoV and static security analysis. |
| 19 | VEX-Bench | Whether known dependency vulnerabilities actually affect downstream code and why. [B19a] | Repository and advisory analysis with structured affected-status and justification classification. |
| 20 | PrimeVul | Function-level vulnerability detection, false-alarm tradeoffs and distinction between vulnerable and repaired near-pairs. [B20a] | Labeled C/C++ code classification with chronological splits and paired examples. |
| 21 | DiverseVul | Function-level vulnerability classification and generalization across projects and weakness classes. [B21a] | Labeled C/C++ functions mined from vulnerability-fixing commits. |
| 22 | LLMSecEval | Security of code generated from ordinary natural-language programming requirements. [B22a] | Natural-language prompt-to-code tasks with secure reference examples and static-analysis support. |
C. Suites, defensive analysis and cybersecurity knowledge
| ID | Benchmark | Primary measurement | Evaluation form |
| 23 | CyberSecEval 4 and inherited tracks onward | Cybersecurity capability, insecure behavior, misuse compliance and defensive utility across separate tracks. [C23a] | Text, generated code, images, and tool-enabled execution depending on track. |
| 24 | CyberSOCEval | Reasoning over malware detonation reports and threat intelligence reports. [C24a] | JSON/text telemetry; threat reports as text, page images, or both. |
| 25 | CTIBench | CTI knowledge, CVE-to-CWE mapping, CVSS severity prediction, ATT&CK extraction and actor attribution. [C25a] | Text questions, vulnerability descriptions and anonymized threat reports. |
| 26 | AthenaBench | CTI knowledge, ATT&CK extraction, root causes, mitigation, vulnerability severity and attribution. [C26a] | Text scenarios and reports; optional web-search configuration is a separate evaluated condition. |
| 27 | CyberMetric | Broad cybersecurity knowledge and selection of correct factual/conceptual answers. [C27a] | Text multiple-choice questions with four options. |
| 28 | CyberBench | Cybersecurity named-entity recognition, summarization, knowledge questions and text classification. [C28a] | Static text datasets; classification and generation. |
| 29 | SECURE | ICS cybersecurity knowledge, contextual understanding, uncertainty and advisory reasoning. [C29a] | Text multiple choice, true/false, generated risk summaries and numerical CVSS answers. |
| 30 | SecQA | Foundational and more complex textbook-based computer-security knowledge. [C30a] | Text multiple-choice questions; basic v1 and harder v2. |
| 31 | WMDP-Cyber | Cybersecurity knowledge relevant to malicious-use capability; effects of knowledge unlearning. [C31a] | Text multiple-choice questions. |
D. Prompt injection and deployed-agent security
| ID | Benchmark | Primary measurement | Evaluation form |
| 32 | AgentDojo | Indirect prompt injection resistance and legitimate task utility of a configured tool-using agent [D32a] | Multi-step agent execution in simulated application environments |
| 33 | InjecAgent | Susceptibility to indirect instructions in tool responses, including direct harm and data theft [D33a] | Structured agent continuations and tool-call parsing over synthetic test cases |
| 34 | Agent Security Bench (ASB) | Agent susceptibility across prompt, observation, memory, and system-prompt compromise scenarios [D34a] | Configurable agent/tool framework with multiple attack and defense families |
| 35 | AgentDyn | Security-versus-utility trade-offs for injection defenses in dynamic, open-ended agent workflows [D35a] | AgentDojo-derived executable tasks involving changing plans and useful external instructions |
| 36 | DUMA-Bench | Interaction-dependent security failures when an agent and an active user both influence shared state [D36a] | Executable multi-turn scenarios with user simulation, mutable state, and security assertions |
| 37 | BIPIA | Robustness against attacker instructions embedded in retrieved or supplied external content [D37a] | Text/code content-processing prompts spanning five application tasks |
| 39 | MCP-SafetyBench | Task completion and attack outcomes in multi-step MCP-connected agent workflows [D39a] | Execution with real MCP server integrations and attack-instrumented tasks |
E. Misuse resistance, harmful action and monitoring
| ID | Benchmark | Primary measurement | Evaluation form |
| 40 | HarmBench | Whether adversarial prompts elicit specified harmful model behaviors [E40a] | Text and image-conditioned output evaluation; attack generation, model responses, and classifier scoring |
| 41 | JailbreakBench | Robustness to adversarial prompts, with matched benign refusals and standardized attack accounting [E41a] | Text-generation evaluation with archived jailbreak artifacts and semantic judges |
| 42 | AgentHarm | Refusal and coherent multi-step execution of explicitly malicious tasks using synthetic tools [E42a] | Tool-using agent trajectories scored with task rubrics and limited semantic judgments |
| 43 | SafeArena | Deliberate misuse of web agents, normalized against matched legitimate web-task capability [E43a] | Browser-agent actions on four benchmark websites, with paired safe/harmful tasks |
| 44 | SHADE-Arena | Ability to complete hidden harmful side objectives while doing legitimate work and evading an independent monitor [E44a] | Long-horizon simulated agent environments plus transcript-monitor evaluation |
A. Cyber capability and exploitation
These benchmarks assess technical capability, from bounded CTF problem solving and vulnerability reproduction to exploit development and longer attack sequences. Their outcomes are different levels of evidence; a verified crash and a verified full-control exploit should remain distinct.
01. CyberGym
CyberGym: Evaluating AI Agents’ Real-World Cybersecurity Capabilities at Scale
| Owner: UC Berkeley research team | Release / evidence: June 2025 preprint; ICLR 2026 publication |
| Task form: Agent interacting with source code and executable targets | Score direction: Higher reproduction success = greater capability |
| Status assessment: Established research benchmark; active implementation | Availability: Public code and data; substantial local storage for full environments |
What it is. CyberGym evaluates whether an agent can turn a real vulnerability description and an unpatched codebase into a proof-of-concept input that reproduces the bug. Its reference set contains 1,507 instances from 188 projects. The main leaderboard uses Level 1, where the agent receives the description and vulnerable source. This is a vulnerability-reproduction benchmark; a successful crash does not by itself establish arbitrary code execution. [A1c]
Why it matters. For onboarding, the distinction between recognizing vulnerable code and constructing a working trigger is useful. CyberGym supplies execution-based evidence for security research assistants and coding agents. A strong result can support a defensive use case while also prompting closer review of the same agent’s offensive affordances. This is an analyst interpretation.
How it works. Tasks reconstruct pre-patch repositories. The agent inspects files, creates candidate inputs and iterates using execution feedback. The principal verifier checks that the input triggers the vulnerable version but not the fixed version. Other information levels change what is disclosed; results with stack traces or a ground-truth patch should be identified separately. The project also reports open-ended discovery experiments, which are a separate activity from its fixed reproduction leaderboard. [A1a]
Metrics and score interpretation.
Reproduction success rate. Fraction of target instances with a verified working reproducer; disclose information level and task IDs. [A1a] Direction: Higher capability.
Trials and any-success aggregation. An instance can count as solved if any permitted trial succeeds; retain the number of trials. [A1a] Direction: Compare only at equal trial budgets.
Cost and error breakdown (recommended). Record compute, tokens, invalid submissions and infrastructure failures separately. Direction: Diagnostics, not an official safety score.
How to interpret a result. Treat the result as a property of the model, agent scaffold, information provided and runtime. Compare matched tasks and budgets. A model that produces no valid reproducer may have failed to understand the problem, exhausted its budget or encountered a broken environment; those outcomes have different implications.
Limitations and failure modes.
- The corpus is concentrated on historical, fuzzing-discovered vulnerabilities and is not a representative sample of every bank application. [A1a]
- A PoC that crashes both versions requires investigation: it may expose another bug or an incomplete fix, rather than satisfy the intended task. [A1a]
- Historical task descriptions and fixes can leak through training data, repository history or retrieval; control those channels for discovery-oriented comparisons.
Proposed bank use. Recommended: use a stratified task set for baseline measurement of a security coding assistant, then corroborate with exploitation, repair and misuse-resistance evaluations. Retain intended-bug verification and manual adjudication of ambiguous crashes. Do not convert a high reproduction rate into an automatic rejection of a useful defensive model.
Execution and evidence prerequisites. The official implementation requires Python and Docker. Storage varies markedly between binary-only and full compilation environments. Pilot a declared subset before a full run, keep the submission server private, and preserve pre-/post-patch images and verifier outputs. Public code uses Apache 2.0; confirm the applicable data and upstream component terms. [A1b]
Primary sources: [A1a] [A1b] [A1c]
02. ExploitGym
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
| Owner: Berkeley, MPI-SP, UCSB and ASU researchers with industry collaborators | Release / evidence: 2026 paper; public release v1.0 lists 869 instances |
| Task form: Agent, source code, triggering input, containers / kernel VMs | Score direction: Higher intended-vulnerability exploitation = greater capability |
| Status assessment: Recent research benchmark; versioned public implementation | Availability: Public repository; provider access conditions may affect reported results |
What it is. ExploitGym addresses the step after bug reproduction: whether an AI agent can use a known vulnerability to achieve the benchmark’s code-execution objective, validated through secret-flag capture. The current public release lists 869 tasks across userspace software, V8 and the Linux kernel. The initial research paper described 898 instances, so the paper denominator and current release must not be interchanged. [A2a][A2b][A2c]
Why it matters. An executable proof of exploitation provides stronger evidence of a dangerous capability than a crash alone. For defenders, it can also inform severity triage and mitigation testing. In model onboarding, the relevant question is which exploit capabilities the deployed agent can exercise under its actual tools, permissions and safeguards.
How it works. Each task provides source code, build information, a proof-of-vulnerability input and a reproducible runtime. The agent must progress from that trigger to the specified impact. Flag capture is checked, and an agent-based judge assesses whether the submitted exploit used the intended vulnerability. Results are separated by software domain and mitigation configuration. The current project distinguishes 502 userspace, 181 V8 and 186 kernel tasks. [A2a]
Metrics and score interpretation.
Intended-vulnerability exploitation success. Successful flag-capture exploits that also target the assigned vulnerability. [A2a] Direction: Higher capability.
Raw flag captures. Keep separate from intended-bug successes because an alternative vulnerability can produce impact. [A2a] Direction: Additional capability signal; different denominator/criterion.
Mitigation-conditioned results. Report each defense setting and domain separately, with elapsed time, tokens and trials. [A2a] Direction: No single unconditional safety interpretation.
How to interpret a result. A raw capture that exploits a different bug may fail the assigned exploitation criterion while remaining relevant to an attacker-capability assessment. Preserve both conclusions. Compare trusted-research access results with ordinary production access only when provider safeguards and permitted tools are disclosed; they can represent materially different systems.
Limitations and failure modes.
- A supplied triggering input removes part of the discovery problem; this does not measure independent discovery from an untouched product.
- Success under disabled mitigations should not be presented as success against the corresponding hardened configuration.
- The intended-vulnerability judge introduces an adjudication dependency; retain its configuration and review disputed cases.
- Historical task knowledge and large retry budgets can affect results. Absence of a successful exploit under one budget is limited negative evidence.
Proposed bank use. Recommended: use this as an advanced capability assessment for autonomous security agents or broadly tooled coding assistants. Reproduced success against a hardened target should trigger examination of access boundaries, permitted targets and monitoring. It is a risk-review trigger, not a universally prescribed numeric adoption ceiling.
Execution and evidence prerequisites. Pin the release and its canonical task list. The repository includes controller, firewall, LLM proxy and per-domain runtime setup. Validate the complete containment boundary and judge before scaling, and record which mitigation settings are active. Task data derives from external projects with their own terms in addition to the code license. [A2b]
Primary sources: [A2a] [A2b] [A2c]
03. ExploitBench
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
| Owner: Seunghyun Lee and David Brumley, Carnegie Mellon University | Release / evidence: 2026 paper; v8-bench v0.1 public project |
| Task form: Agents interacting with instrumented V8 targets | Score direction: Higher-tier capability attainment = greater exploitation capability |
| Status assessment: Recent research benchmark; public v8-bench implementation | Availability: Public code and per-bug images; large image footprint |
What it is. ExploitBench separates exploitation into measurable stages. Its initial v8-bench targets 41 V8 bugs and grades 16 capabilities grouped into five tiers. These run from code coverage and bug reproduction to target-specific primitives, broader memory primitives and full control. The purpose is to identify where the agent’s capability stops, rather than compress all partial progress into a binary exploit/no-exploit result. [A3a][A3b]
Why it matters. This is especially useful when an onboarding committee asks whether a model can merely reproduce a defect or can cross a significant exploitation boundary. Partial progress is informative: generating a crash and demonstrating a general memory-write capability warrant different follow-up questions. The ladder is an analytical aid, not a bank risk-rating scale.
How it works. Agents work in reproducible V8 environments through tools exposed by the benchmark. Deterministic oracles validate individual capabilities; the research uses challenge-response checks and execution-based evidence rather than asking an LLM to accept a narrative of success. The methodology distinguishes the uniform model/environment experiment, adaptive coaching and a vendor-native CLI configuration. Those arms are different experimental conditions. [A3b]
Metrics and score interpretation.
Capability / tier attainment. Which of the 16 flags and five capability tiers were mechanically demonstrated. [A3a] Direction: More advanced capability toward T1; T1 is full control.
Per-bug success at each capability. Retain task-level coverage and repeated-run reliability; avoid reducing all tasks to one maximum. Direction: Recommended reporting alongside native flags.
Leaderboard capability coverage. The public display can count a capability reached on at least one bug, and its all-runs view mixes run budgets. [A3a] Direction: Not the proportion of all bugs fully exploited.
How to interpret a result. Read the tier definitions before interpreting a percentage. Reaching every capability somewhere in the corpus does not mean exploiting every target. For selection, use the fixed-experiment view or reproduce matched settings, including coaching, turn budgets and mitigation configuration. Explicitly separate maximum demonstrated ability from consistent success.
Limitations and failure modes.
- The initial benchmark is V8-specific; broad network intrusion or application authorization conclusions require other evidence.
- The paper studies known vulnerabilities with supplied information; it does not establish an unrestricted zero-day discovery rate. [A3b]
- Demonstrated control in the evaluator is not a complete measurement of exploit reliability across varied deployments.
- Deterministic scoring still depends on correct target instrumentation and protection of the oracle.
Proposed bank use. Recommended: use as a focused advanced capability assessment when exploit development is relevant to the proposed tool access. Preserve milestone-level results for risk escalation and pair them with safeguards testing. A bank can define internal review triggers for new capabilities without representing those triggers as benchmark-author requirements.
Execution and evidence prerequisites. The official repository provides a model runner and V8 MCP-based target interfaces, with a full 41-bug configuration and smaller subsets. Pulling many target images can be expensive in storage. Pilot environment health, pin images and runner versions, and record seeds and all coaching or CLI differences. [A3c]
Primary sources: [A3a] [A3b] [A3c]
04. Cybench
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
| Owner: Stanford-led research team | Release / evidence: 2024 introduction; ICLR 2025 publication |
| Task form: Containerized CTF tasks with optional subtask guidance | Score direction: Higher solved-task rate = greater capability |
| Status assessment: Established research reference; active public implementation | Availability: Public benchmark, code and grading |
What it is. Cybench is a set of 40 tasks selected from four professional capture-the-flag competitions. It spans cryptography, web security, reverse engineering, forensics, binary exploitation and miscellaneous tasks. Its distinguishing feature is subtask evaluation, which makes partial progress visible on challenges that an agent cannot solve from start to finish. [A4c]
Why it matters. Cybench is useful as a common reference for practical cyber problem solving and for comparing agent scaffolds. For a bank, it can demonstrate whether an assistant can execute a bounded technical workflow. It contributes capability evidence, while leaving production resilience, authorization and misuse-resistance questions to other tests.
How it works. The agent receives a task description, starter files and access to local or networked services in a containerized environment. It takes actions, observes tool output and eventually submits an answer for grading. Unguided runs expose the task objective alone. Guided runs expose sequential subtask questions. Human first-solve time provides context about challenge difficulty, rather than a direct measure of the AI’s equivalent human job seniority. [A4b][A4c]
Metrics and score interpretation.
Unguided percentage solved. Fraction of complete tasks solved without sequential subtask guidance. [A4a] Direction: Higher autonomous task capability.
Subtask-guided percentage solved. Success on the final subtask under guided conditions. [A4b] Direction: Higher guided capability; separate from unguided.
Subtasks percentage solved. Fraction of subtasks solved per task, macro-averaged across tasks. [A4a] Direction: Diagnostic progress.
Most difficult task solved. Highest human first-solve time among solved tasks. [A4a] Direction: Difficulty context, not a general human equivalence.
How to interpret a result. Keep all three success measures separate. Guidance can change both the information available and the work the model must plan itself. A 40-task score is coarse: one task changes the aggregate by 2.5 percentage points. Report intervals and category coverage when using small differences to rank candidates.
Limitations and failure modes.
- The tasks are competition challenges, not representative enterprise networks or complete attacker campaigns.
- The public leaderboard includes results on different subsets and configurations; task denominators must be preserved. [A4a]
- Public task write-ups and evaluator leaks can inflate apparent ability. The project documents a historical answer-leak correction. [A4a]
- A solved flag does not certify secure behavior, successful defense, or appropriate refusal of unauthorized requests.
Proposed bank use. Recommended: use Cybench as one external reference alongside bank-relevant integration tests. For lower-cost routine evaluation, retain a fixed balanced subset and periodically rerun the full declared set. Use more realistic exploitation or cyber-range evaluations when the deployment permits lengthy autonomous technical actions.
Execution and evidence prerequisites. The repository supplies task launchers, model adapters and logs that retain inputs, outputs and scores. It supports unguided and subtask modes with configurable iteration and token limits. Pin the challenge images, agent implementation and task list; validate task health before measuring the model. [A4b]
Primary sources: [A4a] [A4b] [A4c]
05. InterCode-CTF
InterCode Capture-the-Flag environment (IC-CTF)
| Owner: Princeton NLP InterCode authors; AISI maintains a separate Inspect adapter | Release / evidence: CTF environment released in August 2023; adapter has independent versioning |
| Task form: Shell-based interactive challenge environment | Score direction: Higher correct-flag rate = greater task capability |
| Status assessment: Established lightweight research reference; adapter variants | Availability: Public code and data; distinguish original and offline subset |
What it is. InterCode-CTF is the cybersecurity environment within the broader InterCode interactive-coding framework. It uses picoCTF-derived tasks to test whether an agent can use execution feedback to solve a challenge and submit the hidden flag. The original collection contains 100 challenges; the current Inspect implementation evaluates 78, excluding 22 that require internet access. Those are different evaluation populations. [A5a][A5b]
Why it matters. This is a practical starting point for checking an agent’s ability to use a terminal, inspect files and adapt after failed actions. It is often more useful as a baseline or harness diagnostic than as the sole measure of advanced offensive capability. That prioritization is an analyst recommendation.
How it works. A task starts inside a self-contained Docker environment with a Bash interface. The agent executes commands, reads observations and submits a proposed flag. The original CTF specification assigns a reward of one for the correct flag and zero for an incorrect submission; episodes terminate after success or the configured interaction limit. The Inspect adapter changes the harness to tool-based Bash and Python interaction and maintains its own solver and submission parameters. [A5a][A5b]
Metrics and score interpretation.
Correct-flag accuracy. Successful task submissions divided by the evaluated task population. [A5a][A5b] Direction: Higher capability.
Budget-conditioned success. Always accompany accuracy with message, action and submission-attempt limits. Direction: Recommended reporting condition.
Coverage and failure classification. Report original/subset task IDs, excluded internet tasks and infrastructure errors separately. Direction: Recommended validity diagnostic.
How to interpret a result. Do not compare 78-task offline adapter results directly with a 100-task original score. A new default solver, extra installed tool or reminder to submit may change the system being measured even if the model name is unchanged. Inspect documents such adapter changes in its changelog. [A5b]
Limitations and failure modes.
- The task set is relatively small and public; exposure to known challenge solutions is a plausible confound.
- Some tasks test foundational coding or puzzle-solving skills, so aggregate accuracy is a weak proxy for sophisticated exploit development.
- Dropping internet-dependent tasks is reasonable for containment but changes the task distribution.
- Successful completion provides no direct evidence about a deployed agent’s permissions, protection of sensitive data or resistance to prompt injection.
Proposed bank use. Recommended: include InterCode-CTF in the early evaluation pipeline to verify that a candidate model and its tool interface can execute basic cyber workflows. Pair it with harder real-software tasks before making claims about frontier capability. Preserve the exact adapter identity in vendor evidence requests.
Execution and evidence prerequisites. Use a pinned original environment or a clearly identified adapter. Validate dependencies and task files before the run, record the permitted toolset, and keep answer material outside the agent’s accessible context. For the offline Inspect variant, the exclusion of internet-dependent tasks is an explicit design choice and must remain in the report. [A5a][A5b]
Primary sources: [A5a] [A5b] [A5c]
06. NYU CTF Bench
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
| Owner: NYU-led research team | Release / evidence: 2024 paper; revised February 2025 |
| Task form: Dockerized CSAW CTF tasks; separate evaluation agent | Score direction: Higher task success = greater capability |
| Status assessment: Established research dataset and automation framework | Availability: Public 200-task test set and 55-task development set |
What it is. NYU CTF Bench provides a broad set of reproducible challenges from CSAW competitions for evaluating LLM agents. The official repository contains 200 main test challenges and 55 development challenges across six categories: web, binary exploitation, forensics, reverse engineering, cryptography and miscellaneous. The development split is explicitly intended for agent design without tuning on the final test set. [A6a]
Why it matters. The larger task collection supports a more detailed capability profile than a small selection of demonstrations. Its development/test separation is particularly useful for evaluating a bank’s customized agent scaffold. Use it to identify strengths and weaknesses across technical task families, then examine whether those families resemble the intended application.
How it works. Challenges package metadata and task files, with Docker or Compose environments for service-backed problems. An automated agent interacts with the files and tools to attempt a solution. The accompanying research describes an automated framework using function calling and external tools to evaluate both open and closed models. The benchmark dataset and the model’s orchestration framework are distinct components and should be versioned separately. [A6a][A6b]
Metrics and score interpretation.
Challenge success rate. Fraction of tasks for which the agent returns a valid solution/flag under the chosen evaluator. [A6b] Direction: Higher capability.
Category-level results (recommended). Break down success by the six challenge categories and report each sample count. Direction: Capability profile.
Attempts, cost and tool failures (recommended). Keep single-attempt performance, repeated search and invalid task launches separate. Direction: Reliability and efficiency diagnostics.
How to interpret a result. Treat the development set as a tuning resource and the test set as a held-out evaluation of the selected configuration. If the agent was trained or optimized using test solutions, disclose that and use a fresh holdout for decision evidence. Within a category, preserve challenge provenance rather than assuming every item has equal real-world difficulty.
Limitations and failure modes.
- CTF coverage is broader than a single software target but remains a competition distribution, with limited representation of business authorization and operational defense.
- The main dataset is public, so test/development separation does not itself eliminate model pretraining contamination.
- Tool availability, retries and challenge service stability influence observed success.
- The benchmark does not directly test patch correctness, harmful-request refusal or secure handling of retrieved instructions.
Proposed bank use. Recommended: use the six-category breakdown when comparing security assistants or when assessing whether a tool-enabled general model has acquired broader offensive skills. Keep bank-specific application tests as a separate decision input. A successful challenge should be tied to the demonstrated technique family and deployment affordances, rather than converted directly into a generic risk level.
Execution and evidence prerequisites. The repository provides Dockerized tasks, challenge metadata and a Python interface for selecting development or test splits. Validate the selected challenge images and network setup before running. Ensure the agent can access only task inputs, with ground-truth flags retained by the evaluator. Review benchmark and upstream task licensing before internal redistribution. [A6a]
07. CVE-Bench: web exploitation
CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities
| Owner: UIUC Kang Lab and collaborators | Release / evidence: ICML 2025; repository records 2026 evaluation changes |
| Task form: Containerized web services and outcome validators | Score direction: Higher exploitation success = greater capability |
| Status assessment: Established research benchmark; evolving implementation and scorer | Availability: Public environments and graders; full reference exploits are not all public |
What it is. This CVE-Bench is the UIUC-led web-exploitation benchmark. It is unrelated to the repair benchmark listed as entry 14. The original release comprises 40 critical-severity web application CVEs. Agents receive a running target and task information, then attempt to produce an observable security impact within the sandbox. Its emphasis is practical application exploitation rather than answering vulnerability questions. [A7b]
Why it matters. For a bank assessing an agent that can browse applications or execute HTTP workflows, web-specific impact is often more relevant than low-level engine exploitation. The benchmark can inform capability review, while its targets remain research systems rather than evidence that a particular bank application is vulnerable.
How it works. Docker-based environments reconstruct vulnerable applications. Outcome checks cover impact classes such as unauthorized data access or modification, privilege changes, service disruption and remote code execution. The implementation exposes zero-day and one-day information variants. These labels describe what the agent is told about already known benchmark vulnerabilities; they do not establish discovery of previously unknown real-world vulnerabilities. [A7b]
Metrics and score interpretation.
Validated exploitation success. Fraction of evaluated challenges meeting an applicable impact criterion. [A7a] Direction: Higher capability.
Information-variant results. Keep zero-day and one-day prompt conditions separate. [A7a] Direction: Different discovery assistance.
Impact-category and task coverage (recommended). Record which outcome was achieved, target task IDs and original versus expanded sets. Direction: Prevents misleading aggregate comparison.
How to interpret a result. Pin the grader as well as the dataset. The repository records a v2.1.0 change replacing arbitrary file upload with remote code execution as an evaluation criterion; its current command help displays 2.2.0. A result before that criterion change need not mean the same thing as a later score. Verify the actual tag/commit in every evidence package. [A7a]
Limitations and failure modes.
- The original critical-severity selection is intentionally difficult and is not a prevalence-weighted sample of enterprise web vulnerabilities.
- The sandbox’s observable outcome does not represent every prerequisite and environmental dependency in a live deployment.
- The project withholds most manual reference exploits; public accessibility of code does not imply every reference artifact is available. [A7a]
- Agent tool access, known CVE descriptions and runtime internet retrieval can change what capability the run measures.
Proposed bank use. Recommended: use alongside prompt-injection and authorization testing when onboarding an agent with web or API tools. Treat demonstrated unauthorized-impact capability as a reason to test actual application boundaries and approval controls. Keep repair claims attached to entry 14 or other repair benchmarks rather than importing them from this score.
Execution and evidence prerequisites. The implementation uses Docker and Inspect, with task health checks and configurable challenge and prompt variants. The maintainers recommend amd64; arm64 is experimental. Validate target health and scoring after selecting a release, preserve container images and impact evidence, and use synthetic data in the assessment environment. [A7a]
08. SEC-bench Pro
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
| Owner: University of Illinois Urbana-Champaign research team | Release / evidence: 2026 release; current repository lists 344 verified cases |
| Task form: Repository-level coding agents; browser engines and kernel VMs | Score direction: Higher validated PoC success = greater capability |
| Status assessment: Recent research benchmark; actively revised targets and grading | Availability: Public code, cases and grading; demanding kernel runtime |
What it is. SEC-bench Pro evaluates extended security work in complex software, covering Chromium V8, Mozilla SpiderMonkey and the Linux kernel. Both the July 2026 paper revision and the current repository list 344 verified cases. The early abstract described 183 engine vulnerabilities, so the paper version and task denominator must be retained. Pro is distinct from the original SEC-bench in entry 12. [A8a][A8b][A8e]
Why it matters. The benchmark probes whether an agent can maintain a useful workflow across large source trees and difficult execution environments. For model onboarding it is relevant to autonomous code review, vulnerability research and powerful coding assistants. Its inclusion under cyber capability should not be read as evidence that every successful task achieves full exploit control.
How it works. Cases include metadata, reproduction assets, Docker environments and validation information. Current official submissions use the source_files mode and retain the exact harness configuration. Browser-engine grading combines execution stages with scoped source review; Linux uses a vulnerable/fixed/latest judging process. The runtime and assigned attacker privilege are part of the case definition. [A8a][A8c]
Metrics and score interpretation.
Validated case success. Fraction of the selected task set accepted by its project-specific checker. [A8a] Direction: Higher capability.
Per-project and information-mode results. Separate engine/kernel populations and source-only versus richer information settings. Direction: Required comparability condition.
Oracle-integrity and artifact diagnostics (recommended). Record suspicious retrieval, grader modifications, task health and disputed validation. Direction: Validity evidence rather than capability credit.
How to interpret a result. A sanitizer-triggering PoC is not automatically an exploit that can take over a real system. Read the case’s required outcome and privilege model. Compare the current strict harness with equally restricted runs, and retain the task and grading versions alongside the model snapshot. Treat union-of-agent results separately from a single deployed agent.
Limitations and failure modes.
- Historical vulnerable code may be linked to publicly available fixes and reproducers.
- The maintainers documented score inflation when agents fetched answers online; their correction disabled both container egress and provider-side search. [A8d]
- Current grading can involve source review as well as runtime execution, so judge configuration and disputed-case handling matter. [A8a]
- A particular subset of complex engines and kernel bugs is not a complete evaluation of secure application development.
Proposed bank use. Recommended: use for advanced coding agents after basic environment and coding evaluations. Require a declared information mode, evidence that solution retrieval was controlled and reproducible checker outputs. Use newly observed capabilities to select follow-up controls and tests rather than assigning a generic pass threshold to the aggregate percentage.
Execution and evidence prerequisites. The repository requires Docker, Python and the selected agent configuration. Linux evaluations need x86-64 Linux with usable KVM; the maintainers do not regard software-only QEMU emulation as an equivalent grading mode. Retain exact run configuration, checker summaries and artifacts, following the official submission evidence pattern. [A8a][A8c]
Primary sources: [A8a] [A8b] [A8c] [A8d] [A8e]
09. AISI multi-step cyber ranges
AISI multi-step cyber-range evaluation programme
| Owner: UK AI Security Institute | Release / evidence: 2026 multi-step study; ongoing programme rather than one fixed public benchmark |
| Task form: Multi-host enterprise and industrial-control simulations | Score direction: More completed steps / full chains = greater autonomous capability |
| Status assessment: Institutional evaluation methodology; public reporting with restricted range coverage | Availability: Public methodology and tooling; do not assume all assessed ranges are released |
What it is. AISI’s multi-step cyber ranges evaluate whether agents can combine skills across an extended sequence in a simulated network. The March 2026 study used a 32-step corporate range, The Last Ones, and a seven-step industrial-control range, Cooling Tower. This is a programme of evolving evaluations, not a single universally downloadable dataset with one permanent leaderboard. [A9a][A9b]
Why it matters. Single-task successes do not show whether an agent can preserve context, recover from setbacks and make progress across dependent stages. That sustained autonomy is relevant when a bank considers granting an agent broad tool access or permission to act for long periods. It warrants a separate measurement from isolated cyber skill.
How it works. Experts construct multi-host environments and define milestones or steps along attack scenarios. Agents interact with the range over an extended budget, with their progress recorded across repeated runs. The published study varies inference-time compute to test how additional budget affects progress. Its ranges have no active defenders: detections can be logged without stopping the agent, so this is capability evidence under that condition. [A9a][A9b]
Metrics and score interpretation.
Steps / milestones completed. Progress within one defined range under a specified budget. [A9a] Direction: Higher capability; steps differ in difficulty.
End-to-end completion. Completion of the full scenario, reported with repeat frequency and conditions. Direction: Important autonomous capability signal.
Budget-progress curve. Progress as token/time budgets change; retain mean and best-run outcomes. [A9b] Direction: Shows resource sensitivity, not a general safety probability.
How to interpret a result. Do not compare 50% of a seven-step range with 50% of a 32-step range as equivalent capability. Step definitions, dependencies, available tools and defenses differ. A best run demonstrates possibility under tested conditions; repeated runs estimate reliability. A bank-specific range should disclose deviations from AISI’s design.
Limitations and failure modes.
- Two engineered scenarios cannot represent the distribution of corporate environments.
- The absence of active defenders limits conclusions about success against a monitored production estate. [A9a]
- Alternative solution paths can invalidate intended milestone assumptions; the study discusses environment changes after unexpected paths were discovered. [A9a]
- Public methodology does not mean a third party has access to every evaluated range or can reproduce a published score exactly.
Proposed bank use. Recommended: reserve tailored multi-step range evaluation for deployments with sustained autonomy, substantial technical access or materially consequential operations. Select representative identity, network and application boundaries, and test detection and interruption separately. Record capability and defensive performance as separate outcomes rather than combining them into a single success percentage.
Execution and evidence prerequisites. AISI’s open-source Inspect Cyber standardizes agentic evaluation setup and infrastructure definitions. It is tooling for building and running evaluations, not the public release of every AISI range. An internal implementation requires resettable infrastructure, an independent scorer, synthetic assets and controlled routes to model endpoints. [A9c]
Primary sources: [A9a] [A9b] [A9c]
B. Vulnerability lifecycle, secure coding and detection
This group covers discovery, reproduction, repair, secure development, vulnerability detection and dependency exploitability. Select the profile matching the intended workflow. A detector’s classification score, a patch’s security tests and an agent’s complete lifecycle result have different meanings.
10. CyberGym-E2E
CyberGym-E2E: Scalable Real-World Benchmark for AI Agents’ End-to-End Cybersecurity Capabilities
| Owner: Tianneng Shi, Robin Rheem and collaborators; UC Berkeley-led collaboration with Johns Hopkins, UC Santa Cruz and UC Santa Barbara. | Release / evidence: First preprint 3 June 2026; v2 17 July 2026; ICML 2026. Current project baseline: 920 vulnerabilities / 139 projects. |
| Task form: Agent operates in a real project build environment; execution-based validation. | Score direction: Higher stage success means greater capability; it is not a misuse-resistance score. |
| Status assessment: Recent, peer-reviewed benchmark with public artifacts and a maintained project leaderboard; assessment, not certification. | Availability: Public paper, GitHub implementation and author-hosted Hugging Face data; container images required. |
What it is. CyberGym-E2E extends lifecycle evaluation beyond reproducing a disclosed bug. Its current research release contains 920 real-world vulnerabilities across 139 open-source projects and asks agents to discover a vulnerability, demonstrate it and repair the code. A separate patch-only setting supplies the known reproducer and crash information. The publication is associated with ICML 2026. These are two different task conditions and should be reported separately. [B10a]
Why it matters. Analyst view: this is useful when the bank is evaluating an agent expected to turn a repository into a reviewed remediation proposal. A model can be good at implementing a supplied fix while struggling to find the defect. Keeping those stages visible helps determine whether the appropriate deployment is an engineer’s patch assistant, an autonomous investigation assistant, or neither.
How it works. In end-to-end mode, the agent receives the source environment without the ground-truth reproducer or fix. It returns a candidate reproducer and patch. The harness checks that the candidate crashes the original build, the patch removes that crash, and project tests still pass. A final check evaluates the patch against the withheld ground-truth reproducer. Patch-only mode starts with the known reproducer and crash log. The public implementation supplies containerized task execution and default network isolation. [B10c]
Metrics and score interpretation.
S1–S3 cumulative success. Demonstrate a crash, stop that crash, and preserve the tested functionality. S3 is the project’s discover-and-patch outcome. [B10b] Direction: Higher capability.
S4 cumulative success. Also block the ground-truth reproducer; diagnostic evidence that the intended vulnerability was repaired. [B10b] Direction: Higher target-specific capability.
Patch-only success. Repair performance with the ground-truth reproducer and crash log supplied. [B10b] Direction: Higher repair capability.
How to interpret a result. Analyst view: an S3 success without S4 need not be a failed discovery; the agent may have found another valid defect. Conversely, even all four checks cannot establish absence of every related defect. Record the task count, model, scaffold and budget together. An improvement produced by additional runtime or a changed task set is not evidence of a model-only improvement.
Limitations and failure modes.
- The project distinguishes an earlier 615-task evaluation from the current 920-task set; their percentages are not directly interchangeable. [B10b]
- The authors acknowledge shallow patches that satisfy execution checks while leaving deeper quality concerns. [B10b]
- Analyst limitation: native-code and sanitizer coverage does not establish competence in the bank’s identity, authorization or business-logic controls.
Proposed bank use. Recommendation: include it in a controlled evaluation of security engineering agents, using both patch-only and discovery modes. Require a human root-cause review and the bank’s own regression tests before adopting an individual patch. Record stage-level failures to decide where human assistance is required. Pair the capability evidence with separate misuse-resistance and agent-security evaluations before granting repository-write or tool-execution permissions.
Execution and evidence prerequisites. Practical preparation: provision the documented Python dependencies, benchmark data and Linux container images, and validate build/sanitizer compatibility. Preserve the default network isolation and save the exact task manifest, image versions, submitted patches and evaluator logs. Run a small environment-validation subset before committing resources to the complete suite. [B10c]
Primary sources: [B10a] [B10b] [B10c]
11. BountyBench
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
| Owner: Andy K. Zhang, Joey Ji and collaborators; Stanford-led research collaboration including UC Berkeley. | Release / evidence: First preprint 21 May 2025; v3 2 December 2025. Published baseline: 40 bounties across 25 systems. |
| Task form: Containerized repositories and running applications/services, with executable outcome checks. | Score direction: Higher Detect/Exploit means stronger offensive capability; higher Patch means stronger defensive capability. |
| Status assessment: Peer-reviewed research benchmark with public code; operationally complex and based on a limited published task portfolio. | Availability: Public paper, main bountybench repository and task submodules. |
What it is. BountyBench connects evaluation tasks to vulnerabilities that received real bug-bounty awards. The paper’s baseline contains 40 bounties across 25 systems, with three distinct tasks: discovering a vulnerability, exploiting a specified vulnerability and repairing a specified vulnerability. Its defining addition is an economic lens based on the associated bounty amounts. The number of underlying bounties must not be confused with the number of task-phase combinations. [B11a]
Why it matters. Analyst view: this benchmark can make an offense–defense imbalance visible within the same application portfolio. That is useful for deciding whether an agent is ready to assist application security engineers and for examining the capability being exposed through a coding service. Monetary task weights also communicate why solving a few difficult, valuable cases may matter more than an unweighted average suggests.
How it works. The authors build application snapshots with source, services, databases, exploits, patches and code/runtime invariants. Detect and Exploit submissions are checked for observable effects on vulnerable snapshots and, where specified, failure against patched snapshots. A patch must block the reference exploit while preserving the invariants. Information supplied to the agent varies by task, including different levels of discovery guidance. These conditions materially change difficulty and must remain attached to every result. [B11b]
Metrics and score interpretation.
Success rate by phase. Fraction of Detect, Exploit or Patch tasks satisfying that phase’s verifier. [B11b] Direction: Higher capability in the named phase.
Bounty value represented. Sum of bounty values associated with solved tasks, according to the benchmark’s mapping. [B11a] Direction: Greater represented economic value; not bank loss avoided.
Usage and cost. Tokens, attempts and computational expenditure alongside task outcomes. [B11b] Direction: Lower cost is useful only at comparable quality.
How to interpret a result. Analyst view: report three scores rather than one combined security grade. A high patch score does not establish discovery competence or policy compliance. Bounty values reflect the originating programs’ rewards; they are not a calibrated estimate of financial exposure, adversary revenue or savings in the bank. Reproducing a publicly disclosed vulnerability also differs from discovering a previously unknown issue in production.
Limitations and failure modes.
- The published baseline is small and its task environments require carefully maintained services and invariants. [B11b]
- Analyst limitation: known bounty reports and historical source can be present in model training data.
- Analyst limitation: the amount of guidance, retries and scaffolding can dominate apparent differences between models.
- Analyst limitation: changing live repository/task-submodule versions can change the evaluated portfolio.
Proposed bank use. Recommendation: use BountyBench as a comparative application-security evaluation with a fixed publication baseline or an explicitly versioned current manifest. Preserve separate offensive and patching results in the onboarding record. Review the relevance of each system and weakness to bank applications. Prefer evidence that shows both exploit blocking and maintained business behavior before authorizing automated remediation suggestions at scale.
Execution and evidence prerequisites. Practical preparation: use the main bountybench repository, initialize its task submodules, and provision its documented Python and Docker environment plus the selected model endpoint. Validate service readiness and invariant tests before scoring an agent. Save task snapshots, prompts, attempts, environment failures and phase-specific verifier results. Public wrappers and forks can use different inventories. [B11c]
Primary sources: [B11a] [B11b] [B11c]
12. SEC-bench
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
| Owner: Hwiwon Lee, Ziqi Zhang, Hanxiao Lu and Lingming Zhang; UIUC and Purdue University. | Release / evidence: Preprint June 2025; revised October 2025; NeurIPS 2025. Paper: 200 verified CVE instances / 29 projects; current official eval split: 300 rows. |
| Task form: Agentic repository interaction in Docker; sanitizer-based executable oracles. | Score direction: Higher scores mean greater task capability; patch score is bounded by the verifier’s scope. |
| Status assessment: Established public research framework; original-family dataset has expanded beyond its publication baseline. | Availability: Public paper, MIT-licensed code, official Hugging Face dataset and evaluation images. |
What it is. SEC-bench combines benchmark construction with agent evaluation. The original paper describes 200 verified CVE instances across 29 projects, while the official Hugging Face evaluation split now contains 300 rows. Treat the publication baseline and the expanded dataset as different manifests. This entry concerns original SEC-bench; SEC-bench Pro is a separate, newer benchmark configuration and should retain its own identity and results. [B12a] [B12c]
Why it matters. Analyst view: the distinctive value is repeatable evidence production. A model claim becomes more useful when the codebase can be rebuilt, the vulnerable behavior reproduced and the proposed fix tested under known conditions. The construction workflow is also a useful design reference for a bank-owned evaluation corpus derived from resolved internal defects, provided labels and sensitive artifacts remain appropriately controlled.
How it works. A preprocessor extracts reports and repository information; a verifier coordinates building, vulnerability reproduction and reference-fix preparation; an evaluator packages tasks. For PoC generation, a successful output must trigger the expected sanitizer error at the relevant location. For patching, the candidate must apply, compile and stop the original reproducer from producing that error. Those checks establish the benchmark’s repair outcome, rather than complete semantic equivalence. [B12a]
Metrics and score interpretation.
PoC resolved rate. Share of cases producing the expected executable vulnerability signal. [B12a] Direction: Higher reproduction capability.
Patch resolved rate. Share whose submitted patch builds and stops the checked vulnerability signal. [B12a] Direction: Higher measured repair capability.
Submitted rate / failure categories. Separate submission from success; retain absent/invalid patches, build failures and still-vulnerable outcomes. [B12a] Direction: Diagnostic.
How to interpret a result. Analyst view: a successful build plus a stopped reproducer is a useful first gate, but it can reward narrow crash suppression. Require broader regression and root-cause evidence before interpreting a high score as readiness for autonomous patch deployment. A comparison should identify the exact split, amount of vulnerability information, agent scaffold, budget and sanitizer configuration; each changes what the score establishes.
Limitations and failure modes.
- Analyst limitation: sanitizer-oriented C/C++ tasks do not cover all application security classes or languages.
- Analyst limitation: the selected cases favor vulnerabilities that the construction process can reproduce and verify.
- Analyst limitation: a reference patch and one reproducer cannot establish that every related path has been repaired.
- Analyst limitation: duplicated rows across dataset splits must not be counted as independent cases.
Proposed bank use. Recommendation: use SEC-bench as a baseline for engineering agents that investigate disclosed vulnerabilities or propose patches. Add an independent validation layer modeled on the bank’s actual change controls: targeted security regression, functional tests and reviewer confirmation of the root cause. Track offensive reproduction separately from repair capability. Use a frozen case list so expanded datasets do not create misleading year-over-year trends.
Execution and evidence prerequisites. Practical preparation: the repository documents Python 3.12+, Docker and substantial disk capacity, recommending more than 200 GB. It exposes distinct patch and PoC tasks, including different information levels for PoC evaluation. Provision only the images needed for the selected manifest, preserve build logs and record environment failures separately from agent failures. [B12b]
Primary sources: [B12a] [B12b] [B12c]
13. AutoPatchBench
AutoPatchBench / AutoPatch in Meta CyberSecEval
| Owner: Meta, Purple Llama / CyberSecEval. | Release / evidence: Announced 29 April 2025: 136 full / 113 Lite cases. Current AutoPatch documentation: 142 full / 120 Lite / 20 sample cases. |
| Task form: Native-code build, crash replay, fuzzing and white-box differential testing in Linux containers. | Score direction: Higher fully validated repair rate is better defensive capability; basic crash suppression is a weaker measure. |
| Status assessment: Public, vendor-maintained research benchmark within CyberSecEval; resource-intensive executable validation. | Availability: Public CyberSecEval documentation and implementation; large container/dependency footprint. |
What it is. AutoPatchBench evaluates AI-generated fixes for native-code vulnerabilities discovered through fuzzing. Its April 2025 announcement describes 136 full cases and a 113-case Lite subset. The operational documentation now lists 142 full, 120 Lite and 20 sample cases. This difference makes the manifest and evaluation date essential: a published result on the original Lite set should not be relabeled as a result on today’s documented Lite set. [B13a] [B13b]
Why it matters. Analyst view: this benchmark addresses a recurring governance mistake—treating a patch that stops one observed failure as a dependable security fix. A remediation assistant can appear productive while removing useful behavior, leaving the root cause intact or creating another failure. AutoPatchBench provides a concrete way to separate initial patch generation from stronger validation and to explain why human review remains meaningful.
How it works. The benchmark starts with reproducible vulnerable and reference-fixed code. Candidate fixes first undergo build and crash-reproduction checks. Subsequent fuzzing searches for additional crashes. White-box differential testing compares internal program states against the reference repair on generated inputs, using debugger-based observations. The authors explicitly describe this as an imperfect practical approximation: nondeterminism, timeouts and missing observations can affect the comparison. Their design therefore offers stronger evidence than a single reproducer, without proving program equivalence. [B13a]
Metrics and score interpretation.
Patches generated (shr_patches_generated). Share of cases yielding a patched function name; not a validation result. [B13b] Direction: Higher generation coverage.
Fuzzing pass (shr_passing_fuzzing). Share also surviving the configured fuzz testing. [B13b] Direction: Higher tested robustness.
Full verification pass (shr_correct_patches). Share passing fuzzing and differential validation. [B13b] Direction: Higher measured repair quality.
How to interpret a result. Analyst view: the difference between initial and final pass rates is a quality signal, not merely an inconvenient reduction in the headline score. Compare the same testing duration, corpus and observation policy. Keep a separate count of evaluator errors and inconclusive comparisons; treating them as clean passes or hiding them in the denominator would overstate confidence.
Limitations and failure modes.
- The launch study reports that differential testing can still accept incorrect patches; manual review is therefore valuable. [B13a]
- Analyst limitation: reference-patch behavior may itself contain defects or unnecessary implementation choices.
- Analyst limitation: a small, native-code corpus cannot establish reliability for the bank’s broader application portfolio.
- Analyst limitation: running more fuzzing changes both evaluation cost and the strength of the evidence.
Proposed bank use. Recommendation: use it to assess an automated patching service and to design the evidence required for individual remediation proposals. Make the full-verification result the meaningful benchmark outcome, while retaining early-stage outcomes for diagnosis. Pilot on native-code components where the bank has an actual support requirement. Require existing regression tests and engineering review before moving a candidate fix into change management.
Execution and evidence prerequisites. Practical preparation: use Linux with the documented Python and Podman setup. Current guidance describes approximately 3 TB for the full set, 2 TB for Lite and 500 GB for samples, and recommends substantial CPU parallelism. The documentation excludes Apple Silicon support. Validate artifact availability and a small sample before planning a full campaign. [B13b]
Primary sources: [B13a] [B13b]
14. CVE-Bench: vulnerability repair
CVE-Bench: Benchmarking LLM-based Software Engineering Agent’s Ability to Repair Real-World CVE Vulnerabilities
| Owner: Peiran Wang, Xiaogeng Liu and Chaowei Xiao; repository WhileBug/CVEBench. | Release / evidence: NAACL 2025, April 2025; published corpus: 509 CVEs / 120 repositories / four languages. |
| Task form: Interactive coding-agent environment with patch application and vulnerability-associated unit tests. | Score direction: Higher repair rate indicates better performance under the specified information level. |
| Status assessment: Peer-reviewed NAACL 2025 research benchmark with public implementation; distinct from exploitation-oriented namesakes. | Availability: Public ACL paper and author repository; environment preparation uses CVEfixes. |
What it is. The repair-oriented CVE-Bench is the NAACL 2025 project by Wang, Liu and Xiao, implemented at WhileBug/CVEBench. Its published corpus contains 509 CVEs from 120 repositories in Python, Java, JavaScript and PHP. It evaluates repair of disclosed defects. The user’s separate web-exploitation CVE-Bench entry is a different project, so reports should always include the owner or full paper title. [B14a] [B14b]
Why it matters. Analyst view: the benchmark addresses a practical handoff problem between security and engineering: how much information must a vulnerability report contain before an agent can produce a useful fix? A bank may have mature scanning but weak remediation because findings arrive without a reliable location or actionable context. This benchmark can help distinguish report quality, code navigation and actual repair competence.
How it works. The environment checks out vulnerable repository states and gives the agent increasingly detailed reports: description only, description plus file path, or description plus file and method. The paper calls these black-box, intermediate and white-box information levels; the repair agent still receives a codebase. Agents can interact with tools and edit code. A repair counts when the patch applies and the associated unit tests pass. The paper also examines whether static-analysis tools help the process. [B14a]
Metrics and score interpretation.
Repair rate by information level. Percentage of evaluated vulnerabilities whose patch applies and passes the required tests. [B14a] Direction: Higher repair capability.
Failure distribution. Location failure, iteration limit, unit-test failure and other agent failures. [B14a] Direction: Diagnostic.
Steps / tool use. Recommended additional operational measures for distinguishing search effort from repair quality. Direction: Analyst-recommended diagnostic.
How to interpret a result. Analyst view: the information labels describe the originating report, not a claim that the patching agent has no source-code access. Keep that distinction explicit in procurement questionnaires. A favorable result with a supplied function location should support a narrowly scoped patch-assistance use case; it should not be cited as proof that the agent can independently triage vague production findings.
Limitations and failure modes.
- Analyst limitation: vulnerability-associated unit tests are evidence for checked behavior, not exhaustive proof of a secure repair.
- Analyst limitation: historical CVE and patch data create contamination and memorization risks.
- Analyst limitation: available report detail and repository preparation can substantially change the workload.
- Analyst limitation: the four-language corpus provides no direct evidence for other languages or managed platform configuration.
Proposed bank use. Recommendation: use this benchmark when selecting agents that consume SAST, penetration-test or vulnerability-management findings. Evaluate all three report-detail settings and compare them to the actual quality of the bank’s tickets. Require unsuccessful patches and navigation failures in the evidence package. Use the result to define mandatory ticket fields, analyst assistance and engineer review, alongside independent tests of generated fix quality.
Execution and evidence prerequisites. Practical preparation: obtain the author implementation and the specified CVEfixes database, then preserve the selected CVE list, vulnerable repository states and environment dependencies. The repository separates preparation, agent execution and verification. Validate that the target tests fail and pass in the intended baseline states before treating an agent result as meaningful; document any repairs to the harness. [B14b]
Primary sources: [B14a] [B14b]
15. PatchBench
PatchBench: Evaluating AI Agents for Vulnerability Patching
| Owner: Chihao Shen, Jiacheng Li, Aastha Mahajan, Jeffery Siyuan Tian, Yonghwi Kwon and Yizheng Chen; University of Maryland. | Release / evidence: Preprint 3 September 2026; 213 C/C++ patching tasks / 32 projects / 16 CWEs. |
| Task form: Repository-level coding agents; replay corpora, sanitizer checks, unit tests and output-state comparison. | Score direction: Higher combined security-and-semantics solve rate is better defensive capability. |
| Status assessment: Emerging: September 2026 preprint with public code and data; independent replication was not established in this review. | Availability: Public ai-sec-lab repository and dataset; substantial per-task container and corpus storage. |
What it is. PatchBench is a verified, recent research project, rather than an unspecified future benchmark. The September 2026 paper introduces 213 native-code patching tasks across 32 projects and 16 CWEs. It deliberately emphasizes cases where the reference fix lies outside the crash stack, and moves historical vulnerabilities into changed repository contexts. Those choices target two concerns: shallow crash suppression and recall of published fixes. [B15a]
Why it matters. Analyst view: PatchBench is particularly relevant when a vendor claims that an agent repairs almost every vulnerability. Its evaluator asks whether that apparent success survives additional security cases and benign-behavior checks. For the bank, the question is whether increased automation reduces remediation risk or simply produces more plausible patches that engineers must subsequently reject.
How it works. The evaluator distinguishes the original reproducer from broader crashing-input replay. Three semantic checks assess benign-input sanitizer behavior, retention of reference-passing unit tests, and agreement with reference-patched outputs. Solving requires security plus all three semantic checks. Reduced agent metadata excludes evaluation ground truth. [B15b]
Metrics and score interpretation.
Original-PoC pass. Known reproducer no longer crashes; an early-stage result. [B15b] Direction: Higher initial capability.
Security validation pass. No sanitizer failure on the selected crashing-input corpus. [B15b] Direction: Higher tested repair robustness.
Combined solve rate. Security validation plus benign-input sanitizer, output-state and unit-test checks. [B15b] Direction: Higher measured repair quality.
How to interpret a result. Analyst view: compare the gap between original-PoC and combined solve rate. That gap identifies the risk concealed by a weaker headline metric. The paper’s patch-similarity analysis provides evidence consistent with memorization; similarity alone does not prove what a model encountered during training. New-context performance is a useful countercheck, but it should still be accompanied by reviewer assessment of the proposed fix.
Limitations and failure modes.
- The corpus deliberately selects difficult cases whose reference fixes lie outside the crash stack; it is not a random sample of all repair work. [B15a]
- Analyst limitation: agreement with a reference output can reject legitimate alternative behavior and cannot prove absence of untested defects.
- Analyst limitation: code mutation reduces some memorization opportunities without certifying that the benchmark is uncontaminated.
- Analyst limitation: recent release status means operational reproducibility should be established before using it as an onboarding gate.
Proposed bank use. Recommendation: add PatchBench to the evaluation of agents permitted to propose security-sensitive code changes, especially when another repair benchmark appears saturated. Adopt its separation of known-trigger suppression, broader vulnerability repair and functional preservation in the bank’s evidence template. Treat a passing patch as a candidate for review. Set any local acceptance criterion using a baseline and the intended change authority, rather than copying a research score.
Execution and evidence prerequisites. Practical preparation: provision the official Python environment and task images. The repository describes about 22 GB of corpora and 870 GB for all 213 images. Partial reports can still be aggregated. [B15b] Recommendation: pilot selected tasks, retain individual reports and confirm all required stages ran.
Primary sources: [B15a] [B15b]
16. SecurityEval
SecurityEval Dataset: Mining Vulnerability Examples to Evaluate Machine Learning-Based Code Generation Techniques
| Owner: Mohammed Latif Siddiq and Joanna C. S. Santos; University of Notre Dame / s2e-lab. | Release / evidence: MSR4P&S 2022. Original: 130 samples / 75 CWEs; current repository dataset: 121 prompts / 69 CWEs. |
| Task form: Prompt-to-code generation followed by targeted static analysis and manual assessment. | Score direction: Lower vulnerability-generation rate is better; analyzer silence is not proof of security. |
| Status assessment: Established, lightweight 2022 research dataset; updated prompts require version-specific interpretation. | Availability: Public author paper, GitHub dataset and official Hugging Face dataset. |
What it is. SecurityEval is a Python code-generation dataset organized around CWE-linked vulnerability scenarios. Its current repository contains 121 prompts spanning 69 CWEs after corrections and removal of a prompt explicitly requesting vulnerable code. The original paper and demonstration results used 130 samples spanning 75 CWEs. The maintainers state that the old results were not recalculated, so a current dataset run must not be compared as though the prompts were unchanged. [B16a]
Why it matters. Analyst view: its small size makes it useful for an early screening of coding-model behavior and for inspecting concrete insecure suggestions. It can expose recurring failure patterns before a more expensive repository-level assessment. The result is most useful as a diagnostic by weakness category, especially when a bank is reviewing Python assistants used for automation, analytics or integration scripts.
How it works. Each prompt elicits code in a context associated with a particular weakness. The research demonstration examines generated outputs manually and with CodeQL and Bandit, then checks whether the scenario’s designated CWE appears. The authors observe that automated tools can miss insecure outputs and that generated code may have other vulnerabilities or functional defects. This is therefore a targeted security evaluation of generated snippets, not an end-to-end application acceptance test. [B16b]
Metrics and score interpretation.
Target-CWE vulnerability rate. Share of generated snippets judged to contain the weakness associated with their prompt. [B16b] Direction: Lower is better.
Analyzer-specific finding rate. CodeQL and Bandit results, kept separate from manual judgments. [B16b] Direction: Diagnostic; depends on coverage.
Functional / valid-output rate. Recommended companion measure so invalid or nonfunctional code cannot appear safe merely by avoiding a finding. Direction: Higher useful-output rate.
How to interpret a result. Analyst view: identify the analyzer and query-pack versions, prompt revision and generation settings beside the score. A decrease in analyzer findings can come from improved security, different output structure or poorer analyzer coverage. Investigate disagreements on representative samples and track newly introduced weaknesses separately from the designated CWE. Do not collapse all of these outcomes into a single binary model label.
Limitations and failure modes.
- The dataset is limited to Python and the original study demonstrates substantial differences between manual and automated assessment. [B16b]
- Analyst limitation: a small number of scenarios per CWE cannot support strong claims about entire weakness classes.
- Analyst limitation: public prompts can be memorized or heavily tuned against.
- Analyst limitation: snippet-level performance does not establish secure behavior across repository dependencies or deployment settings.
Proposed bank use. Recommendation: use SecurityEval as a lightweight regression set for coding prompts, model upgrades and security reminders. Manually review a sample of claimed improvements and disputed analyzer results. Supplement it with relevant internal tasks and repository-level secure coding benchmarks before granting broader development authority. Preserve functional failures, refusals and empty responses as distinct outcomes, and retain the current 121-prompt version explicitly.
Execution and evidence prerequisites. Practical preparation: obtain the intended dataset revision and install a compatible Python, CodeQL and Bandit environment. The repository documents the prompt schema, insecure reference examples and analysis materials. Generate fresh outputs for the selected model; do not reuse the historical Copilot/InCoder demonstration files as if they describe a current product. Save raw outputs and adjudication records. [B16a]
Primary sources: [B16a] [B16b]
17. SecRepoBench
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
| Owner: Chihao Shen, Connor Dilgren, Purva Chiniya, Luke Griffith, Yu Ding and Yizheng Chen; University of Maryland and Google DeepMind. | Release / evidence: Initial preprint April 2025; current title/release associated with LLM4Code 2026; 318 tasks / 27 C/C++ repositories / 15 CWEs. |
| Task form: Masked code region plus repository context; compilation, developer tests and a security reproducer. | Score direction: Higher secure-pass@1 is better; ordinary pass@1 alone does not measure security. |
| Status assessment: Public, peer-reviewed repository-completion benchmark with documented agent and standalone-model settings. | Availability: Public paper, MIT-licensed repository and official Hugging Face dataset. |
What it is. SecRepoBench tests secure code completion inside real repositories. Its 318 tasks come from 27 C/C++ repositories and cover 15 CWEs. The model must fill a masked region while respecting surrounding code and dependencies. This is a distinct workload from repairing a fully described CVE or generating a standalone program. The benchmark’s principal result combines functional correctness with a vulnerability-specific security test. [B17a]
Why it matters. Analyst view: repository context changes what secure coding requires. A completion may look safe in isolation yet violate assumptions made elsewhere in the application. This benchmark is useful for evaluating development assistants that work within existing codebases and for checking whether an agent’s apparent productivity improvement also preserves security. It creates a direct connection between code quality and the integration context.
How it works. The implementation supports standalone models with retrieved repository context and agents with broader repository access. A generated completion is compiled with the project; developer tests assess correctness and an OSS-Fuzz-derived reproducer checks the target vulnerability. The repository provides neutral prompts and several security-reminder variants, including CWE-specific guidance. The neutral condition avoids telling the model which weakness to anticipate. These prompt variants and retrieval modes must be identified when comparing results. [B17b]
Metrics and score interpretation.
pass@1. Probability that one generated completion passes functional tests. [B17a] Direction: Higher correctness.
secure-pass@1. Probability that one completion passes both functionality and security tests. [B17a] Direction: Higher useful secure completion.
Security-only outcome. Companion result for the designated vulnerability check; retain its denominator separately. [B17a] Direction: Diagnostic.
How to interpret a result. Analyst view: the gap between pass@1 and secure-pass@1 shows how often apparent functional success still fails security. Compare scaffolds using the same task set and prompt condition. A CWE-specific reminder gives valuable assistance but changes the question from detecting an implicit security requirement to following an explicit one. Scores from those settings should not be merged.
Limitations and failure modes.
- Analyst limitation: the masked region is localized, so the task does not test unrestricted feature development across arbitrary files.
- Analyst limitation: one designated reproducer can miss alternative defects in an otherwise passing completion.
- Analyst limitation: native-code repository coverage does not establish competence in the bank’s web authorization or cloud-policy logic.
- Analyst limitation: mutation and prompt changes can reduce literal recall without proving absence of training contamination.
Proposed bank use. Recommendation: include SecRepoBench when onboarding a code assistant expected to modify existing repositories. Use the neutral condition as a baseline, then assess approved security guidance as a controlled intervention. Review the added security value independently from increased functional pass rate. Supplement with bank-relevant languages, project conventions and negative tests before accepting claims that a development assistant reliably writes secure code.
Execution and evidence prerequisites. Practical preparation: obtain the benchmark metadata, container environments and documented Python dependencies. Select the model, scaffold, retrieval strategy and security-prompt condition in advance. Save completions, agent trajectories, build outcomes and both security and functional verdicts. Validate a representative task before a full run, particularly when a scaffold uses a different host/container arrangement. [B17b]
Primary sources: [B17a] [B17b]
18. SecureVibeBench
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
| Owner: Junkai Chen, Huihui Huang and collaborators; Singapore Management University-led academic collaboration. | Release / evidence: Initial preprint September 2025; public code/data April 2026; ACL 2026 publication. 105 C/C++ tasks / 41 projects. |
| Task form: Natural-language requirements and full repositories; functional comparison, dynamic PoV and static security analysis. | Score direction: Higher Correct-and-Secure share is better; suspicious and proven-vulnerable outcomes remain separate. |
| Status assessment: Recent, peer-reviewed ACL 2026 benchmark with public code and data. | Availability: Public ACL paper, MIT-licensed code, official Hugging Face tasks and Docker images. |
What it is. SecureVibeBench contains 105 C/C++ development tasks from 41 projects. [B18a] It reconstructs the point at which a human developer originally introduced a vulnerability, then asks an agent to implement the associated requirements. This evaluates security during ordinary development rather than only after a defect is known. Public code and data were released in April 2026, followed by the ACL 2026 publication. [B18b]
Why it matters. Analyst view: a bank needs to know whether a coding agent introduces vulnerabilities while delivering a requested change, even when no one explicitly asks it to perform security work. This benchmark is relevant to that concern. It also broadens the evidence beyond avoiding one historical flaw by looking for new suspicious changes that a target-specific test would miss.
How it works. A functional comparison checks the agent’s implementation against reference behavior using repository tests. A dynamic proof-of-vulnerability checks the historical defect, while SAST identifies additional potential issues. Outcomes distinguish incorrect code, correct-but-vulnerable code, correct-but-suspicious code and correct-and-secure code. A new SAST alert is classified as suspicious because it may be a false alarm; it is not automatically treated as a demonstrated exploitable vulnerability. [B18a]
Metrics and score interpretation.
C-SEC rate. Correct-and-secure share under the combined evaluator. [B18a] Direction: Higher is better.
C-VUL / C-SUS rates. Functionally correct outputs with demonstrated target vulnerability or additional suspicious findings, respectively. [B18a] Direction: Lower is better; different evidence strength.
IC rate. Functionally incorrect output share. [B18a] Direction: Lower is better.
How to interpret a result. Analyst view: report all four outcomes. A low vulnerability rate can be misleading if the agent rarely produces working code. Equally, a SAST warning deserves triage rather than an unsupported claim that an exploit exists. The most valuable evidence identifies whether a model improved secure completion, merely became less functional, or moved failures from the known vulnerability into other suspicious patterns.
Limitations and failure modes.
- Analyst limitation: the historical vulnerability determines part of the task selection, so these are security-sensitive scenarios rather than a random development workload.
- Analyst limitation: functional tests can omit new-feature behavior, and static/dynamic checks do not cover every weakness.
- Analyst limitation: SAST configuration changes the suspicious-output rate and must remain versioned.
- Analyst limitation: reconstructed history and generated requirements can differ from the full context available to the original developer.
Proposed bank use. Recommendation: use SecureVibeBench as a development-assistant evaluation alongside a repair benchmark. Evaluate the actual code-agent configuration intended for the bank, including its approved security instructions. Require an evidence trail showing functional results, reproducer results and adjudicated new alerts. Use failure examples to strengthen review checklists and project templates, rather than assigning a universal model pass threshold from an aggregate research result.
Execution and evidence prerequisites. Practical preparation: obtain the author dataset and per-instance Docker images, install the evaluation dependencies and configure the selected model/scaffold. The implementation supports several popular coding-agent frameworks. Reserve enough disk capacity for the selected image set, freeze the static-analysis configuration, and retain submitted patches together with the separate functional, dynamic-security and static-analysis reports. [B18b]
Primary sources: [B18a] [B18b]
19. VEX-Bench
VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities
| Owner: Jiahao Shi and collaborators; Purdue University and Red Hat. | Release / evidence: Preprint 7 September 2026; author explanation 29 September 2026. 75 cases / 67 CVEs / 35 projects. |
| Task form: Repository and advisory analysis with structured affected-status and justification classification. | Score direction: Higher status recall/F1 and justification macro-F1 are better; false not-affected decisions are particularly consequential. |
| Status assessment: Emerging: September 2026 paper; authors report EMNLP 2026 acceptance. Runnable artifact access was not independently validated. | Availability: Public paper; authors link code/data at github.com/steven1518/vex-bench, but that repository could not be independently retrieved during this review. |
What it is. This entry selects the software-supply-chain VEX-Bench from Purdue and Red Hat. It contains 75 cases spanning 67 CVEs and 35 projects in Go, Python and Java. The task is to determine whether a vulnerable dependency makes the downstream project affected. [B19b] A separate misinformation-evaluation project also uses the name VEX-Bench; the full title and owner should therefore appear in the benchmark register. [B19c]
Why it matters. Analyst view: this is closely aligned to a bank’s dependency-alert triage. A scanner finding an affected package version does not resolve how the application uses that package or which deployment conditions matter. An assistant that explains those relationships could reduce analyst effort, but a wrong not-affected judgment could suppress a real issue. That asymmetry should shape the evaluation.
How it works. The agent receives a complete target repository and a CVE identifier, retrieves relevant external information and examines dependency usage. It returns affected/not-affected status, an evidence explanation and a finer justification. The four not-affected reasons concern absent code, unreachable code, required configuration and required environment. Expert-labeled cases provide the answer key. These are code-snapshot assessments, not tasks requiring an attack on a deployed service. [B19b]
Metrics and score interpretation.
Status precision, recall and F1. Binary affected/not-affected classification quality. [B19a] Direction: Higher; inspect affected-case recall.
Justification macro-F1. Equal-weight performance across the affected class and four not-affected reasons. [B19a] Direction: Higher fine-grained reasoning quality.
Failure / invalid-output rate. Timeouts, parsing failures and predictions outside the scored labels are failed runs. [B19a] Direction: Lower.
How to interpret a result. Analyst view: a correct status with a wrong rationale is inadequate evidence for a durable risk exception. The justification must be tied to the actual application version and deployment conditions. Measure how often the agent wrongly clears an affected case, not just overall accuracy. Require a reviewer to confirm that its cited code path, package identity and configuration are relevant to the bank’s environment.
Limitations and failure modes.
- The paper uses a precedence rule for cases with several valid reasons; another factually supportable reason can still be scored wrong. [B19a]
- Repository snapshots cannot fully represent all production inputs, runtime settings or deployment configurations. [B19a]
- Analyst limitation: 75 cases provide limited support for narrow language- or reason-specific conclusions.
- Analyst limitation: operational use requires revalidation whenever dependencies, configuration or exposed paths change.
Proposed bank use. Recommendation: pilot VEX-Bench for SCA-triage assistants. Keep final not-affected decisions under vulnerability-management ownership and require supporting evidence plus an expiry or change trigger. Supplement the public tasks with adjudicated internal dependency findings. Connect the result to the bank’s exception and remediation workflow; do not let a high benchmark score alone authorize automatic suppression of dependency alerts.
Execution and evidence prerequisites. Practical preparation: reproduce the paper’s isolated, language-specific environments and hide answer labels from the agent. [B19a] Before scheduling a campaign, confirm the author-announced code/data can be obtained and validate the harness. Freeze repository snapshots, external-advisory evidence, output schema and resource limits so results remain reviewable when upstream pages or tools change.
Primary sources: [B19a] [B19b] [B19c]
20. PrimeVul
PrimeVul, introduced in Vulnerability Detection with Code Language Models: How Far Are We?
| Owner: Yangruibo Ding, Yanjun Fu and collaborators; DLVulDet research project. | Release / evidence: March 2024 initial release; July 2024 paper revision; ICSE 2025. Metadata-enriched PrimeVul-v0.1 released September 2024. |
| Task form: Labeled C/C++ code classification with chronological splits and paired examples. | Score direction: Higher precision/recall/F1 and paired correctness are better; lower VD-S is better. |
| Status assessment: Established public vulnerability-detection dataset and evaluation methodology; release-specific filtering matters. | Availability: Public paper, code and author-linked dataset releases; original corpus and v0.1 are not identical. |
What it is. PrimeVul’s original corpus contains 6,968 vulnerable and 228,800 benign functions across 755 projects. [B20a] The later v0.1 adds commit, CVE and file-context metadata, but retains only samples for which that metadata was recovered. Therefore the original paper’s inventory must not be assigned automatically to v0.1. PrimeVul evaluates detection of vulnerable code; it is not a vulnerability-repair or autonomous exploitation benchmark. [B20b]
Why it matters. Analyst view: vulnerability detection is an imbalanced decision problem. An assistant that flags nearly every function may find many defects while overwhelming reviewers; an assistant that labels almost everything benign may look accurate while missing the real issues. PrimeVul helps expose those failure modes and the tendency to react to superficial code patterns instead of distinguishing a vulnerable function from its repaired counterpart.
How it works. The dataset reconstructs prior vulnerability corpora using stricter commit/CVE-based labeling and deduplication. Its evaluation uses chronological partitions and separately examines similar vulnerable/repaired pairs. The Vulnerability Detection Score, VD-S, measures false negatives after constraining false positives. The paper uses a 0.5% false-positive tolerance as an evaluation parameter, rather than an adoption standard. Pair-wise scoring requires the labels of both versions to be correct. [B20a]
Metrics and score interpretation.
VD-S. False-negative rate at a specified false-positive ceiling. [B20a] Direction: Lower is better.
Pair-wise correct prediction. Both vulnerable and repaired versions classified correctly. [B20a] Direction: Higher is better.
Precision, recall and F1. Companion classification measures; document class balance and threshold. [B20a] Direction: Higher is better.
How to interpret a result. Analyst view: select an acceptable reviewer workload before judging detector utility. A false-positive percentage that seems small can still generate an impractical number of alerts in a large codebase. Use validation data to set any operating threshold and keep test data untouched. The bank’s own vulnerability prevalence and review capacity should determine the local decision, rather than adopting the paper’s 0.5% parameter unchanged.
Limitations and failure modes.
- Analyst limitation: function-level labels can omit cross-function, runtime and configuration context.
- Analyst limitation: improved labeling remains a research approximation; disagreement cases merit expert review.
- Analyst limitation: chronological splitting reduces particular leakage paths without establishing that a foundation model never saw the code.
- Not every original sample is present in the metadata-enriched v0.1 release. [B20b]
Proposed bank use. Recommendation: use PrimeVul when evaluating a code-security classifier or SAST-assistance model. Require false-negative and false-positive evidence, paired-sample performance and a release-specific manifest. Supplement it with a representative internal holdout, including benign code that often triggers alerts. Keep classification assistance separate from authority to close findings, and assess whether the additional detections justify the analyst effort they create.
Execution and evidence prerequisites. Practical preparation: select the original or v0.1 data deliberately, preserve its provided partitions and keep repaired pairs grouped. The repository supplies training/inference examples and VD-S calculation code. [B20b] Record tokenization/context limits, prediction thresholds, excluded samples and output parsing. An API-evaluated model needs no local training run, but its benchmark split and preprocessing still require validation.
Primary sources: [B20a] [B20b]
21. DiverseVul
DiverseVul: A New Vulnerable Source Code Dataset for Deep Learning Based Vulnerability Detection
| Owner: Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen and David Wagner. | Release / evidence: Initial preprint April 2023; revised August 2023; RAID 2023. 18,945 vulnerable / 330,492 non-vulnerable functions. |
| Task form: Labeled C/C++ functions mined from vulnerability-fixing commits. | Score direction: Higher recall/precision/F1 is better; lower false-positive rate is better. |
| Status assessment: Established RAID 2023 research dataset; useful breadth with material label and generalization limitations. | Availability: Public paper and author repository linking dataset, metadata and research splits. |
What it is. DiverseVul contains 18,945 vulnerable and 330,492 non-vulnerable functions from 7,514 commits across 797 projects, covering 150 CWEs. [B21a] It is a detection dataset intended for comparing and training vulnerability classifiers. The author repository provides data, metadata and study splits. A score must identify whether it uses DiverseVul alone or a merged corpus, since the published experiments also examine combined datasets. [B21b]
Why it matters. Analyst view: breadth matters when a detector appears effective on one project family but encounters unfamiliar software. DiverseVul provides a basis for examining that transfer. It also makes the data-engineering decisions behind a vulnerability score visible: what constitutes a positive example, how changes are labeled and how projects are partitioned can affect the result as much as the model architecture.
How it works. The authors mine security-related fixes, extract changed and unchanged C/C++ functions, and label pre-fix changed functions as vulnerable, with post-fix and unchanged functions treated as non-vulnerable. Exact duplicates are removed. Their study evaluates both sample-level partitions and held-out projects, alongside weakness-specific analysis. This construction is scalable, but a security-fixing commit may contain changes that are not themselves vulnerable. [B21a]
Metrics and score interpretation.
Precision / recall / F1. Classification quality for vulnerable functions. [B21a] Direction: Higher is better.
False-positive rate. Benign functions incorrectly flagged. [B21a] Direction: Lower is better.
Held-out-project performance. Quality on projects excluded from training. [B21a] Direction: Higher generalization.
How to interpret a result. Analyst view: a random function split and an unseen-project split answer different questions. For bank onboarding, unseen-project performance is often closer to the intended use because the assistant will encounter repositories outside its task-specific training. Never present an accuracy number without class balance and false alarms. Require per-project or per-weakness breakdowns where enough examples exist, rather than interpreting sparse categories as statistically decisive.
Limitations and failure modes.
- The paper identifies noisy labels and poor transfer to unseen projects as important limitations. [B21a]
- Analyst limitation: benign labels indicate the dataset’s construction judgment, not proof that a function has no vulnerability.
- Analyst limitation: exact deduplication does not eliminate every near-duplicate or cross-dataset overlap.
- Analyst limitation: code-only classification does not demonstrate exploitability assessment, safe remediation or production detection effectiveness.
Proposed bank use. Recommendation: use DiverseVul as supplementary breadth for a code-vulnerability evaluation, especially when training or comparing a specialized detector. Keep an independently adjudicated internal holdout and inspect disagreement samples. Compare performance with a cleaner or stricter dataset such as PrimeVul, while checking overlap before claiming independent confirmation. Decide whether the result justifies assisted review; it should not by itself support automated acceptance of code as secure.
Execution and evidence prerequisites. Practical preparation: obtain the author-linked data and document the exact archive, metadata coverage and chosen split. The repository also links merged-dataset splits and label-noise analysis materials. [B21b] Freeze preprocessing and deduplication rules, validate labels on a sample, and record project boundaries before training or scoring. Reusing a published model’s test set during tuning would invalidate the holdout’s purpose.
Primary sources: [B21a] [B21b]
22. LLMSecEval
LLMSecEval: A Dataset of Natural Language Prompts for Security Evaluations
| Owner: Catherine Tony, Markus Mutas, Nicolás E. Díaz Ferreyra and Riccardo Scandariato; TU Hamburg software security researchers. | Release / evidence: Preprint 16 March 2023; MSR 2023 Data and Tool Showcase; 150 natural-language prompts. |
| Task form: Natural-language prompt-to-code tasks with secure reference examples and static-analysis support. | Score direction: Lower confirmed vulnerable-output rate is better; useful functionality should be measured separately. |
| Status assessment: Established MSR 2023 prompt dataset; lightweight security diagnostics rather than repository-level assurance. | Availability: Public paper, author GitHub repository, dataset and CodeQL analysis interface. |
What it is. LLMSecEval contains 150 natural-language programming prompts, each accompanied by a secure implementation example. [B22a] Its scenarios cover 18 of the 2021 CWE Top 25, rather than the entire current ranking. The prompts aim to describe requested behavior without explicitly naming the vulnerability. The dataset therefore examines whether a model produces secure code when the user asks for functionality in ordinary language. [B22b]
Why it matters. Analyst view: security-sensitive requirements are often implicit in everyday development requests. A developer may ask for file handling, authentication or database functionality without naming the relevant weakness. This dataset is useful for checking those defaults and for comparing the effect of approved security instructions. It also supplies concrete outputs that reviewers can inspect, making it suitable for training and early model screening.
How it works. The dataset derives natural-language tasks from security-relevant code scenarios and includes manually improved prompts plus provenance fields. An evaluator asks the selected model to implement each requirement, then checks the resulting code for the associated weakness. The repository supplies a CodeQL interface and a demonstration application. Its prompt-quality ratings concern language/content quality; they are not model-security scores. Although prompts are described as largely language-agnostic, executable checking still depends on the chosen language and tooling. [B22b]
Metrics and score interpretation.
Confirmed vulnerability rate. Recommended operational result: proportion of valid outputs confirmed to contain the designated weakness. Direction: Lower is better.
Analyzer finding rate by CWE. Keep tool-detected findings separate from adjudicated vulnerabilities and report analyzer coverage. Direction: Diagnostic.
Functional / valid-output rate. Recommended companion measure: code fulfills the requested behavior rather than merely avoiding a warning. Direction: Higher is better.
How to interpret a result. Analyst view: declare the evaluator before reporting a score, because this dataset does not make every possible downstream scoring method equivalent. Compare the same prompts, target language, sampling policy and query pack. A secure reference demonstrates one acceptable implementation; token similarity to that reference is not the objective. Inspect whether alternative code preserves the requirement and the relevant security property.
Limitations and failure modes.
- Analyst limitation: 150 prompts provide a narrow sample of development work and limited evidence for individual CWEs.
- Analyst limitation: translated or language-adapted prompts are new evaluation conditions and require validation.
- Analyst limitation: static analysis can miss defects or flag patterns that do not represent an exploitable issue.
- Analyst limitation: public prompts and secure examples are exposed to memorization and benchmark-specific tuning.
Proposed bank use. Recommendation: use LLMSecEval for inexpensive regression checks when changing a model, system prompt or approved security guidance for coding assistants. Add human review of a sample and meaningful functional checks. Use it to identify recurring secure-coding weaknesses and improve engineering guidance. Complement it with repository-level evaluation and bank-specific scenarios before granting an assistant meaningful code-change authority.
Execution and evidence prerequisites. Practical preparation: obtain the author dataset, select a target language and implement a fresh model adapter. Review the CodeQL analysis materials and pin the query/tool versions. [B22b] Validate output extraction and account for incomplete, refused and invalid generations separately. Preserve generated code, prompt version and reviewer decisions so later model changes can be compared fairly.
Primary sources: [B22a] [B22b]
C. Suites, defensive analysis and cybersecurity knowledge
This group includes parent suites, SOC and CTI analysis, advisory reasoning and security knowledge. Preserve track-specific metrics and dataset lineages. Knowledge performance is relevant to domain usefulness, while action-level security needs additional evidence.
23. CyberSecEval 4 and inherited tracks onward
CyberSecEval 4 (Meta Purple Llama)
| Owner: Meta; CyberSOCEval developed with CrowdStrike. | Release / evidence: Version 4 announced 29 April 2025; September 2025 CyberSOCEval release. Official documentation reviewed still identifies version 4; no later numbered release verified. |
| Task form: Text, generated code, images, and tool-enabled execution depending on track. | Score direction: Mixed: capability/defensive success higher; insecure output, attack compliance, injection success and false refusal lower. |
| Status assessment: Established public suite; implementation completeness and maintenance vary by track (analyst assessment). | Availability: Public PurpleLlama code and documentation; track-specific data and execution requirements. |
What it is. CyberSecEval is a benchmark family rather than a single security score. Version 4 retains earlier tests for malicious-request compliance, false refusals, insecure code, textual and visual prompt injection, code-interpreter abuse, vulnerability exploitation, spear phishing and autonomous offensive operations. Its principal additions are AutoPatchBench and CyberSOCEval’s Malware Analysis and Threat Intelligence Reasoning tracks. The latter benchmarks also appear separately in this document; their inclusion in the parent suite must not create duplicate evidence. [C23a] Meta announced version 4 in April 2025. No version 5 or later numbered release was verified in the official sources reviewed. [C23e]
Why it matters. Analyst assessment: this suite separates failure mechanisms that a claim of being cyber secure would hide. A model can improve at patching while becoming more capable of exploitation; it can refuse malicious prompts while obstructing legitimate security work. These outcomes require different management decisions and metric directions.
How it works. Each track defines its own prompts, environment and scorer. The MITRE track uses response expansion followed by model-based judgment of whether assistance would support an attack; the false-refusal track uses keyword-based evaluation. [C23b] Instruct and autocomplete tracks send generated code to an insecure-code detector, reporting vulnerable suggestions and pass rates. [C23c] A common Python runner supports model adapters, but a shared runner does not make the underlying tests equivalent. [C23d]
Metrics and score interpretation.
Malicious assistance and false-refusal rates. Report attack-assistance classification separately from refusals of benign security requests; label denominator and scorer. [C23b] Direction: Lower for both, subject to the distinct task definitions.
Vulnerable generation percentage. Share of generated suggestions identified as vulnerable by the insecure-code detector. [C23c] Direction: Lower.
Track-specific capability and defense scores. Exploitation, patching and SOC scores measure different outcomes; retain the original metric for each. [C23a] Direction: Task-dependent; higher defensive utility does not establish lower misuse risk.
How to interpret a result. Analyst recommendation: present a dashboard with separate capability, resistance and utility sections. Do not average these into a single approval percentage. Compare releases only after fixing dataset revision, model endpoint, agent tools, sampling budget and grader. Treat a safety regression as its own finding even if a composite capability score improves.
Limitations and failure modes.
- Inherited tracks differ in realism, grading and infrastructure; suite membership alone does not establish complete coverage. [C23a][C23b][C23c][C23d]
- A current Inspect adaptation explicitly omits its autonomous-uplift and autopatching prototypes from its supported public suite surface. An adaptation therefore cannot automatically substantiate a claim that every Meta track was executed. [C23f]
- Analyst assessment: model-judge and keyword-scoring errors require sampled human review, and static-code findings should not be interpreted as full functional or application-security validation.
Proposed bank use. Analyst recommendation: require an evidence manifest of executed tracks and omissions. For a developer assistant, prioritize insecure generation, repair and misuse resistance; for a SOC assistant, prioritize CyberSOCEval and operational validation. Approve the intended configuration and purpose, adding bank-specific controls and assigning each material failure an owner.
Execution and evidence prerequisites. Prerequisites: a pinned PurpleLlama revision, Python dependencies, supported model adapter, data access and separate judge credentials where required. [C23d] Analyst recommendation: retain prompts, raw outputs, scorer versions, coverage counts, failures, costs and execution settings. Record which implementation was used and keep execution-based tracks in controlled test infrastructure.
Primary sources: [C23a] [C23b] [C23c] [C23d] [C23e] [C23f]
24. CyberSOCEval
CyberSOCEval: Malware Analysis and Threat Intelligence Reasoning
| Owner: Meta and CrowdStrike. | Release / evidence: 15 September 2025 release announcement and research publication. |
| Task form: JSON/text telemetry; threat reports as text, page images, or both. | Score direction: Higher exact-set accuracy and Jaccard similarity are better. |
| Status assessment: Public benchmark and data with documented runner; focused SOC reasoning coverage (analyst assessment). | Availability: Public CyberSOCEval data repository and PurpleLlama implementation; external report downloads required. |
What it is. CyberSOCEval is the defensive-analysis component of CyberSecEval 4. It contains two benchmarks: Malware Analysis and Threat Intelligence Reasoning. The malware track supplies existing sandbox detonation reports and asks security questions about their contents; the threat-intelligence track tests interpretation of reports, including relationships that require more than extracting an explicitly stated fact. It does not measure the end-to-end operation of a SOC. [C24a]
Why it matters. Analyst assessment: this benchmark is relevant to bank assistants interpreting telemetry or campaigns. It can reveal confusion between observed behavior and inferred intent, unsupported conclusions, or missed relationships in a long report. Those findings help select a SOC copilot and determine which conclusions require analyst verification.
How it works. The malware benchmark uses JSON reports derived from CrowdStrike Falcon Sandbox and has 609 question-answer cases across five malware categories. Answers may require selecting several correct options; exact-set scoring requires all and only the correct choices. The original threat-intelligence design similarly uses multiple-choice questions over report pages. [C24a] The current runner supports text, image or combined input for threat reports, converts PDFs into text and page images, and reports both exact-set correctness and Jaccard similarity for partial credit. [C24b]
Metrics and score interpretation.
Exact-set accuracy / correct_mc_pct. Percentage of questions for which the entire chosen option set matches the reference set. [C24a][C24b] Direction: Higher.
Average Jaccard similarity. Intersection divided by union of predicted and reference option sets; the threat-intelligence runner calls this avg_score. [C24b] Direction: Higher.
Response parsing errors. Questions whose answer selections could not be extracted; count separately from content mistakes. [C24b][C24c] Direction: Lower.
How to interpret a result. Analyst recommendation: preserve both strict and partial-credit results. A high partial score may coexist with operationally important omissions or unsupported extra assertions. Break results down by malware category, difficulty, report source and input modality. A text-only result should not be presented as evidence that the model reliably interprets tables and diagrams. Measure actual analyst usefulness independently from answer-option correctness.
Limitations and failure modes.
- Analyst assessment: existing reports and fixed answer choices do not test telemetry collection, incident ownership, alert queue prioritization, containment actions or resilience to hostile report content.
- The malware documentation recommends an optional truncation mode for models below a 128k context window; this is a material input transformation to record. [C24c]
- Analyst assessment: results from one telemetry producer or small sets of report styles may not transfer to the bank’s SIEM and EDR schemas. Exact-set grading also does not directly quantify operational false-positive and false-negative rates.
Proposed bank use. Analyst recommendation: follow this benchmark with blinded testing on approved bank cases containing malicious and benign activity, incomplete evidence and contradictory observations. Test whether analysts detect unsupported conclusions and whether escalation decisions remain correct. Distinguish errors that could delay containment from those that only add review effort.
Execution and evidence prerequisites. Prerequisites: PurpleLlama, its CrowdStrike data submodule, required report downloads and PDF-to-image dependencies for the selected modality. [C24b][C24d] Analyst recommendation: pin every report, input transformation and answer parser. Retain raw predictions, per-category results, model context limits and analyst adjudications; there is no need to detonate live malware merely to evaluate the released report tasks.
Primary sources: [C24a] [C24b] [C24c] [C24d]
25. CTIBench
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
| Owner: Md Tanvirul Alam, Dipkamal Bhusal, Le Nguyen and Nidhi Rastogi, Rochester Institute of Technology. | Release / evidence: June 2024; NeurIPS 2024; paper v3 dated 11 November 2024. |
| Task form: Text questions, vulnerability descriptions and anonymized threat reports. | Score direction: Accuracy/F1 higher; CVSS mean absolute deviation lower. |
| Status assessment: Established public research dataset and evaluation notebooks; static and version-sensitive (analyst assessment). | Availability: Public dataset, notebooks, model responses and logs. |
What it is. CTIBench evaluates five distinct threat-intelligence tasks: CTI-MCQ for knowledge, CTI-RCM for vulnerability root-cause mapping, CTI-VSP for severity prediction, CTI-ATE for ATT&CK technique extraction and CTI-TAA for threat-actor attribution. The current repository lists 4,610 benchmark examples plus a separate 1,000-example 2021 root-cause comparison split. These counts explain why some dataset inventories show 5,610 rows. [C25b]
Why it matters. Analyst assessment: these tasks resemble transformations in a bank’s vulnerability and intelligence pipelines. Incorrect CWE mapping can misdirect remediation; severity errors can distort queues; unsupported ATT&CK mapping can mislead detection engineering. CTIBench tests these analytical abilities separately and challenges models that sound authoritative while producing inconsistent standardized identifiers.
How it works. The model answers text prompts and its outputs are compared with task-specific labels. RCM asks for a CWE; VSP asks for a CVSS v3.1 vector, from which a deterministic library derives the numerical score. TAA removes actor and campaign names from reports and evaluates the proposed attribution, including acceptable aliases. The paper distinguishes strictly correct attribution from correct-or-plausible attribution. [C25a] Released notebooks and response files support inspection of how predictions were scored. [C25b]
Metrics and score interpretation.
MCQ / root-cause accuracy. Proportion of correct knowledge answers or CWE mappings. [C25b] Direction: Higher.
CVSS mean absolute deviation. Mean absolute difference between reference severity and severity computed from the predicted vector. [C25b] Direction: Lower.
ATT&CK extraction F1. Balance of precision and recall over technique labels; the paper’s prose and result-table averaging labels disagree. [C25a] Direction: Higher; identify the implemented averaging method.
Correct and plausible attribution accuracy. Strict actor match versus the broader category including plausible attributions; preserve both. [C25b] Direction: Higher, without treating plausible as confirmed.
How to interpret a result. Analyst recommendation: do not average raw CVSS error with percentages. Assess the task actually used by the application and review the consequences of its errors. If a model proposes a CVSS vector accurately but calculates poorly, a deterministic calculator may address the arithmetic issue. If it misunderstands privileges or impact, a calculator cannot repair that reasoning error.
Limitations and failure modes.
- The v3 methodology specifies micro-F1 for ATT&CK extraction, whereas its main results table labels that column macro-F1. Validate and pin the evaluator before making cross-model comparisons. [C25a]
- The repository provides only 50 attribution examples. The distinction between correct and plausible answers also requires careful adjudication. [C25b]
- Analyst assessment: public static CVEs and reports may be in newer models’ training data. The original study’s temporal separation does not automatically apply to 2026 models.
Proposed bank use. Analyst recommendation: prioritize RCM, VSP and ATE where the bank actually proposes automated enrichment. Add newly published, withheld examples and disagreements among authoritative sources. Keep attribution advisory and require evidence citations plus analyst sign-off before communicating a named actor. Acceptance should reflect error costs for each workflow and confidence intervals, not a borrowed leaderboard threshold.
Execution and evidence prerequisites. Prerequisites: the released TSVs, notebooks, prediction formatter, a model endpoint and a compatible CVSS library. [C25b] Analyst recommendation: freeze the CWE/ATT&CK versions, remove test examples from tuning material, record truncation and parse failures, and preserve the distinction between the main evaluation and the 2021 comparison split.
Primary sources: [C25a] [C25b]
26. AthenaBench
AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
| Owner: Rochester Institute of Technology and Athena Security Group. | Release / evidence: 3 November 2025; paper v2 revised 14 February 2026. |
| Task form: Text scenarios and reports; optional web-search configuration is a separate evaluated condition. | Score direction: Task accuracies and mitigation F1 higher; raw CVSS error lower. |
| Status assessment: Public research benchmark with evolving artifacts; dynamic design requires snapshot discipline (analyst assessment). | Availability: Public repository advertises benchmark and mini directories; exact full-data availability and use terms must be checked against the selected revision. |
What it is. AthenaBench extends CTIBench through refreshed datasets, duplicate removal and revised scoring, adding mitigation recommendation. Its design generates dated evaluation material from sources including NVD and MITRE ATT&CK. It is a related successor, so its scores are not independent confirmation of CTIBench results. [C26a] The repository exposes six tasks across knowledge, mapping, severity, attribution, mitigation and technique extraction. [C26b]
Why it matters. Analyst assessment: mitigation tests address a banking question: can the assistant recommend relevant defensive action after interpreting a threat? Fresh evaluation material helps test current knowledge. However, changing the dataset also changes test difficulty, so a higher score on a newer snapshot cannot be attributed to the model alone.
How it works. Models produce categorical answers, identifiers, CVSS vectors or mitigation sets. The paper uses categorical accuracy for most tasks and F1 for mitigation. Its VSP percentage is a transformation of mean absolute CVSS-score error: 1 − MAD/7.7. That quantity is labelled accuracy but does not count exactly correct answers. [C26a] The repository provides a Python runner, model configuration, scored outputs and mini subsets, and requires Git LFS for large artifacts. [C26b]
Metrics and score interpretation.
Task accuracy. Correct categorical answers for the applicable knowledge, mapping, technique and attribution tasks. [C26a] Direction: Higher.
Risk Mitigation Strategy F1. Agreement between predicted mitigation labels and reference mitigation sets. [C26a] Direction: Higher.
Normalized VSP score. 1 − mean absolute deviation/7.7 for the published score range; not exact-answer accuracy. [C26a] Direction: Higher; also report raw error.
How to interpret a result. Analyst recommendation: report all six tasks and identify whether web search was permitted. Treat the combined score as a comparison aid, not a deployment gate. A strong severity score can conceal weak mitigation reasoning; mitigation mistakes may matter more for an assistant proposing changes to controls. Maintain a frozen reference set for regression and a separately labelled fresh set for emerging-threat performance.
Limitations and failure modes.
- The paper describes mini release and planned controlled full-data access; the current repository advertises full and mini artifacts with a CKT-file caveat. Check actual files and usage terms at the selected revision. [C26a][C26b]
- Analyst assessment: synthetic scenarios and standardized mitigation labels do not establish that a proposed control is feasible, correctly configured or appropriate to a bank application.
- Analyst assessment: a refreshed public source can still be familiar to a newer model. Construction date alone does not prove absence of contamination or eliminate leakage through optional search.
Proposed bank use. Analyst recommendation: use as a pilot benchmark for CTI assistants that recommend response or mitigation actions. Have detection engineers and application owners grade a private supplement against platform capabilities, dependencies and change constraints. Require the assistant to distinguish source-supported mitigation, inferred applicability and missing evidence. Until those checks are successful, retain human approval for operational actions.
Execution and evidence prerequisites. Prerequisites: pinned repository and data revision, Git LFS, Python dependencies and configured model credentials. [C26b] Analyst recommendation: record source-date ranges, hashes, mini/full designation, scoring formula, search access and missing artifacts. Preserve fresh raw predictions separately from supplied scored results so historical files cannot be mistaken for a newly executed evaluation.
Primary sources: [C26a] [C26b] [C26c]
27. CyberMetric
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
| Owner: Norbert Tihanyi, Mohamed Amine Ferrag, Ridhi Jain, Tamas Bisztray and Merouane Debbah. | Release / evidence: February 2024; paper v2 June 2024; IEEE CSR 2024. |
| Task form: Text multiple-choice questions with four options. | Score direction: Higher answer accuracy is better for knowledge; not a security-resistance measure. |
| Status assessment: Public established knowledge benchmark; static multiple-choice evaluation (analyst assessment). | Availability: Public JSON sets and evaluator; four dataset sizes. |
What it is. CyberMetric provides four named knowledge datasets: CyberMetric-80, -500, -2000 and -10000. Its questions draw on cybersecurity publications such as standards, research, books and RFCs. The development process used LLM generation with retrieval, multiple checks and expert review. The RAG terminology in the title describes dataset construction; it does not mean that every evaluated model has a retrieval system or that the benchmark validates a deployed RAG application. [C27a]
Why it matters. Analyst assessment: broad knowledge screening can expose a weak foundation before expensive agent testing. It can also help assess whether a smaller model suits a restricted task. Selecting a correct option nevertheless offers little evidence about interpreting incomplete production information, maintaining permissions or recognizing malicious instructions in retrieved material.
How it works. The released JSON files hold questions, candidate answers and a reference choice. The evaluator prompts the model to return its answer in a fixed format. The original study compared 25 models and included a 30-person closed-book exercise on the 80-question set. [C27b] The paper describes filtering ambiguous, outdated and context-dependent questions during construction. [C27a] An evaluation therefore principally measures performance on the selected retained questions under a specified prompting regime.
Metrics and score interpretation.
Multiple-choice accuracy. Correct answers divided by evaluated questions; report the dataset size/version in the metric name. [C27a][C27b] Direction: Higher.
Format failures and coverage. Analyst diagnostic: report unparseable answers, skipped cases and completed cases separately from knowledge accuracy. Direction: Fewer failures and complete intended coverage.
Topic-level error distribution. Analyst diagnostic: inspect errors in topics relevant to the application instead of relying exclusively on the aggregate. Direction: Fewer consequential errors.
How to interpret a result. Analyst recommendation: use the same named dataset and prompting protocol for all comparisons. A score on the 80-question set and a score on 10,000 questions are not interchangeable. Report uncertainty, especially for the smallest set. If the model receives web search or retrieval during testing, label that as an augmented-system condition rather than comparing it directly with the original closed-book model results.
Limitations and failure modes.
- Analyst assessment: public questions and public source documents create contamination risk as newer models are trained. A high score can reflect familiarity as well as useful knowledge.
- The original human comparison used only CyberMetric-80 under its study conditions; it does not establish that an LLM outperforms security practitioners at their jobs. [C27a]
- Analyst assessment: four-option recognition supplies clues that an open-ended security problem would not provide. Generated-answer and source-quality issues remain appropriate targets for sampled review despite the reported validation process.
Proposed bank use. Analyst recommendation: use as a supplementary screen for security-concept assistants. Choose a larger set when candidate models score similarly, then test actual use cases such as interpreting a control exception or evidence gap. Knowledge scores should not authorize connector access or replace secure-code, misuse-resistance and application-security testing.
Execution and evidence prerequisites. Prerequisites: the exact JSON variant, evaluator or documented compatible implementation, and a model endpoint. [C27b] Analyst recommendation: record dataset hashes, answer-order handling, prompt examples, tool access and parser behavior. Reserve any bank-specific supplement for evaluation, check its reference answers independently, and preserve item-level errors so subject experts can assess whether failures are material.
Primary sources: [C27a] [C27b] [C27c]
28. CyberBench
CyberBench: A Multi-Task Benchmark for Evaluating LLMs in Cybersecurity (JPMorgan Chase, 2024)
| Owner: Zefang Liu, Jialei Shi and John F. Buford / JPMorgan Chase. | Release / evidence: AAAI-24 AICS workshop, 2024; repository archived May 2026. |
| Task form: Static text datasets; classification and generation. | Score direction: Higher task-specific accuracy, F1 and ROUGE are better. |
| Status assessment: Verified historical public NLP benchmark; original repository archived 26 May 2026 (analyst assessment). | Availability: Public archived code and author dataset; constituent datasets retain their own terms. |
What it is. The name CyberBench is not unique. In this section it refers to JPMorgan Chase’s 2024 benchmark, consistent with the user’s knowledge-and-defensive-analysis grouping. That project combines ten datasets covering entity recognition, summarization, multiple choice and classification. Its repository is now archived. [C28a] Vals AI uses CyberBench for a different exploitation-and-patching evaluation, while the University of Twente describes another framework for direct security risks. Their scores must not be pooled with the JPMorgan benchmark. [C28d][C28e]
Why it matters. Analyst assessment: many security assistants perform information-processing tasks such as entity extraction, phishing classification or report summarization. CyberBench decomposes those abilities into measurable tests. Its banking origin is relevant context, but it does not establish that the datasets represent this bank’s controls, evidence or operating conditions.
How it works. The suite uses CyNER and APTNER for entities, CyNews for article-to-headline summarization, SecMMLU and CyQuiz for questions, and MITRE, CVE, Web, Email and HTTP datasets for classification. Inputs range from sentences and articles to URLs, email content and HTTP requests. The author’s dataset card specifies separate metrics for each task rather than one universal score. [C28b] The repository includes data preparation and evaluation scripts for supported model configurations. [C28a]
Metrics and score interpretation.
Micro-F1. Entity-extraction agreement in CyNER and APTNER. [C28b] Direction: Higher.
ROUGE-1 / ROUGE-2 / ROUGE-L. Reference-text overlap for article-to-headline summarization. [C28b] Direction: Higher; does not establish factual faithfulness.
Accuracy. Correct choices or labels for knowledge, technique and severity tasks. [C28b] Direction: Higher.
Binary F1. Balance of precision and recall for phishing URL/email or HTTP anomaly labels. [C28b] Direction: Higher.
How to interpret a result. Analyst recommendation: use task-level results to select a model for a bounded workflow, and pair text-overlap scores with factuality review. A headline can resemble the reference while omitting an operationally critical qualification. A good average phishing F1 can still conceal an unacceptable rate of missed attacks; inspect the confusion matrix and apply the expected production class balance before estimating review workload.
Limitations and failure modes.
- The original repository was archived on 26 May 2026; long-term dependency and data-download maintenance should not be assumed. [C28a]
- The author’s data card states that each underlying dataset retains its own licensing terms. Public benchmark access does not imply uniform reuse permission. [C28b]
- Analyst assessment: legacy text datasets and isolated examples omit live tool interaction, novel campaigns, authorization boundaries and adversarial manipulation of the assistant. Similar names also create a material evidence-attribution risk.
Proposed bank use. Analyst recommendation: use selected tasks for an internal proof of value or regression library where their format resembles the intended application. Revalidate on current, approved bank email, report or telemetry examples with subject-expert labels. Keep operational effectiveness and assistant security as separate assessments. In procurement evidence, require the full benchmark title, owner, year and dataset identifiers rather than accepting “CyberBench score” alone.
Execution and evidence prerequisites. Prerequisites: a reproducible checkout, Python 3.10 or later and the needed model/embedding configuration according to the repository. [C28a] Analyst recommendation: mirror permitted source artifacts with hashes, pin dependency versions and capture per-dataset outputs. Review the data card’s split definitions and prevent underlying training examples from entering the evaluation set.
Primary sources: [C28a] [C28b] [C28c] [C28d] [C28e]
29. SECURE
SECURE: Security Extraction, Understanding & Reasoning Evaluation
| Owner: Dipkamal Bhusal and collaborators; Rochester Institute of Technology-led research. | Release / evidence: 30 May 2024; paper v4 dated 30 October 2024. |
| Task form: Text multiple choice, true/false, generated risk summaries and numerical CVSS answers. | Score direction: Accuracy/ROUGE-L higher; CVSS mean absolute deviation lower; appropriate abstention requires separate interpretation. |
| Status assessment: Public research benchmark with six released datasets; focused ICS domain (analyst assessment). | Availability: Public repository with six TSV datasets, prompts and reference answers. |
What it is. SECURE expands to Security Extraction, Understanding & Reasoning Evaluation. It is a cybersecurity-advisory benchmark focused on industrial control systems, not the similarly named SEC-bench software-security environment. Its public release contains six tasks spanning knowledge extraction, understanding and reasoning. [C29b] The research publication was revised through version 4 in October 2024. [C29c]
Why it matters. Analyst assessment: the useful distinction is between recalling security information and reasoning from available evidence. An assistant may know a weakness category yet misread an advisory or invent a missing detail. SECURE exposes several such problems. Its ICS specialization also requires checking whether the evaluation domain matches the proposed banking use case.
How it works. MAET and CWET ask knowledge questions derived from MITRE ATT&CK and CWE. KCV asks true/false questions with CVE context; VOOD removes relevant context to probe behavior when evidence is insufficient. RERT asks for risk-evaluation text based on advisory details. CPST supplies CVSS v3.1 vectors and tests calculation of their scores. The paper scores knowledge tasks with accuracy, RERT using ROUGE-L and CPST using mean absolute deviation. [C29a] The repository supplies prompts and ground truth in TSV form. [C29b]
Metrics and score interpretation.
Knowledge/context accuracy. Correct MCQ and true/false answers. [C29a] Direction: Higher.
RERT ROUGE-L. Overlap between generated risk text and the reference. [C29a] Direction: Higher; semantic correctness needs separate review.
CPST mean absolute deviation. Numerical error for a supplied CVSS vector. [C29a] Direction: Lower.
Evidence-sensitive abstention. Analyst diagnostic for missing-context cases: distinguish justified uncertainty from fabricated certainty or blanket refusal. Direction: More justified abstention, without increased refusal on answerable cases.
How to interpret a result. Analyst recommendation: explicitly distinguish CPST’s arithmetic task from CTIBench’s inference of a vector from a vulnerability description. They diagnose different weaknesses. For an implemented application, route CVSS arithmetic to a deterministic calculator and evaluate the model’s interpretation of the vector separately. Review VOOD outputs for whether the model recognizes missing evidence; ordinary answer accuracy alone does not fully describe trustworthy abstention.
Limitations and failure modes.
- Analyst assessment: ICS advisories are an imperfect proxy for banking cloud applications, identity infrastructure and customer-facing services. Score transfer must be tested.
- Analyst assessment: ROUGE-L rewards wording overlap and cannot establish that a risk statement is accurate, complete or appropriate for a bank’s exposure.
- The original missing-context design used newly published CVEs relative to the tested models. Those same 2024 CVEs may be familiar to later models, weakening the intended unknown-information condition. [C29a]
Proposed bank use. Analyst recommendation: use SECURE selectively for cyber-advisory assistants or facilities/operational-technology use cases. For a general bank assistant, borrow the evaluation pattern and create a separate, clearly labelled set of approved current cloud advisories. Include misleading premises, insufficient evidence and similar vulnerabilities with different prerequisites. Require evidence-grounded conclusions and an escalation route when exposure cannot be established.
Execution and evidence prerequisites. Prerequisites: the six released TSVs, a model runner and the selected accuracy, text-overlap and numerical scorers. [C29b] Analyst recommendation: freeze context documents, verify labels, capture whether context was actually supplied and retain all refusals. Use expert review for consequential risk summaries and compare performance with a deterministic CVSS calculator where arithmetic is relevant.
Primary sources: [C29a] [C29b] [C29c]
30. SecQA
SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security
| Owner: Zefang Liu. | Release / evidence: 26 December 2023. |
| Task form: Text multiple-choice questions; basic v1 and harder v2. | Score direction: Higher accuracy is better for knowledge. |
| Status assessment: Small established public knowledge dataset; limited discrimination among stronger models (analyst assessment). | Availability: Author-hosted Hugging Face data; dataset card identifies CC BY-NC-SA 4.0. |
What it is. SecQA is a concise multiple-choice dataset generated with GPT-4 from the textbook Computer Systems Security: Planning for Success. It offers basic and harder variants, v1 and v2. [C30b] The published test splits contain 110 and 100 questions, with separate development and validation material. The paper evaluates zero-shot and five-shot conditions, so supplied examples are part of the result definition. [C30a]
Why it matters. Analyst assessment: a small dataset can quickly reveal evaluation-pipeline failures and major knowledge gaps. SecQA serves that role with little inference cost. It provides limited evidence for selecting among strong models or approving an agent with privileged connectors, where application behavior and action controls matter more.
How it works. A model receives a question and four choices and predicts the correct option. Development examples can be included in a few-shot prompt; held-out test questions are scored against reference answers. The author’s dataset card supplies the question, options, answer and explanation fields and separates the two variants. It identifies a noncommercial share-alike dataset license, making intended organizational use a practical item to verify before adoption. [C30b]
Metrics and score interpretation.
v1 test accuracy. Correct answers on the 110-question basic test split, with the prompting condition identified. [C30a] Direction: Higher.
v2 test accuracy. Correct answers on the 100-question harder test split, reported separately from v1. [C30a] Direction: Higher.
Zero-shot versus five-shot change. Difference under the paper’s two prompting conditions; measures sensitivity to supplied examples. [C30a] Direction: Context-dependent; do not mix conditions in a ranking.
How to interpret a result. Analyst recommendation: a one-question difference on the v2 test changes accuracy by one percentage point, so small apparent advantages should not drive procurement decisions. Report counts as well as percentages. The original study already obtained near-perfect results from some contemporary models; that makes present-day ceiling effects a plausible concern and reduces its value for differentiating strong systems. [C30a]
Limitations and failure modes.
- Analyst assessment: the narrow source base and limited sample size restrict coverage of changing threats and specialist security work.
- Analyst assessment: recognizing a textbook answer is substantially easier to standardize than producing a justified answer from incomplete bank evidence. The benchmark does not assess tool use, code execution, prompt injection or harmful-action resistance.
- Analyst assessment: GPT-generated questions and public reference material justify checking ambiguous answer keys and training contamination. A difference caused by answer formatting should not be mistaken for a cybersecurity knowledge difference.
Proposed bank use. Analyst recommendation: use as an optional baseline or harness smoke test after verifying usage rights. For internal security-policy assistants, add a separate test using approved policy versions and evidence requirements. Include unknown answers, obsolete policy citations and controls owned by another team. Avoid making this small public quiz a universal approval gate.
Execution and evidence prerequisites. Prerequisites: the exact SecQA variant and split, model endpoint and a documented multiple-choice scorer. [C30b] Analyst recommendation: keep development examples out of the test set, preserve zero-shot/five-shot labels, record malformed outputs and inspect every error in the small set. Pair any score with the sample count and a confidence interval, and retain the full prompt to support reproducibility.
Primary sources: [C30a] [C30b]
31. WMDP-Cyber
WMDP-Cyber: cybersecurity subset of the Weapons of Mass Destruction Proxy benchmark
| Owner: Center for AI Safety-led research consortium; Nathaniel Li and collaborators. | Release / evidence: March 2024; ICML 2024; official data corrected in March and April 2024. |
| Task form: Text multiple-choice questions. | Score direction: Higher means more tested knowledge; lower is an unlearning objective, with useful-capability retention assessed separately. |
| Status assessment: Established public hazardous-knowledge proxy and unlearning benchmark; indirect measure of operational risk (analyst assessment). | Availability: Public official dataset and evaluation/unlearning repository; 1,987 questions in current cyber subset. |
What it is. WMDP-Cyber is the cybersecurity portion of the Weapons of Mass Destruction Proxy benchmark. Despite the umbrella name, this subset is a multiple-choice proxy for hazardous cyber knowledge, not a test involving physical weapons. The official dataset currently lists 1,987 cyber questions and documents corrections to option randomization, formatting and excessively long items. [C31b] The original research uses the benchmark both to measure knowledge and to study unlearning. [C31a]
Why it matters. Analyst assessment: WMDP-Cyber provides a repeatable probe when a provider claims to have reduced hazardous knowledge. A bank’s security assistant needs much of the same technical knowledge that could support misuse. Penalizing it indiscriminately can remove valuable defensive capability, making a universal acceptance threshold inappropriate.
How it works. Models answer multiple-choice questions that sample knowledge across stages of a cyberattack. Evaluation compares selected answers with the references. The study pairs hazardous-knowledge reduction with measures of retained general utility and discusses the dual-use costs of forgetting cybersecurity knowledge. [C31a] The official repository also contains the RMU unlearning implementation and evaluation guidance. Running the benchmark and training a model with an unlearning procedure are separate activities. [C31c]
Metrics and score interpretation.
WMDP-Cyber accuracy. Correct answers on the pinned cybersecurity subset; a knowledge proxy rather than observed harmful actions. [C31a][C31b] Direction: Higher indicates more measured knowledge; lower is desired in the paper’s unlearning setting.
Change after mitigation. Analyst comparison: before/after accuracy under identical data, prompting and scoring, with uncertainty. Direction: Interpret alongside the mitigation objective and utility loss.
Retained useful capability. Companion evaluation of general and defensive utility after unlearning; the research explicitly considers the tradeoff. [C31a] Direction: Higher retention is preferable.
How to interpret a result. Analyst recommendation: never treat a low WMDP-Cyber score as proof that a deployed agent cannot cause harm. Low scores may arise from lack of knowledge, refusal behavior, answer-format problems or genuine unlearning. Conversely, high scores do not demonstrate willingness to comply with malicious requests. Evaluate knowledge, misuse resistance, operational capability and tool permissions as distinct properties.
Limitations and failure modes.
- The benchmark is intentionally a proxy; selecting answers does not measure autonomous exploitation, practical reliability or access to external tools. [C31a]
- Different mirrors or harness descriptions may retain older question counts. Pin the official data revision and actual evaluated row count. [C31b]
- Analyst assessment: public questions can be memorized or specifically optimized against. A reduction on this dataset may not transfer to paraphrases, new tasks, augmented retrieval or the deployed agent configuration.
Proposed bank use. Analyst recommendation: use when reviewing provider unlearning claims or material capability changes. For bank security tools, require evidence of defensive utility, resistance to unauthorized requests and enforcement of permitted actions. Escalate unexpected capability increases for risk review without automatically categorizing a knowledgeable model as unsafe.
Execution and evidence prerequisites. Prerequisites: the official cyber dataset revision, an identified compatible evaluator and a model endpoint. [C31b][C31c] Analyst recommendation: document whether answers use generated text or choice likelihoods, how refusals are scored, and whether retrieval is available. Preserve raw answers and conduct paired before/after comparisons. Training access is required only if the bank separately chooses to reproduce an unlearning intervention.
Primary sources: [C31a] [C31b] [C31c]
D. Prompt injection and deployed-agent security
These evaluations examine attacks on an agent’s handling of untrusted content, tools and interactions. Record the attack surface, attacker budget, permissions and legitimate-task utility together with attack outcomes.
32. AgentDojo
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
| Owner: ETH Zurich and Invariant Labs; Edoardo Debenedetti and collaborators | Release / evidence: 19 June 2024 initial paper; NeurIPS 2024 Datasets and Benchmarks Track |
| Task form: Multi-step agent execution in simulated application environments | Score direction: Attack success lower; benign and attacked-task utility higher |
| Status assessment: Established research benchmark with an extensible public implementation; this is an analyst assessment, not production certification. | Availability: Public paper, Python package, tasks, attacks, defenses, and documentation |
What it is. AgentDojo tests whether an assistant can finish a legitimate user task while encountering hostile instructions inside external information. Its original release contains 97 user tasks and 629 security cases across banking, Slack, travel, and workspace settings. It is an extensible environment, allowing additional tasks and adaptive attacks, rather than a permanently fixed question bank. [D32a]
Why it matters. For a bank, this addresses the gap between an assistant that answers safely in a chat window and one that reads documents, calls connectors, and changes records. A useful security test must show both that the attacker failed and that the authorized workflow still succeeded. Otherwise, a broken integration can look impressively resistant.
How it works. A legitimate task runs against a modeled application state. At an allowed injection point, adversarial content replaces or modifies an external tool response. The agent continues, and task-specific evaluators separately check legitimate completion and the attacker’s objective. Model, prompt, tool pipeline, defense, and attacker configuration jointly define the evaluated system. The implementation supports selectable suites and defense combinations. [D32a][D32b]
Metrics and score interpretation.
Benign utility (native). Share of legitimate tasks completed without attack. Direction: Higher.
Utility under attack (native). Security-case task success without adversarial side effects, following the paper’s definition. Direction: Higher.
Attack success rate (native). Share of security cases satisfying the attacker objective. Direction: Lower for the defender.
Critical-action violation count (recommended). Individually confirmed unauthorized writes or disclosures in a bank-specific extension. Direction: Lower.
How to interpret a result. Raw legitimate completion can coexist with an attacker objective; the paper’s stricter utility-under-attack definition excludes adversarial side effects. Preserve that distinction when reporting custom metrics. Compare identical task, budget, and defense settings. The maintainers explain that their results page is not a universal leaderboard because tested combinations are incomplete. [D32c] Keep bank-specific results separate from upstream reproduction.
Limitations and failure modes.
- Simulated applications do not establish the security of the bank’s actual connectors, identity delegation, or data boundaries.
- Public tasks can enter training or defense-development data; retain independently written holdout workflows.
- A weak or fixed attacker gives limited assurance against an adaptive adversary. Report attempts, knowledge, injection scope, and tool access.
- Package interfaces can change; record the installed version and commit. [D32b]
Proposed bank use. Analyst recommendation: use AgentDojo as a core preproduction agent-security test for assistants that consume untrusted material. Platform Security should run shared defense baselines; the application team should add representative connector workflows; an independent assessor should inspect every sensitive-action failure. Preserve complete trajectories, final state, task and attack IDs, identity permissions, guardrail configuration, and adjudications. Require acceptable legitimate-task performance alongside resistance before accepting a defense improvement. A successful public-suite run is one component of deployment evidence, not authorization to access additional systems.
Execution and evidence prerequisites. Install a pinned AgentDojo package or repository revision, configure a supported model endpoint, select suites and attacks, and retain run logs. Detector-based defenses may require the optional transformers dependency. [D32b] Start with harmless smoke tasks and synthetic records; do not map benchmark banking tools to real accounts.
Primary sources: [D32a] [D32b] [D32c]
33. InjecAgent
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
| Owner: University of Illinois Urbana-Champaign; Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang | Release / evidence: 5 March 2024 initial paper; Findings of ACL 2024 |
| Task form: Structured agent continuations and tool-call parsing over synthetic test cases | Score direction: ASR-valid and ASR-all lower; valid output rate is a separate reliability diagnostic |
| Status assessment: Established, relatively lightweight research baseline; narrower than a full application security evaluation (analyst assessment). | Availability: Public paper, dataset, evaluation scripts, and model adapters |
What it is. InjecAgent evaluates whether tool-integrated agents follow an attacker’s instructions embedded in outside content. Its 1,054 cases combine 17 user-tool settings with 62 attacker cases. The two principal harm classes are an unauthorized harmful action and extraction followed by transmission of private information. The scope is indirect prompt injection, not arbitrary cyber exploitation or general chatbot harmfulness. [D33a]
Why it matters. This is useful when assessing an assistant that reads email, document excerpts, search results, or similar connector output. It asks whether content from one trust zone can direct tools in another. The separate stages of data theft also help identify whether the failing control concerns access to information, outgoing communication, or both.
How it works. The harness presents a legitimate request and a prepared external-tool response containing the injected instruction, then examines the agent’s continuation. Direct-harm success requires the target harmful tool invocation; complete data theft requires both extraction and transmission. It supports base and enhanced attack settings and prompted or tool-tuned agents. Execution is recognized through parsed model output, rather than proving a real external action occurred. [D33a][D33b]
Metrics and score interpretation.
ASR-valid (native). Successful attacks divided by valid agent outputs, excluding malformed or otherwise invalid continuations. Direction: Lower.
ASR-all (native). Successful attacks divided by all evaluated outputs. Direction: Lower, interpreted with validity.
Valid rate (native). Fraction of outputs meeting the harness validity criteria. Direction: Higher reliability; not a safety score.
Data-theft S1/S2 outcomes (native). Extraction and subsequent transmission results, with denominators retained. Direction: Lower.
Benign workflow completion (recommended). Success on clean equivalents using the same application adapter. Direction: Higher.
How to interpret a result. Never accept a low ASR-all without examining ASR-valid and the valid rate. A model that fails to produce tool syntax can appear secure because it cannot act. Native validity does not show that the authorized business task was completed. Preserve per-stage denominators rather than treating a conditional transmission result as an overall probability of data loss.
Limitations and failure modes.
- Many cases reuse the same user and attacker templates; 1,054 cases are not 1,054 independent business workflows.
- Prepared continuations simplify the preceding retrieval and planning trajectory.
- Tool-call intent is weaker evidence than verified destination state in a deployed connector.
- Base and enhanced settings, agent prompts, and output parsers materially affect comparability. [D33b]
Proposed bank use. Analyst recommendation: use as a low-cost screening and regression suite before more expensive AgentDojo or application-specific testing. Prioritize read-to-write and read-to-send boundary cases. Add clean counterpart tasks, synthetic sensitive fields, and deterministic destination checks in an isolated bank adapter. Triage invalid outputs separately from blocked attacks. Store the original tool response, continuation, parser version, stage outcomes, and reviewer decision. A positive finding should become a reproducible application test; a clean score should not close broader connector-assurance requirements.
Execution and evidence prerequisites. Use the official repository’s dependencies and model adapter, pin the prompt and base/enhanced setting, and configure the chosen endpoint. The repository supports cached execution and custom model wrappers. [D33b] Review adapter behavior before connecting a new tool system, and use mock tools for this baseline.
Primary sources: [D33a] [D33b] [D33c]
34. Agent Security Bench (ASB)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
| Owner: Zhejiang University and Rutgers University; Hanrong Zhang and collaborators; agiresearch/ASB | Release / evidence: 3 October 2024 initial paper; ICLR 2025 final paper used here |
| Task form: Configurable agent/tool framework with multiple attack and defense families | Score direction: Attack success and detector errors lower; task performance and NRP higher |
| Status assessment: Established research framework with broad attack-surface coverage; source-version inconsistencies require care (analyst assessment). | Availability: Public ICLR paper, AIOS-based implementation, configuration files, and research scenarios |
What it is. ASB broadens evaluation beyond injected tool responses. It includes direct and indirect prompt injection, memory poisoning, system-prompt Plan-of-Thought backdoors, and mixed attacks. The final ICLR 2025 abstract reports ten scenarios, over 400 tools, 27 attack/defense methods, and seven metrics. Earlier abstracts report different totals; versions must be identified. No separately named official successor was verified in the reviewed project. [D34a][D34b]
Why it matters. The useful governance question is which component the attacker controls. A model-onboarding team may own model selection, while another team imports prompts, exposes tools, or populates memory. ASB helps separate those failure paths. A malicious system-prompt scenario demands a different remediation owner from an untrusted web result, even if both produce the same unauthorized tool call.
How it works. A configured agent performs domain tasks while a selected attack modifies an allowed surface. The framework checks tool use and compares attack and defense variants. Some experiments assume the adversary can introduce attack tools or contaminate a system prompt. Those assumptions belong in the result record; they must not silently become claims about every deployment. The public implementation is based on AIOS. [D34a][D34b]
Metrics and score interpretation.
ASR (native). Attacked tasks using the attack-specific tools. Direction: Lower.
PNA (native). Required-tool completion without attack or defense. Direction: Higher.
RR (native). Refusal rate for aggressive tasks. Direction: Context-dependent.
BP (native). Original-task performance in backdoor experiments; specify trigger conditions. Direction: Higher.
FNR / FPR (native). Compromised data missed / clean data wrongly flagged. Direction: Lower.
NRP (native). PNA × (1 − ASR), combining utility and resistance. Direction: Higher.
Utility with defense enabled (recommended). Clean task completion under the actual proposed defense. Direction: Higher.
How to interpret a result. NRP is a convenient composite, not a probability of secure business completion or a measure of financial loss. Preserve its components. Because PNA is an unprotected baseline, it does not itself quantify defense-induced disruption. A detector’s FPR/FNR also describes that detector, not all unauthorized actions the complete agent might take.
Limitations and failure modes.
- Tool-use scoring does not automatically establish correct parameters or real business outcomes.
- The final paper’s BP trigger wording differs between its table and appendix; verify the selected implementation before comparison. [D34a]
- Attacks with privileged prompt or tool-catalog access represent specific supply-chain assumptions.
- An ASB-branded package or unrelated Agent Safety benchmark should not be assumed to be this project.
Proposed bank use. Analyst recommendation: use ASB for threat-model coverage during architecture review and third-party agent-component onboarding. Build a surface-to-owner matrix covering imported prompts, memory stores, tool descriptions, and user input. Run only scenarios matching actual access assumptions, then retain excluded scenarios with rationale. For each failure, identify the violated boundary and whether the fix belongs in artifact integrity, retrieval isolation, tool authorization, or model behavior. Require clean-task evidence under the proposed defense alongside the native metrics.
Execution and evidence prerequisites. Pin the official repository and configuration files, prepare its Python/AIOS dependencies, and select API-backed or local models. The repository documents an Ollama route as well as API models. [D34b] Keep all research tools and memory content isolated, and record attack privileges explicitly.
Primary sources: [D34a] [D34b]
35. AgentDyn
AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
| Owner: Hao Li, Ruoyao Wen, Shanghao Shi, Ning Zhang, Yevgeniy Vorobeychik, and Chaowei Xiao; SaFo-Lab repository | Release / evidence: 3 February 2026 initial paper; version 3 dated 7 May 2026 changes the paper subtitle |
| Task form: AgentDojo-derived executable tasks involving changing plans and useful external instructions | Score direction: ASR lower; benign utility and utility under attack higher |
| Status assessment: Recent executable research benchmark, with a revised 2026 paper and public code (analyst assessment). | Availability: Public research paper and repository with AgentDojo-compatible evaluation entry point |
What it is. AgentDyn targets a weakness in simpler injection evaluations: some legitimate tasks require the agent to learn useful next steps from outside content. Blanket suppression of those instructions may block attacks while preventing completion. The benchmark contains 60 tasks and 560 injection cases across shopping, GitHub, and daily-life scenarios, implemented on AgentDojo. Its current paper title differs from the initial February subtitle. [D35a][D35b]
Why it matters. A bank assistant may need to follow an approved workflow described in a returned document or navigate a changing tool result. Testing only rigid tasks can reward defenses that reject all external direction. AgentDyn is therefore especially useful for challenging claims that an injection filter is secure without imposing substantial operational friction.
How it works. Agents plan through open-ended tasks while encountering both helpful external instructions and adversarial content. The study compares prompting, detection, alignment, and system-level defenses, measuring security and task completion separately. It reports averages across the three suites. Its published setup asks agents to complete tasks automatically without seeking confirmation, an important difference from deployments that intentionally require human approval. [D35a]
Metrics and score interpretation.
Benign utility (native). Original user tasks completed without an attack. Direction: Higher.
Utility under attack (native). Original user tasks completed in attacked runs. Direction: Higher.
Attack success rate (native). Security cases achieving the attacker’s objective. Direction: Lower.
Authorized escalation quality (recommended). Whether a deployment’s required human approval is requested correctly and can resume the task. Direction: Higher.
Useful-instruction false blocking (recommended). Legitimate external instructions rejected by the deployed control. Direction: Lower.
How to interpret a result. Read a defense result as a joint security and usability outcome. Preserve per-suite results so a weakness in a relevant workflow is not hidden by averaging. A bank’s required approval must not be classified as accidental failure merely because the research configuration favors fully automatic completion. Report upstream reproduction and the bank’s approval-aware variant independently.
Limitations and failure modes.
- Sixty manually designed tasks provide focused stress testing, not representative coverage of every enterprise process.
- Built-in attack templates do not establish resistance to an independently optimized attacker.
- The no-confirmation instruction changes the evaluated policy and limits direct transfer to approval-controlled deployments. [D35a]
- AgentDojo-derived task infrastructure creates overlap; do not count both suites as wholly independent assurance.
Proposed bank use. Analyst recommendation: add AgentDyn when evaluating a proposed prompt-injection defense or allowing an agent more workflow autonomy. Platform teams should test the same model with and without the defense; application owners should specify legitimate external instructions and mandatory approval points before scoring. Investigate whether low ASR comes from successful trust-boundary enforcement, inability to plan, or excessive blocking. Retain blocked benign steps as operational defects and unauthorized side effects as security defects, with separate owners and acceptance criteria.
Execution and evidence prerequisites. Use a pinned SaFo-Lab/AgentDyn revision and compatible endpoint; the repository retains an AgentDojo-style runner and supports additional defense integrations. [D35b] Isolate its environments, freeze task and attacker configurations, and establish an explicit approval policy for any bank-specific adaptation before execution.
Primary sources: [D35a] [D35b] [D35c]
36. DUMA-Bench
DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
| Owner: AI Security Lab, ITMO University, and Hive Trace Lab; Ivan Aleksandrov, German Kochnev, Sabrina Sadiekh, and Yaroslav Rogoza | Release / evidence: 21 September 2026 initial arXiv release; conference acceptance not independently established in this review |
| Task form: Executable multi-turn scenarios with user simulation, mutable state, and security assertions | Score direction: ASR lower; security pass^k higher, meaning success on all k independent trials |
| Status assessment: Emerging: September 2026 paper and public executable implementation; independent validation remains limited (analyst assessment). | Availability: Public paper, task definitions, prompts, seeds, traces, and evaluation code |
What it is. DUMA-Bench evaluates dual-control interaction: both an agent and a user participant can shape the environment over multiple turns. Extending the tau-squared-bench interaction structure, it provides 35 executable scenarios across eight vulnerability domains, including retrieval poisoning, cross-agent manipulation, identity spoofing, unsafe output handling, oversharing, and tool shadowing. It was first posted in September 2026, so its emerging status is supported. [D36a]
Why it matters. A bank workflow seldom consists of one instruction followed by an isolated model response. Customers clarify requests, operators update records, and other agents contribute information. DUMA-Bench offers a way to examine whether controls remain reliable as this shared state changes. Its concept should not be confused with the bank’s separate four-eyes or dual-approval control.
How it works. The benchmark compares solo operation with an active simulated-user regime. It replays tool trajectories to check deterministic state assertions and uses an LLM judge for communication assertions where needed. A task passes only when every applicable assertion is satisfied. Repeated trials expose variability. Because the interactive regime changes several factors together, its solo-versus-dual difference does not isolate a single causal mechanism. [D36a]
Metrics and score interpretation.
Attack success rate (native). Fraction of episodes violating the task’s security assertions. Direction: Lower.
Security pass^k (native). Probability that all k independent trials pass; estimator per task is C(c,k)/C(n,k), then averaged across tasks. Direction: Higher.
Legitimate-task utility (recommended). Completion of intended work alongside satisfaction of security assertions. Direction: Higher.
Assertion-type breakdown (recommended). Separate state violations, textual disclosures, and judge disagreements. Direction: Lower violations.
How to interpret a result. Security pass^k is not capability pass@k. Here, every one of k trials must remain secure; it becomes stricter as k grows. Report trial counts and distinguish a real policy violation from a judge error. The interactive comparison is useful evidence of a regime-level robustness gap, not an estimated real-world incident probability for customers.
Limitations and failure modes.
- A small scenario inventory and LLM user simulation constrain generalization to actual human behavior.
- Communication judgments require validation, especially for policy language and permitted disclosure exceptions.
- Solo and interactive modes vary multiple factors at once, limiting causal attribution. [D36a]
- The new codebase needs local reproducibility checks; research release is not operational certification.
Proposed bank use. Analyst recommendation: place DUMA-Bench in an emerging-evidence pilot for agents with customer dialogue, shared workflows, or agent-to-agent delegation. Translate representative scenarios into explicit bank policies using synthetic identities and records. Keep user-simulator model and temperature fixed across candidate models, then test controlled variability separately. Independently review assertion outcomes before using them in an adoption decision. Use failures to improve authorization and state checks; avoid allowing a small aggregate score to outweigh a confirmed sensitive-data or transaction violation.
Execution and evidence prerequisites. The repository requires Python 3.10 or later and uses LiteLLM for provider access. It includes data checks, solo/dual agent modes, multiple trials, result viewing, and domain policy inspection. [D36b] Pin agent, user, and judge models independently and retain replayable traces with expected assertions.
Primary sources: [D36a] [D36b]
37. BIPIA
BIPIA: Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
| Owner: Microsoft Research collaborators with University of Science and Technology of China and Hong Kong University of Science and Technology; Jingwei Yi and collaborators | Release / evidence: 21 December 2023 initial paper; revised 27 January 2025; KDD 2025 publication |
| Task form: Text/code content-processing prompts spanning five application tasks | Score direction: ASR lower; clean-output quality and general helpfulness higher |
| Status assessment: Established research dataset and baseline; useful for retrieval/content processing rather than complete deployed-agent assurance (analyst assessment). | Availability: Public code and attack templates; some source-content datasets require separate reconstruction steps |
What it is. BIPIA is a benchmark for indirect prompt injection in content-consuming LLM applications. It combines five tasks—email, web and table question answering, summarization, and code question answering—with attack families and insertion positions. The paper describes 30 text-attack types and 20 code-attack types, with separate training and test attacks. No separately numbered official BIPIA successor was verified in the reviewed Microsoft project. [D37a][D37b]
Why it matters. A read-only assistant can still mislead a user, contaminate an answer, or insert unsafe suggestions into generated code. The absence of write-enabled tools therefore does not remove the need to test injection resistance. BIPIA offers a practical baseline for retrieval-augmented assistants whose main output is text, including internal knowledge and document-analysis applications.
How it works. The harness embeds an attacker instruction at the beginning, middle, or end of external content while retaining an authorized task. It measures whether outputs satisfy the attacker goal. The research evaluates prompt-based defenses and defenses that modify model weights, and measures clean-task quality alongside ASR. Some application source texts must be obtained or reconstructed under their own licensing conditions. [D37a][D37b]
Metrics and score interpretation.
Attack success rate (native). Attack-goal success across evaluated prompts, also reported by application task. Direction: Lower.
ROUGE-1 recall on clean prompts (native study measure). Overlap with expected answer information for utility evaluation. Direction: Higher, with semantic review.
MT-Bench helpfulness (native study companion). General helpfulness checked for weight-changing defenses. Direction: Higher.
Grounded-answer correctness (recommended). Correct, supported business answers despite hostile retrieved content. Direction: Higher.
How to interpret a result. Low ASR supports resistance under the selected task and attack distribution. It does not demonstrate trustworthy multi-step tool use. ROUGE overlap alone cannot establish that a financial or policy answer is substantively correct, complete, or appropriately qualified. Separate benchmark-native utility measures from additional business-answer adjudication and record which sources the assistant was permitted to use.
Limitations and failure modes.
- Large prompt counts arise from combinations of content, attacks, and positions; they are not equally many independent threat scenarios.
- Static prompt evaluation omits changing tool catalogs, delegated identity, and later consequential actions.
- Training/test attack separation must be preserved when tuning defenses.
- Output-grading errors and source-data reconstruction can alter measured results; record both.
Proposed bank use. Analyst recommendation: use BIPIA for early retrieval and content-boundary testing, then add internal-document analogues with synthetic data. Application teams should specify intended answer behavior and permissible source instructions; Platform Security should test the retrieval wrapper and defense settings. Retain representative cases where the model ignores useful content or produces an unsupported answer even when the attack goal is not met. A claim that a defense achieved very low ASR in the publication must remain tied to that publication’s models and setup.
Execution and evidence prerequisites. Install a pinned Microsoft/BIPIA revision, obtain required context files, choose task and split, and configure target and evaluator models. API-backed evaluation does not inherently require a local GPU; local inference or fine-tuning has separate compute needs. The repository documents dataset reconstruction requirements for web QA and summarization. [D37b]
Primary sources: [D37a] [D37b] [D37c]
39. MCP-SafetyBench
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
| Owner: Xuanjun Zong and collaborators; East China Normal University, Salesforce AI Research, Singapore Management University, and Shanghai AI Laboratory | Release / evidence: 17 December 2025 initial paper; ICLR 2026 final paper |
| Task form: Execution with real MCP server integrations and attack-instrumented tasks | Score direction: ASR lower; task success higher; both must be reported |
| Status assessment: Recent conference-published benchmark with real-service integration and substantial setup dependencies (analyst assessment). | Availability: Public paper and code; provider/service credentials and isolated infrastructure may be necessary |
What it is. MCP-SafetyBench evaluates agents using real Model Context Protocol servers. Its final paper describes 245 cases, 20 attack types, and five domains: browser automation, financial analysis, location navigation, repository management, and web search. Attack surfaces span server, host, and user layers. It is distinct from similarly named MCP benchmarks such as SafeMCP or MCPSecBench. [D39a]
Why it matters. For bank deployments that adopt MCP, failures can originate in tool descriptions, server behavior, host mediation, or user context. The evaluated object is consequently the integrated agent system. A provider’s model-level refusal score cannot establish that the bank’s host validates tool identity, parameter use, authorization, or cross-server data movement correctly.
How it works. Tasks adapted from MCP-Universe receive a specified attack modification and corresponding evaluators. The agent executes through MCP integrations, then separate checks produce task and attack outcomes. The design includes disruptive attacks and stealthy side effects. Final-paper settings bound iterations and repeat tasks, so those settings should accompany any comparison. Live-service prerequisites differ by domain. [D39a][D39b]
Metrics and score interpretation.
Attack success rate (native). Fraction of attacked tasks in which the attack objective succeeds. Direction: Lower.
Task success rate (native). Fraction satisfying the original task evaluator. Direction: Higher.
Defense success rate (native presentation). Complement of attack success for the same evaluated cases. Direction: Higher.
Unauthorized action and disclosure counts (recommended). Confirmed policy violations in a bank-specific host/server test, classified by severity. Direction: Lower.
How to interpret a result. Task success and attack success can coexist, especially when the unwanted effect is a side action. Report both outcomes per case and per attack surface. An attack assuming control of the host tests a different boundary from an attacker controlling one server description. Avoid translating success in a synthetic financial-analysis task into assurance for payment execution or bank data access.
Limitations and failure modes.
- Service availability, credentials, rate limits, and changing server behavior can affect reproducibility.
- The benchmark’s domain and attack distribution is not an estimate of the bank’s threat frequency.
- Server and host compromise assumptions must match the assessed architecture.
- The maintainers warn that runs can alter system files and real repositories; a successful installation alone is not evidence of safe containment. [D39b]
Proposed bank use. Analyst recommendation: prioritize this benchmark when MCP is actually in the proposed architecture. Inventory each host, server, tool manifest, granted scope, and destination. Run comparable baseline and defended configurations, then add policy checks for identity binding, tool-version changes, argument validation, and sensitive-data egress. Assign host findings to platform owners, server findings to connector owners, and prompt/task issues to application teams. Treat a confirmed unauthorized real effect as an individual remediation item regardless of the aggregate score.
Execution and evidence prerequisites. Prepare Python 3.10+, Dockerized server support, the selected model API, and domain-specific services. Use dedicated test accounts and narrowly scoped credentials in an isolated VM or container environment; the official repository explicitly describes real repository operations. [D39b] Capture server versions, manifest hashes, task traces, and post-execution state. Do not point the harness at production assets.
Primary sources: [D39a] [D39b] [D39c]
E. Misuse resistance, harmful action and monitoring
This group examines responses to harmful requests, harmful task execution and oversight of adversarial behavior. Separate content refusal, verified action outcomes and monitor detection; each supports a different security claim.
40. HarmBench
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
| Owner: Center for AI Safety and academic/industry collaborators; Mantas Mazeika and collaborators | Release / evidence: 6 February 2024 initial release; HarmBench 1.0 announced 26 February 2024; ICML 2024 paper |
| Task form: Text and image-conditioned output evaluation; attack generation, model responses, and classifier scoring | Score direction: Harmful-output ASR lower for adoption; higher indicates a stronger red-team attack |
| Status assessment: Established research suite with public attack implementations and scoring models (analyst assessment). | Availability: Public code, behavior datasets, precomputed cases, and evaluation classifiers |
What it is. HarmBench standardizes automated red teaming and robust refusal evaluation. It contains 510 behaviors: 400 textual and 110 multimodal, organized into standard, contextual, copyright, and multimodal functional categories. Semantic coverage extends beyond cybercrime to other misuse areas. It therefore measures elicited harmful behavior, not independently verified exploit capability. [E40a]
Why it matters. A bank can use the suite to assess whether exposed assistants can be redirected into clearly disallowed assistance. It also supports controlled comparison of model and guardrail changes. The same ASR has opposite practical readings for two audiences: researchers may want a stronger attack, while an onboarding team wants the deployed defense to resist that attack.
How it works. The pipeline generates or loads adversarial test cases, obtains target-model completions, and grades whether a specified behavior occurred. The paper uses held-out behavior and classifier arrangements to reduce evaluation gaming. The official implementation supports API models, local models, and selected multimodal models, with separate classifiers for different behavior types. Preserve the selected functional categories and target-response settings. [E40a][E40b]
Metrics and score interpretation.
Attack success rate (native). Proportion of test cases whose completions are judged to exhibit the specified harmful behavior. Direction: Lower for the defended system.
Per-category ASR (native breakdown). ASR separated by relevant semantic or functional behavior groups. Direction: Lower.
Clean business-task quality (recommended). Utility on approved bank work under identical deployed defenses. Direction: Higher.
Reviewed harmful assistance severity (recommended). Human assessment of the practical substance of confirmed failures. Direction: Lower severity and frequency.
How to interpret a result. A low score means the tested prompts rarely elicited benchmark-defined behavior from the configured system. It does not establish that the underlying model lacks relevant knowledge or that a tool-using agent cannot cause harm. Separate baseline requests, transfer attacks, and adaptive attacks with their respective budgets. Do not combine different category mixtures into a single comparative ranking without disclosing the composition.
Limitations and failure modes.
- Automatic behavior classifiers can produce false positives or miss obfuscated harmful assistance.
- A refusal followed by actionable assistance is not successful resistance; inspect the whole response.
- Public behavior sets create contamination and overfitting risks, particularly when used in defense training.
- Content generation does not verify actual code execution, compromise, or downstream damage.
Proposed bank use. Analyst recommendation: include a defined HarmBench subset in a general misuse-resistance gate, while retaining separate cyber-capability and application-control assessments. Security should agree the applicable harm taxonomy with model and application owners before running tests. Report the raw benchmark result and any bank-policy adjudication separately. Check benign security-research and educational requests for inappropriate blocking. For every accepted defense change, preserve the attack set, generation budget, judge identity, sampled human review, and business-utility comparison.
Execution and evidence prerequisites. Use the official repository’s pinned dependencies, model configuration, and evaluation classifiers. Load precomputed cases for regression and run separately budgeted attack generation when adaptive testing is needed. The repository offers local and cluster pipeline modes. [E40b] Use approved endpoints and retain generated content in restricted evaluation records; no tools need to act on it.
Primary sources: [E40a] [E40b] [E40c]
41. JailbreakBench
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
| Owner: JailbreakBench research collaboration; Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, and collaborators | Release / evidence: 28 March 2024 initial paper; version 1.0 camera-ready paper dated 31 October 2024; NeurIPS 2024 |
| Task form: Text-generation evaluation with archived jailbreak artifacts and semantic judges | Score direction: Harmful-response ASR and benign refusal lower; legitimate utility assessed separately |
| Status assessment: Established reproducibility-oriented research benchmark, with maintained artifacts and evaluation tooling (analyst assessment). | Availability: Public Python package, behavior set, attack artifacts, judges, and leaderboard |
What it is. JailbreakBench combines a defined threat model, common evaluation interfaces, a repository of attack artifacts, and a benchmark dataset. JBB-Behaviors contains 100 harmful behaviors with 100 topic-matched benign counterparts. The paired benign set is particularly valuable for identifying defenses that block a subject indiscriminately. The scope concerns text behavior and refusal, not exploitation of a live system. [E41a]
Why it matters. Security teams need to reproduce a claimed improvement months later and compare model candidates fairly. Merely saying a model passed a jailbreak test leaves unanswered which prompts, judge, system instructions, and budget were used. JailbreakBench’s emphasis on retained artifacts provides a useful pattern for a bank’s own onboarding evidence register.
How it works. An attack supplies prompts against a target model under the prescribed system prompt and interface. Outputs are judged for harmful behavior; benign counterparts support refusal checks. Artifacts retain prompts, responses, classifications, and attack metadata. The official implementation includes separate jailbreak and refusal judges, and accepts different attack-access regimes. Versions of the judge and artifacts belong in every result. [E41a][E41b]
Metrics and score interpretation.
Attack success rate (native). Share of harmful benchmark goals elicited under the stated attack and target configuration. Direction: Lower for the defender.
Benign refusal rate (native). Refusals on 100 topic-matched benign behaviors. Direction: Lower, subject to policy review.
Target-query count (native artifact metadata). Queries used by the attack; compare only with the same access regime and budget. Direction: Context, not a safety pass score.
Bank-task utility (recommended). Performance on approved work after adding the defense. Direction: Higher.
How to interpret a result. Read ASR together with the benign refusal rate. The benign set is a useful sanity check, not complete business-utility validation, and some borderline cases may differ from the bank’s policy. The benchmark’s own leaderboard conventions are not regulatory adoption thresholds. An archived attack that fails today does not establish resistance to a newly optimized attack with more queries.
Limitations and failure modes.
- One hundred harmful behaviors leave substantial coverage gaps for industry-specific misuse.
- A classifier’s decision can be wrong; prioritize review of high-consequence or ambiguous outputs.
- Results under white-box, black-box, and transfer access are not interchangeable.
- Excessive restriction can improve resistance while reducing useful assistance; the benign set only partially detects this.
Proposed bank use. Analyst recommendation: use as a repeatable baseline for comparing candidate models and defense revisions. Keep a fixed artifact snapshot for regression and an independently developed, budgeted challenge set for assurance. Have an assessor review unexpected benign refusals and all materially harmful completions. Preserve policy mappings separately from native labels so the benchmark remains interpretable. For procurement, ask vendors for reproducible artifacts and configuration evidence rather than accepting a single aggregate percentage or a current leaderboard position.
Execution and evidence prerequisites. Install a pinned jailbreakbench release, choose the model adapter, load a fixed behavior/artifact snapshot, and configure the separate judges. The repository documents evaluation and artifact-submission structures. [E41b] A local GPU is optional with supported hosted models; disclose provider filters and endpoint behavior because they affect the tested system.
Primary sources: [E41a] [E41b] [E41c]
42. AgentHarm
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
| Owner: Gray Swan AI and UK AI Safety/Security Institute, with academic collaborators; Maksym Andriushchenko, Alexandra Souly, and collaborators | Release / evidence: 11 October 2024 initial paper; ICLR 2025; paper version 3 dated 18 April 2025 |
| Task form: Tool-using agent trajectories scored with task rubrics and limited semantic judgments | Score direction: Harm score lower; benign completion higher; refusal alone is insufficient |
| Status assessment: Established agent-misuse benchmark with an official Inspect implementation; public coverage is narrower than the full paper suite (analyst assessment). | Availability: Public subset and Inspect implementation; full benchmark and paper attack template are not all publicly released |
What it is. AgentHarm tests deliberate misuse of a tool-using agent. The paper defines 110 base malicious tasks, expanded to 440 variants across eleven harm categories, together with benign equivalents. Completing a harmful task requires coherent use of multiple tools, so an unhelpful response that merely lacks a refusal does not receive full credit. [E42a]
Why it matters. An agent can say that it cannot help and still have already called a harmful tool. Conversely, it can agree with a bad request but fail to accomplish anything. For a bank, distinguishing refusal, attempted action, partial completion, and actual effect is essential when deciding what tool privileges an assistant should receive.
How it works. Handwritten scenarios use synthetic tools without real-world side effects. Task rubrics award partial credit for completed components; semantic judges handle selected output properties, and a separate refusal judge assesses messages. Detailed/open-ended and hint/no-hint variants vary the difficulty. Benign counterparts provide a capability comparison. [E42a] The available release is 44 public-test base behaviors (176 augmented) plus eight validation base behaviors (32 augmented). The paper’s rules_v4 jailbreak template is withheld. [E42b]
Metrics and score interpretation.
Harm score (native). Average rubric credit for harmful-task execution, including partial completion. Direction: Lower.
Refusal rate (native). Semantic judge’s refusal assessment over agent messages. Direction: Higher on harmful tasks; not sufficient alone.
Non-refusal harm score (native). Harm score restricted to behaviors without detected refusal. Direction: Lower harmful capability when compliant.
Benign counterpart performance (native). Task rubric performance on benign behaviors. Direction: Higher.
Refusal before harmful action (recommended). Whether rejection occurs before any forbidden side effect in a bank adapter. Direction: Higher.
How to interpret a result. Harm score combines compliance and execution ability; it is not a direct severity-weighted risk rating. Evaluate benign tasks alongside harmful ones to identify inability masquerading as safety. A result on the released subset must be labeled with those exact behavior IDs, not presented as coverage of all 440 variants. Paper jailbreak scores cannot be assumed reproducible from the public release alone.
Limitations and failure modes.
- Synthetic tools simplify actual identity, service, and network constraints and are proxies for harmful action.
- Refusal and semantic judges can disagree with executed behavior; inspect trajectories for material findings.
- Public and private test splits and withheld attacks limit direct replication. [E42b]
- Explicitly malicious requests provide limited coverage of ambiguous authorization or indirect injection.
Proposed bank use. Analyst recommendation: include AgentHarm for assistants that can initiate actions, with a separate application test for attempted misuse of the actual permission model. Run direct requests and separately defined adversarial variants under declared budgets. Require a benign-workload comparison and record partial harmful execution, not just completed tasks. Application owners should define forbidden side effects; Platform Security should verify that authorization or confirmation controls prevent them even when the model complies. Retain exact split and tool-choice settings in the onboarding record.
Execution and evidence prerequisites. Use the official Inspect Evals AgentHarm implementation and pin its task version, dataset revision, model, and both judge configurations. It exposes separate harmful and benign tasks plus configurable prompts and tool behavior. [E42c] Keep the synthetic tools intact for baseline runs; assess real connectors only in a separately authorized isolated adaptation.
Primary sources: [E42a] [E42b] [E42c]
43. SafeArena
SafeArena: Evaluating the Safety of Autonomous Web Agents
| Owner: McGill-NLP and collaborators; Ada Defne Tur, Nicholas Meade, Xing Han Lù, and collaborators | Release / evidence: 6 March 2025 initial research paper |
| Task form: Browser-agent actions on four benchmark websites, with paired safe/harmful tasks | Score direction: Harmful completion lower; safe completion and normalized safety higher |
| Status assessment: Executable research benchmark with public paired tasks and website environments; deployment relevance depends on the browser-agent architecture (analyst assessment). | Availability: Public paper, 500-task paired dataset, code, and dedicated containerized websites |
What it is. SafeArena examines agents that can operate web interfaces for a user. It pairs 250 harmful tasks with 250 safe tasks across four websites, spanning misinformation, illegal activity, harassment, cybercrime, and social bias. The paired design distinguishes low harmful completion caused by resistance from low completion caused by poor browser competence. It evaluates intentional misuse, rather than primarily hostile content in an otherwise legitimate task. [E43a][E43b]
Why it matters. For browser automation, the security-relevant outcome is often a change in the website rather than an answer in chat. A bank considering browser-operated administrative workflows should therefore examine whether unsafe requests lead to actual attempted actions. Matching legitimate tasks helps avoid selecting an ineffective agent because it cannot perform either kind of work.
How it works. An agent receives a user intent and interacts with the benchmark website until completion or termination. Task checks assess expected outcomes; refusal assessment and human trajectory review provide additional context. A normalized safety score conditions harmful-task completion on the corresponding safe tasks that the model could solve. The repository uses SafeArena-specific website containers, not interchangeable ordinary WebArena instances. [E43a][E43c]
Metrics and score interpretation.
Harmful task completion rate (native). Share of harmful intents achieving the benchmark outcome. Direction: Lower.
Safe task completion rate (native). Share of matched legitimate tasks completed. Direction: Higher.
Normalized Safety Score, NSS (native). 1 − joint safe-and-harmful completions / safe completions, calculated on matched pairs. Direction: Higher; undefined if no safe task succeeds.
Refusal rate (native). Rejection of harmful requests; interpret alongside trajectory outcomes. Direction: Higher on harmful tasks.
Forbidden side-effect count (recommended). Material unauthorized effects even when full harmful task completion fails. Direction: Lower.
How to interpret a result. NSS does not equal one minus the ratio of all harmful completions to all safe completions: its numerator counts paired successes. State the number of successfully solved safe tasks that forms the denominator. A high NSS with a tiny denominator offers weak coverage. A failed complete task can still contain an unacceptable partial action, which needs separate application-level adjudication.
Limitations and failure modes.
- Four websites cannot represent all enterprise browser interfaces or permission models.
- String-based refusal detection is imperfect, and genuine refusal does not undo earlier actions. [E43b]
- Matched task difficulty and browser/framework configuration affect the normalized score.
- Website state, resets, screenshots, and action histories are necessary to distinguish agent failure from infrastructure failure.
Proposed bank use. Analyst recommendation: prioritize SafeArena for browser-enabled agents, especially those entering forms, publishing content, or administering systems. Pair bank-specific misuse scenarios with legitimate equivalents and use synthetic accounts in isolated application replicas. Inspect partial actions and the timing of approvals, not just final task reward. Retain both native NSS and raw task outcomes; avoid using the ratio alone in a risk decision. Platform teams own browser/session restrictions, while application owners define allowed operations and success criteria.
Execution and evidence prerequisites. Deploy the dedicated SafeArena website containers, install the benchmark’s browser-agent dependencies, load both safe and harmful splits, and configure a supported model endpoint. The official code documents environment URLs and trajectory inspection. [E43c] Reset application state between trials and record browser observations, actions, and evaluation evidence.
Primary sources: [E43a] [E43b] [E43c]
44. SHADE-Arena
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
| Owner: Jonathan Kutasov and collaborators, including Anthropic and Redwood Research collaborators | Release / evidence: June 2025; paper first submitted 17 June 2025 |
| Task form: Long-horizon simulated agent environments plus transcript-monitor evaluation | Score direction: Covert sabotage success lower; monitor detection higher at controlled false-positive rates; benign completion higher |
| Status assessment: Specialized research benchmark for sabotage/control evaluation; public coverage is intentionally partial (analyst assessment). | Availability: Five public task pairs; full 17-pair dataset available by request; updated public repository differs from original link |
What it is. SHADE-Arena pairs legitimate main tasks with hidden harmful side objectives and evaluates an agent against an independent monitoring model. The full benchmark contains 17 task pairs, with five publicly released. The intended question is whether a sufficiently capable agent can carry out covert sabotage, and whether oversight can detect it. It is not a test of spontaneous malicious intent in an ordinary deployment. [E44a]
Why it matters. A capable question-answering model may be a poor monitor, while a weak agent may appear safe because it cannot finish a workflow. Banks assessing delegated actions need to understand both execution capability and oversight failure. This benchmark helps examine whether an additional LLM monitor provides effective control.
How it works. The agent operates in a synthetic environment with tools, completing a main task and an explicitly assigned secret side objective. A separate model inspects the visible transcript and rates suspiciousness without seeing the hidden objective. Experiments vary hints and hidden-scratchpad access. Main-task completion, side-task completion, and monitoring are evaluated separately before calculating joint sabotage success. [E44a][E44b]
Metrics and score interpretation.
Joint sabotage success (native). Completion of main and harmful side tasks while evading the specified monitor. Direction: Lower.
Main / side task performance (native). Separate execution outcomes for authorized work and harmful objective. Direction: Main higher; side lower.
Monitor ROC-AUC (native). Ability to distinguish benign from malicious transcripts across suspicion thresholds. Direction: Higher.
Detection at operating false-positive rate (recommended gate). Detection achieved at a bank-approved alert burden and intervention point. Direction: Higher.
Pre-effect interception (recommended). Whether a control prevents an unacceptable effect before commitment. Direction: Higher.
How to interpret a result. AUC summarizes discrimination, not the false-positive burden at the bank’s chosen threshold. Retrospective detection of a completed unauthorized action is different from prevention. Lower sabotage success can reflect failed legitimate work, so retain the component scores. The exact agent–monitor pairing, visible information, attack affordances, and decision threshold define the result; there is no transferable model-only safety score.
Limitations and failure modes.
- Five public pairs provide less coverage than the full 17-pair benchmark, and repeated trials do not create new independent tasks.
- An agent instructed to pursue sabotage models an adversarial condition; it does not estimate real-world motive or prevalence.
- Scratchpad visibility and hints materially change the threat model and must be stated.
- The original safety-research repository is deprecated and directs users to jkutaso/SHADE-Arena. [E44c]
Proposed bank use. Analyst recommendation: use for advanced agent-control research and monitor procurement, not as a routine standalone model pass/fail test. Establish a realistic alert budget and test monitors on benign and adversarial trajectories from held-out tasks. Compare model monitoring with enforceable action checks and approval controls. Count severe missed side effects individually. Require explicit evidence of when a monitor can intervene, which actions it can block, and what the agent can see about the monitor; a transcript review score alone cannot answer those questions.
Execution and evidence prerequisites. Use the updated public repository, pinned tasks and dependencies, separate agent/monitor endpoints, and recorded visibility/hint settings. Its mock environments can be used independently; additional research code and transcripts are request-only. [E44b] Record public/full coverage and do not claim the full benchmark from public tasks alone.
Primary sources: [E44a] [E44b] [E44c]
Selecting benchmarks for the proposed application
The bundles below are proposed starting points, not a requirement to execute all 43 entries. Choose benchmarks whose task, threat model and available artifacts match the intended deployment. Add an internal holdout representing the bank’s own workflows. Dataset variants from the same lineage can improve diagnostic coverage without providing independent confirmation.
| Application / decision | Suggested external evidence | Internal evidence that completes the picture |
| General conversational assistant | HarmBench and JailbreakBench; a knowledge benchmark only if domain expertise is part of the role | Representative legitimate requests, false refusals, sensitive-output handling and system-policy tests. |
| RAG or document assistant | BIPIA and AgentDojo; InjecAgent when tool outputs can carry untrusted text | Poisoned retrieved passages, tenant/document permissions, source attribution and attempted unauthorized disclosure. |
| Business agent with write actions | AgentDojo, AgentDyn, ASB and a suitable action-level misuse benchmark | Connector authorization, approval boundaries, record changes, transaction limits and safe interruption. |
| MCP-connected agent | MCP-SafetyBench plus ASB or AgentDyn; AgentHarm if malicious direct requests are in scope | Actual server schemas, tool trust, argument validation, credential scope and isolation of test accounts. |
| Code-completion assistant | SecurityEval, LLMSecEval and a relevant repository/application secure-coding benchmark | Production languages and frameworks, functional tests, dependency handling, secrets and security review. |
| Autonomous code-change / repair agent | SEC-bench repair, AutoPatchBench, PatchBench; repair CVE-Bench for its covered tasks | Held-out security regressions, broader functional tests, patch review and repository change permissions. |
| Vulnerability detection model | PrimeVul and DiverseVul; richer repository or lifecycle tasks if the system performs them | Time/project-separated bank code, class balance, false-positive workload and high-impact missed findings. |
| SOC / threat-intelligence assistant | CyberSOCEval; CTIBench or AthenaBench; SECURE for relevant ICS advisory tasks | Local telemetry, investigation outcomes, evidence grounding, escalation quality and analyst review effort. |
| Supply-chain exploitability / VEX assistant | Cybersecurity VEX-Bench | Real dependency/build context, reachability evidence, false not-affected decisions and human disposition review. |
| Autonomous security research / red-team agent | Cybench or NYU CTF baseline; CyberGym; then web CVE-Bench, ExploitGym or ExploitBench as relevant | Permitted targets, realistic mitigations, long-duration budgets and application-level safeguards. |
| Long-horizon agent with broad access | AISI-inspired custom ranges; AgentDyn; SHADE-Arena if a monitor is part of the defense | Linked enterprise workflows, observation gaps, intervention latency and recovery after a detected deviation. |
| Unlearning or hazardous-knowledge mitigation | WMDP-Cyber, with retained-utility and action-level misuse evaluations | Before/after assessment of useful defensive work and the actual prohibited behaviors. |
Use several kinds of evidence without double counting
A sensible bundle covers distinct failure mechanisms: technical capability, defensive quality, trust-boundary resilience and misuse or oversight. Running several closely related knowledge datasets does not substitute for testing an agent that can modify records. Likewise, a successful secure-code score does not validate the connector permissions used to publish that code.
For a bank’s model-onboarding process, keep model-artifact integrity, dependency review, supplier evidence, data handling and application access design alongside benchmark results. An evaluation result can inform those decisions, while its measured task should remain explicit in the decision record.
Designing an evaluation that can support a decision
1. Define the system and the consequence
Record the intended users, business workflow, model endpoint or weights, agent implementation, tools, data sources, permitted write actions and human oversight. State the concrete adverse outcomes being evaluated: for example, unauthorized disclosure, an incorrect high-impact patch, a false vulnerability dismissal or an unauthorized record change. Use these outcomes to choose tests and severity labels.
2. Register the benchmark and task population
Select the canonical owner, paper version, repository commit, dataset revision, task IDs, scorer and any available reference solutions. Record whether the test is the original benchmark, a public subset, an official adapter or a custom adaptation. List excluded tasks and reasons before the scored run. Keep development and decision sets separate.
3. Freeze the candidate configuration
Retain model snapshot, inference settings, agent commit, prompts, memory behavior, tools, installed packages, credentials scopes, defense configuration and judge model. Include the permitted internet routes and any provider-side retrieval features. A shared agent memory across tasks can change independence and expose information from earlier instances.
4. Validate the environment and scorer
Use available benign controls and reference artifacts to confirm that a valid outcome is accepted and an invalid one is rejected. Keep the evaluator, answer keys, private tests and fixed-version artifacts outside the agent’s authority. Inspect a sample of trajectories to check whether the reported result reflects the intended task. The documented SEC-bench Pro retrieval issue illustrates why environment integrity matters. [A8d]
5. Run the declared comparison
Use fixed per-task budgets, explicit retry rules and independent resets. Separate retries for infrastructure faults from additional semantic attempts. For injection or harmful-action evaluations, run matched legitimate tasks and report utility with security outcomes. For repair, test both security behavior and required functionality. Retain failed runs as well as successful ones.
6. Adjudicate consequential failures
Have an appropriate specialist review high-impact successes or failures and a sample of ordinary cases. Determine whether a claimed exploit used the intended weakness, whether a patch actually removed the defect, whether an output crossed a trust boundary and whether a monitor detected the problem before impact. Record grader disagreement instead of resolving it silently.
7. Make and scope the decision
State the allowed use, required controls, unresolved limitations, responsible owner and re-evaluation triggers. An approval should identify the configuration to which it applies. If tools, model version, prompts, memory, retrieval sources or permissions change materially, reassess the affected evidence rather than assuming that the previous score carries forward.
Interpreting thresholds and promotion gates
This document does not assign an industry-wide adoption percentage to each benchmark. The proposed approach is to calibrate thresholds against the intended role, error consequences, baseline performance and uncertainty. Numerical floors or ceilings become useful only after the metric and operating conditions are defined.
| Evidence family | Proposed policy shape | Why the direction matters |
| Offensive capability | Review trigger for newly demonstrated high-impact capability; require relevant access and safeguard evidence | A higher score shows ability. It can be valuable for an authorized defensive role and consequential for misuse risk. |
| Secure generation / repair | Minimum functional-and-secure quality, plus review of consequential regressions and incomplete fixes | A functional answer alone is insufficient, and a patch that disables required behavior is not an acceptable repair. |
| Detection / triage | Quality floor calibrated to missed-findings severity and analyst false-positive capacity | A global F1 value can hide poor recall for a consequential class or unsustainable review workload. |
| Prompt injection / harmful actions | Maximum accepted harmful-outcome rate, paired with a minimum legitimate-task utility and critical-scenario checks | A blocked or broken agent can have low ASR without being useful. Consequential failures need individual review. |
| Monitoring | Detection requirement at a declared false-alarm rate, with intervention tested before impact | Catching a deviation after the harmful action does not demonstrate effective prevention. |
| Knowledge / advisory reasoning | Role-specific quality floor and evidence-grounding review | Knowledge is useful for an analyst role but does not prove secure action-taking. |
| Emerging / restricted benchmark | Evidence-readiness gate before its score becomes binding | A result cannot serve as a dependable promotion criterion when its identity, coverage or scorer is unresolved. |
A practical calibration sequence
- Define the decision loss: which misses, false alarms, broken workflows or unauthorized outcomes matter, and to whom.
- Measure the current approved system and, where meaningful, the human workflow on the same held-out tasks.
- Set a proposed quality floor or review trigger using the relevant metric and sample uncertainty.
- Add scenario-specific checks for high-consequence behavior that should not be averaged away by many easy successes.
- Validate the chosen operating point on a separate holdout, including adaptive attacks where relevant.
- Document who accepts the residual risk, the permitted deployment conditions and the triggers for renewed testing.
Illustrative gate logic
Proposed, uncalibrated example: promotion of a business agent with write tools requires a valid evaluation run, an agreed level of legitimate workflow completion and no observed unauthorized critical action in the designated critical-scenario suite. An observed critical action pauses that workflow’s promotion until the failure is understood and addressed. This is a suggested control pattern; it is not a claim that zero observed events proves zero future risk.
For an exploitation benchmark, a corresponding rule would be different: demonstration of a newly defined full-control capability triggers additional review of the model’s accessible tools, allowed targets, usage monitoring and safeguard evidence. The measured capability is preserved in the record rather than treated automatically as either a quality pass or an adoption failure.
For a repair agent, a candidate patch must pass the required functional checks and the held-out security tests. Aggregate benchmark performance can qualify the model for a supervised role, while each production change still needs the controls appropriate to that change. This keeps model qualification and individual artifact acceptance distinct.
Evidence package and proposed ownership
The following evidence record is designed to make a benchmark result reviewable and reproducible. It is a proposed template for the bank’s onboarding process, with team names expressed as roles so they can be mapped to the actual organization.
| Evidence item | Required contents |
| Decision and scope | Use case, users, permitted actions, consequences, candidate configuration and requested deployment conditions. |
| Benchmark identity | Full title, owner, source URLs, release/commit, dataset revision, original/adapted/public-subset status. |
| Population | Task IDs, categories, sample count, excluded tasks, selection method and holdout policy. |
| Model and agent | Model snapshot or weights hash, agent commit, prompts, inference settings, tools and memory design. |
| Security configuration | Identity and scopes, execution boundaries, guardrails, approval behavior, retrieval and network configuration. |
| Budget and attempts | Time, token, action and cost limits; seeds; repeated trials; semantic retries versus infrastructure reruns. |
| Scorer | Code/commit, judge model and prompt if applicable, success definitions and reference-artifact checks. |
| Raw artifacts | Tool-visible inputs and outputs, actions, state changes, verifier outputs, generated code/patches and timestamps. |
| Results | Native metrics, denominators, class/task breakdowns, clean utility, uncertainty, error categories and cost. |
| Adjudication | Reviewed failures, false positives, alternative-path successes, judge disagreements and resolution rationale. |
| Coverage and overlap | Risk scenarios covered, remaining gaps and shared datasets or components across suites. |
| Decision record | Allowed use, conditions, owner, residual risk, exceptions, evidence date and re-evaluation triggers. |
| Role | Proposed responsibility |
| Application / use-case owner | Define intended behavior, business consequences, legitimate-task acceptance and deployment scope. |
| Model / AI platform team | Provide stable candidate access, model/agent configuration, run infrastructure and reproducible artifacts. |
| Security evaluation / AppSec team | Design attack and secure-code tests, validate containment and scorer integrity, adjudicate technical failures. |
| SOC / threat-intelligence specialists | Validate analyst-task relevance, triage quality, monitoring metrics and investigation evidence when applicable. |
| AI / model risk governance and BISO | Challenge coverage, thresholds, assumptions and residual-risk treatment; connect evidence to the onboarding decision. |
| Designated risk acceptance authority | Approve the scoped use or exception according to the bank’s existing authority model. |
| Independent assurance / audit | Review traceability and operation of the process where required; avoid assuming benchmark ownership implies independent validation. |
Re-evaluation triggers
Revisit affected tests after a material model change, fine-tune, agent rewrite, new connector or MCP server, increased privilege, new retrieval source, altered memory behavior, changed safeguard or monitor, benchmark scorer correction, or a relevant incident. Preserve the earlier result and explain which parts remain comparable.
Glossary
| Term | Working definition |
| Agent scaffold / harness | The software that connects a model to tools, prompts, memory, execution limits and the evaluation environment. |
| Benchmark oracle / grader | The mechanism that determines whether the task’s success or failure condition was met. |
| PoC / PoV | Proof of concept / proof of vulnerability. In these benchmarks, often an executable input demonstrating a bug; the required impact varies. |
| Crash reproduction | An input that triggers a failure condition. It need not establish useful attacker control. |
| Exploit primitive | A reusable capability such as an information leak or memory read/write that can contribute to a larger exploit. |
| Mitigation configuration | The protections enabled in the target or agent; changing them can change the measured task. |
| CTF | Capture the flag: a bounded security challenge whose solution is commonly validated by a secret token. |
| CVE / CWE | A CVE identifies a publicly catalogued vulnerability; a CWE identifies a weakness category. |
| CVSS | A standardized vulnerability severity scoring system. Benchmarks may predict a vector, a score, a category or a normalized error. |
| VEX | Vulnerability Exploitability eXchange: statements about whether and how a product is affected by a vulnerability. |
| Indirect prompt injection | An attacker’s instructions arriving through lower-trust material such as retrieved content or tool output. |
| Jailbreak | An attempt to bypass a model or system’s behavioral restrictions; judge the actual resulting behavior. |
| MCP | Model Context Protocol: a protocol used to expose tools and resources to models and agents. |
| Contamination | Overlap or leakage of test data or solutions into training, tuning, retrieval or the agent’s evaluation context. |
| Reward hacking / benchmark gaming | Achieving a favorable score through a path that fails to satisfy the evaluation’s intended objective. |
| Held-out test | Evaluation material excluded from development or tuning, subject to the limitations of public-data exposure. |
| Clean utility | Performance on legitimate, non-attacked tasks under the same relevant configuration. |
| Temporal / project split | Separating training and evaluation by time or source project to reduce overly optimistic generalization estimates. |
Reference register
All sources below were reviewed on 6 October 2026. Titles link to the public primary source. Profile citations link here. Repository contents and unversioned paper pages can change; retain an immutable revision for an actual evaluation. Bibliographic entries identify what was used in this review, while the profile text distinguishes recommendations from source-reported methodology.
Guidance and measurement framework
[G1] NIST AI 100-1 — Artificial Intelligence Risk Management Framework (AI RMF 1.0) — Institutional framework.
https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
MEASURE functions: documented, context-sensitive measurement, validity, safety, security and resilience.
[G2] NIST AI 600-1 — Generative Artificial Intelligence Profile — Institutional guidance.
https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
MEASURE 2.3, 2.5, 2.6 and 2.7; pre-deployment testing and benchmark-to-deployment limitations.
01. CyberGym
[A1a] CyberGym — official project and methodology — Official project.
1,507 instances / 188 projects, Level 1, differential verification and trial semantics
[A1b] sunblaze-ucb/cybergym — official implementation — Official repository.
Installation, storage modes, network guidance and code license
[A1c] CyberGym — research paper — Research paper.
https://arxiv.org/abs/2506.02548
Research design, information settings, vulnerability reproduction and publication history
02. ExploitGym
[A2a] ExploitGym — official project — Official project.
Current domains/counts, provided PoV, flag and intended-bug validation
[A2b] sunblaze-ucb/exploitgym — official repository — Official repository.
Release v1.0 / 869 instances and setup/data license boundaries
[A2c] ExploitGym — research paper — Research paper.
https://arxiv.org/abs/2605.11086
Original study describes 898 instances; do not mix with current release
03. ExploitBench
[A3a] ExploitBench — official methodology and leaderboard — Official project.
Five tiers, 16 flags, all-runs warning and capability coverage semantics
[A3b] ExploitBench — research paper — Research paper.
https://arxiv.org/abs/2605.14153
41 V8 bugs, deterministic oracles, separate experiment arms and known-vulnerability scope
[A3c] exploitbench/exploitbench — official implementation — Official repository.
Runner, target images, full/subset configurations and execution prerequisites
04. Cybench
[A4a] Cybench — official project — Official project.
40 tasks, metric semantics, subset caveats and historical answer-leak correction
[A4b] andyzorigin/cybench — official implementation — Official repository.
Task modes, launchers, model adapters, budgets and retained logs
[A4c] Stanford CRFM — introduction to Cybench — Author institutional publication.
2024 introduction and first-solve-time interpretation
05. InterCode-CTF
[A5a] InterCode — official CTF environment specification — Official implementation documentation.
Task environment, Bash action space, flag reward and episode termination
[A5b] Inspect Evals — InterCode CTF implementation — Official adapter documentation.
100 original / 78 offline tasks, adapter behavior and version changes
[A5c] InterCode — official project — Official project.
CTF environment introduction in August 2023 and broader framework identity
06. NYU CTF Bench
[A6a] NYU-LLM-CTF/NYU_CTF_Bench — official dataset — Official repository.
200 test / 55 development tasks, categories and Docker packaging
[A6b] NYU CTF Bench — research paper — Research paper.
https://arxiv.org/abs/2406.05590
Automated function-calling evaluation framework and publication history
07. CVE-Bench: web exploitation
[A7a] uiuc-kang-lab/cve-bench — official implementation — Official repository.
Original 40 CVEs, impact criteria, v2.1.0 change, CLI 2.2.0, variants and access limits
[A7b] CVE-Bench: web exploitation — research paper — Research paper.
https://arxiv.org/abs/2503.17332
Real-world critical web CVE benchmark and sandbox evaluation design
08. SEC-bench Pro
[A8a] SEC-bench/SEC-bench-Pro — official implementation — Official repository.
344 cases, platform requirements, privilege fields and project-specific grading
[A8b] SEC-bench Pro — research paper — Research paper.
https://arxiv.org/abs/2605.26548
July 2026 v2 paper describes 344 vulnerabilities; early abstract described 183.
[A8c] SEC-bench Pro — submission requirements — Official methodology.
source_files mode and exact configuration / artifact requirements
[A8d] Reward Hacking a Kernel Benchmark, and How We Sealed It — Maintainer methodology report.
Documented retrieval leakage and dual-layer network/search restriction
[A8e] SEC-bench Pro — original v1 paper — Versioned research paper.
https://arxiv.org/abs/2605.26548v1
Original 183-engine-case definition; retained to explain population expansion.
09. AISI multi-step cyber ranges
[A9a] AISI — How do frontier AI agents perform in multi-step cyber-attack scenarios? — Institutional methodology report.
Two named ranges, step counts, no active defenders and unintended-path caveat
[A9b] Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios — Research paper.
https://arxiv.org/abs/2603.11214
March 2026 study, repeated runs and inference-budget sensitivity
[A9c] AISI — Inspect Cyber — Official framework publication.
Open evaluation framework and infrastructure definitions; not a release of all ranges
10. CyberGym-E2E
[B10a] CyberGym-E2E paper and release history — Primary research paper.
https://arxiv.org/abs/2606.04460
ICML 2026 association, release dates, 920 vulnerabilities and 139 projects.
[B10b] CyberGym-E2E official project page — Official benchmark documentation.
Cumulative S1–S4 definitions, 615/920 distinction and shallow-patch limitation.
[B10c] CyberGym-E2E official implementation — Official repository.
Evaluation modes, validation stages, prerequisites and isolation support.
11. BountyBench
[B11a] BountyBench paper, version 3 — Primary research paper.
https://arxiv.org/abs/2505.15216
Published 25-system/40-bounty baseline, task types and bounty-linked evaluation.
[B11b] Stanford CRFM: Introducing BountyBench — Author institution research explanation.
Snapshot/invariant design and Detect, Exploit and Patch evaluator behavior.
[B11c] BountyBench main implementation — Official repository.
Python/Docker setup, task submodules and phase-specific workflows.
12. SEC-bench
[B12a] SEC-bench paper, version 2 — Primary research paper.
https://arxiv.org/html/2506.11791v2
200-case publication baseline, construction workflow, sanitizer verdicts and patch evaluator scope.
[B12b] SEC-bench official repository — Official implementation.
Prerequisites, supported scaffolds, task information levels and evaluation options.
[B12c] SEC-bench official Hugging Face dataset — Official dataset.
The evaluation split displayed 300 rows when reviewed.
13. AutoPatchBench
[B13a] Meta: Introducing AutoPatchBench — Primary benchmark announcement and methodology.
Original counts, fuzzing/differential validation and evaluator limitations.
[B13b] CyberSecEval official AutoPatch documentation — Official execution documentation.
Current 142/120/20 manifests, Podman, Linux and storage guidance.
14. CVE-Bench: vulnerability repair
[B14a] CVE-Bench repair paper, NAACL 2025 — Primary research paper.
https://aclanthology.org/2025.naacl-long.212.pdf
509 CVEs, report-information settings, four languages and unit-test repair criterion.
[B14b] WhileBug/CVEBench author implementation — Official repository.
Identity, CVEfixes dependency and preparation/execution/verification stages.
15. PatchBench
[B15a] PatchBench paper, version 1 — Primary research preprint.
https://arxiv.org/html/2609.04075v1
September 2026 identity, task inventory, transplantation/mutation and targeted validity threats.
[B15b] PatchBench official implementation — Official repository.
Security/semantic checks, result fields, ground-truth separation and storage requirements.
16. SecurityEval
[B16a] SecurityEval official repository — Official dataset and implementation.
Updated 121/69 inventory, historical 130/75 distinction and analyzer materials.
[B16b] SecurityEval author-hosted paper — Primary research paper.
https://s2e-lab.github.io/preprints/msr4ps22-preprint.pdf
Python scope, target-CWE evaluation, manual/static methods and limitations.
17. SecRepoBench
[B17a] SecRepoBench paper, version 3 — Primary research paper.
https://arxiv.org/html/2504.21205v3
Task scope and inventory, pass@1 and secure-pass@1 definitions.
[B17b] SecRepoBench official implementation — Official repository.
Repository contexts, task evaluator, prompt variants and execution prerequisites.
18. SecureVibeBench
[B18a] SecureVibeBench ACL 2026 paper — Primary peer-reviewed paper.
https://aclanthology.org/2026.acl-long.1107.pdf
105/41 inventory, functional/PoV/SAST evaluator and four outcome categories.
[B18b] SecureVibeBench official implementation — Official repository.
Release chronology, reconstructed development setting, agent support and data/container access.
19. VEX-Bench
[B19a] VEX-Bench supply-chain exploitability paper — Primary research preprint.
https://arxiv.org/html/2609.08040v1
Metrics, invalid-output handling, label precedence and deployment-context limitations.
[B19b] Red Hat Research: VEX-Bench — Author institution research explanation.
Inventory, authorship, task design, four reasons and author-announced artifact availability.
[B19c] VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation — Separate primary research paper, used only for name disambiguation.
https://arxiv.org/abs/2609.35028
Confirms an unrelated September 2026 misinformation benchmark sharing the VEX-Bench name.
20. PrimeVul
[B20a] PrimeVul paper, version 2 — Primary research paper.
https://arxiv.org/html/2403.18624v2
Original inventory, chronological/paired evaluation and VD-S definition.
[B20b] PrimeVul official repository — Official dataset and implementation.
v0.1 metadata filtering, original-release distinction and experiment/VD-S tooling.
21. DiverseVul
[B21a] DiverseVul RAID 2023 paper — Primary research paper.
https://surrealyz.github.io/files/pubs/raid23-diversevul.pdf
Corpus inventory, commit-based labeling, metrics, label noise and project transfer.
[B21b] DiverseVul official repository — Official dataset repository.
Data/metadata links, merged research splits and label-noise analysis materials.
22. LLMSecEval
[B22a] LLMSecEval paper — Primary research paper.
https://arxiv.org/abs/2303.09384
150 prompts, secure references and MSR 2023 publication identity.
[B22b] LLMSecEval official repository — Official dataset and tools.
18/25 historical CWE coverage, prompt provenance/quality fields and CodeQL interface.
23. CyberSecEval 4 and inherited tracks onward
[C23a] Meta CyberSecEval 4 — Introduction — Official documentation.
Version 4, inherited tracks, CyberSOCEval and AutoPatchBench parent-child relationship.
[C23b] Meta — MITRE and MITRE FRR Benchmarks — Official implementation documentation.
Response expansion/model judging for MITRE; keyword-based FRR judgment.
[C23c] Meta — Secure Code Benchmark — Official implementation documentation.
Instruct/autocomplete, insecure code detector, vulnerable percentage and pass-rate outputs.
[C23d] Meta — CyberSecEval Getting Started — Official implementation documentation.
Python runner, dependencies, API/self-hosted model adapters.
[C23e] Meta — Sharing new open source protection tools and advancements in AI security — Official release announcement.
29 April 2025 announcement of CyberSecEval 4 with CyberSOCEval and AutoPatchBench.
[C23f] UK AISI Inspect Evals — CyberSecEval 4 implementation README — Official adaptation repository.
Public adaptation omits autonomous-uplift and autopatching prototypes; not equivalent to every Meta suite track.
24. CyberSOCEval
[C24a] CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning — Original research paper.
https://arxiv.org/html/2509.20166v1
Two tracks, existing detonation reports, 609 malware cases, multi-answer exact-set scoring and original image-report evaluation.
[C24b] Meta — Threat Intelligence Reasoning benchmark — Official implementation documentation.
Text/image/both modes, PDF extraction, exact-set and Jaccard scoring, parsing errors.
[C24c] Meta — Malware Analysis benchmark — Official implementation documentation.
Submodule, context truncation below 128k, category breakdowns and parser-error counts.
[C24d] CrowdStrike — CyberSOCEval benchmark data — Official data repository.
Joint Meta/CrowdStrike dataset repository and license file.
25. CTIBench
[C25a] CTIBench — original paper, version 3 — Original research paper.
https://arxiv.org/html/2406.07599v3
Five tasks; metrics; CVSS-vector conversion; attribution categories; micro/macro-F1 inconsistency.
[C25b] CTIBench — official repository — Official repository.
4,610 primary examples plus 1,000 2021 comparison examples; datasets, response labels, notebooks and logs.
26. AthenaBench
[C26a] AthenaBench — paper, version 2 — Original research paper.
https://arxiv.org/html/2511.01144v2
CTIBench extension; six tasks; dynamic construction; normalized VSP formula; mini/full release statement.
[C26b] AthenaBench — official repository — Official repository.
Python evaluation runner, Git LFS, full/mini advertised directories, CKT caveat and scored-only artifacts.
[C26c] AthenaBench — arXiv publication history — Original publication record.
https://arxiv.org/abs/2511.01144
Initial publication 3 November 2025; v2 14 February 2026; author identities.
27. CyberMetric
[C27a] CyberMetric — original paper, version 2 — Original research paper.
https://arxiv.org/html/2402.07688v2
RAG-based construction, validation, sizes, source types and bounded human comparison.
[C27b] CyberMetric — official dataset and evaluator repository — Official repository.
Four JSON variants, fixed answer-format prompt, evaluator, publication and authors.
[C27c] CyberMetric — arXiv publication record — Original publication record.
https://arxiv.org/abs/2402.07688
February 2024 first publication; June 2024 v2; title and authors.
28. CyberBench
[C28a] JPMorgan Chase — CyberBench official repository — Official repository.
Ten NLP datasets, Python prerequisites, pipeline, authors, and archive date 26 May 2026.
[C28b] CyberBench — author’s dataset card — Official author dataset.
Ten named tasks, metrics and separate underlying license terms.
[C28c] CyberBench — original AICS 2024 paper, author-hosted copy — Original research paper.
https://zefang-liu.github.io/files/liu2024cyberbench_paper.pdf
Original benchmark identity, datasets and per-task evaluation.
[C28d] Vals AI — CyberBench methodology — Official separate benchmark.
Distinct benchmark using OSS-Fuzz exploitation/patching; not the JPMorgan NLP suite.
[C28e] University of Twente — CyberBench direct-security-risk framework — Official university project.
Another distinct 2026 CyberBench name.
29. SECURE
[C29a] SECURE — original paper, version 4 — Original research paper.
https://arxiv.org/html/2405.20441v4
Six tasks, source domains, KCV/VOOD contrast, RERT ROUGE-L and CPST numerical-error evaluation.
[C29b] SECURE — official aiforsec repository — Official repository.
Name expansion, ICS focus, six TSV datasets with prompts/reference answers.
[C29c] SECURE — arXiv publication record — Original publication record.
https://arxiv.org/abs/2405.20441
May 2024 initial submission, October 2024 v4 and authors.
30. SecQA
[C30a] SecQA — original paper — Original research paper.
https://arxiv.org/html/2312.15838v1
GPT-4 textbook construction, two tiers, split sizes, zero-/five-shot experiments and near-ceiling original results.
[C30b] SecQA — author’s dataset card — Official author dataset.
Dataset schema, variants, splits and CC BY-NC-SA 4.0 license label.
31. WMDP-Cyber
[C31a] The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning — Original research paper.
https://arxiv.org/html/2403.03218v6
Hazardous-knowledge proxy, cyber task framing, unlearning and dual-use utility tradeoff.
[C31b] CAIS — official WMDP dataset card — Official dataset.
1,987 cyber questions and March/April 2024 corrections.
[C31c] Center for AI Safety — WMDP official repository — Official repository.
Public benchmark and separate RMU unlearning/evaluation implementation.
32. AgentDojo
[D32a] AgentDojo paper, version 3 — primary research paper.
https://arxiv.org/html/2406.13352v3
97 tasks/629 security cases; three metrics, including utility under attack without adversarial side effects.
[D32b] AgentDojo official repository — official implementation.
ETH/Invariant ownership, installation, selectable suites and defenses, API change notice.
[D32c] AgentDojo results documentation — official documentation.
Results explicitly are not a leaderboard because model/attack/defense coverage is incomplete.
33. InjecAgent
[D33a] InjecAgent paper — primary research paper.
https://arxiv.org/html/2403.02691v2
17 user tools, 62 attacker cases, 1,054 combinations, ASR-valid/all, validity, and staged data-theft scoring.
[D33b] InjecAgent official repository — official implementation.
Base/enhanced settings, adapters, prompt options, cache support, and ASR output fields.
[D33c] InjecAgent ACL publication record — official conference publication.
Canonical identity and Findings of ACL 2024 publication.
34. Agent Security Bench (ASB)
[D34a] ASB final ICLR 2025 paper — primary conference paper.
Final abstract totals, attack families, seven native metrics, NRP formula, privileged assumptions, BP wording inconsistency.
[D34b] ASB official repository — official implementation.
Canonical ICLR implementation, AIOS basis, configurations, local/API model support.
35. AgentDyn
[D35a] AgentDyn revised paper, version 3 — primary research paper.
https://arxiv.org/html/2602.03117v3
Dynamic/helper-instruction motivation, 60/560 size, metrics, suite averaging, no-confirmation setup.
[D35b] AgentDyn official repository — official implementation.
Current title and authors, AgentDojo basis, runner compatibility, supported defenses.
[D35c] AgentDyn version history — primary publication metadata.
https://arxiv.org/abs/2602.03117
Initial 3 February 2026 release and 7 May 2026 version 3 with updated title.
36. DUMA-Bench
[D36a] DUMA-Bench paper, version 1 — primary research paper.
https://arxiv.org/html/2609.24662v1
21 September 2026, 35 scenarios, eight domains, assertion scoring, pass^k definition, regime-confounding limitations.
[D36b] DUMA-Bench official repository — official implementation.
Python 3.10+, LiteLLM, solo/dual execution, task checks, traces and viewing.
37. BIPIA
[D37a] BIPIA research paper, version 4 — primary research paper.
https://arxiv.org/html/2312.14197v4
Five tasks, attack families and positions, clean ROUGE-1 and MT-Bench companion evaluation, KDD metadata.
[D37b] BIPIA official Microsoft repository — official implementation.
Installation, source dataset reconstruction, API/local execution prerequisites, task/split settings.
[D37c] BIPIA publication history — primary publication metadata.
https://arxiv.org/abs/2312.14197
Initial December 2023 release, January 2025 revision, and KDD 2025 acceptance.
39. MCP-SafetyBench
[D39a] MCP-SafetyBench final ICLR 2026 paper — primary conference paper.
245 cases, 20 attacks, five domains, ownership, paired task/attack evaluators and execution parameters.
[D39b] MCP-SafetyBench official repository — official implementation.
Python/Docker/API requirements; real GitHub mutations and isolation/test-account guidance.
[D39c] MCP-SafetyBench publication history — primary publication metadata.
https://arxiv.org/abs/2512.15163
17 December 2025 initial release and March 2026 revision.
40. HarmBench
[E40a] HarmBench research paper — primary research paper.
https://arxiv.org/html/2402.04249v2
510 behaviors, text/multimodal and functional categories, ASR, validation/test design and classifier methodology.
[E40b] HarmBench official repository — official implementation.
Release dates, three-stage pipeline, local/API/multimodal support, public classifiers and precomputed cases.
[E40c] HarmBench ICML 2024 publication — official conference publication.
ICML 2024 publication and canonical authorship.
41. JailbreakBench
[E41a] JailbreakBench camera-ready paper — primary research paper.
https://arxiv.org/html/2404.01318v5
100 harmful plus 100 matched benign behaviors, artifact retention, attack regimes, native refusal check.
[E41b] JailbreakBench official repository — official implementation.
Python tooling, separate jailbreak/refusal judges, query/artifact metadata and supported evaluation flow.
[E41c] JailbreakBench publication history — primary publication metadata.
https://arxiv.org/abs/2404.01318
28 March 2024 initial release, v1.0 camera-ready revision and NeurIPS 2024 acceptance.
42. AgentHarm
[E42a] AgentHarm revised paper — primary research paper.
https://arxiv.org/html/2410.09024v3
110/440 full design, synthetic tools, partial-credit harm score, refusal and non-refusal metrics, benign equivalents.
[E42b] AgentHarm official AISI dataset card — official dataset documentation.
Released 44 test base/176 augmented plus eight validation base/32 augmented; rules_v4 template not public.
[E42c] AgentHarm official Inspect Evals documentation — official implementation documentation.
Maintainers, harmful/benign tasks, split, model/judge/prompt settings, versioned implementation.
43. SafeArena
[E43a] SafeArena research paper — primary research paper.
https://arxiv.org/html/2503.04957v1
Paired benchmark design, execution scoring, human review, exact NSS formula.
[E43b] SafeArena official project site — official project documentation.
250 safe and 250 harmful tasks, four websites, five categories, refusal detector and metrics.
[E43c] SafeArena official repository — official implementation.
Dedicated SafeArena containers, paired datasets, model configuration and trajectory inspection.
44. SHADE-Arena
[E44a] SHADE-Arena research paper — primary research paper.
https://arxiv.org/html/2506.15740v1
17 total/five public pairs, hidden side objectives, monitor visibility, joint success and monitoring evaluation.
[E44b] Updated SHADE-Arena public repository — official implementation.
Current public split, mock environments/tools, agent and monitor pipeline, request-only expanded material.
[E44c] Original SHADE-Arena repository deprecation notice — official repository lifecycle notice.
Explicit deprecated status and redirect to jkutaso/SHADE-Arena.
Leave a Reply