Field guide · snapshot 5 October 2026

Language models for finding vulnerabilities

A reading of the community index Awesome LLMs for Vulnerability Detection: 90 papers, six tools, and three agent skills. The short version is that a function-level score stopped being believable, and the work moved toward context, then toward agents.

Curated papers
90
From 2025–26
36
With code
39
Tools + skills
9
Overhead photograph of a notebook, a brass magnifying glass, and a coil of copper wire on a wooden desk.

01Five things the shelf is saying

The index is a bibliography with three tracks: papers, projects, and agent skills. Read as a stack, the papers argue with each other in a fairly clear order.

01

A 7B code model scored 68.26% F1 on BigVul and 3.09% F1 on PrimeVul. In the strictest PrimeVul settings, GPT-3.5 and GPT-4 sat near chance. The paper is indexed as ICSE 2025.

02

Nine in ten machine-learning detection papers ask whether one function is vulnerable. Risse et al. argue that call context usually decides the answer, and that word counts alone can produce a high score.

03

More context moves the number. The CORRECT paper reports an example of 0.7 F1 and 0.8 precision on key CWEs once models see richer context, and also reports weak generalization and overthinking.

04

Fine-tuning can learn the surface of a patch. The 2026 Semantic Trap study finds strong scores when buggy code is paired with unrelated code, and failure when it is paired with the actual fix.

05

The 2026 list, plus six projects and three vendor skills, is about agents on a repository: scanning, filtering false positives, and checking their own findings.

02The measurement that reset the field

3.09%

F1 for a state-of-the-art 7B model on PrimeVul. The same model scored 68.26% F1 on BigVul. Ding et al., How Far Are We?, indexed as ICSE 2025. PrimeVul deduplicates, splits by time so later fixes do not leak into training, and scores the task more strictly.

Same 7B model. Bars are scaled to 70. The paper also reports that further training, and GPT-3.5 and GPT-4, stayed near random guessing in the most stringent setting.

What the abstracts actually report

These figures are the authors’ own, taken from the abstracts. This page has not re-run the experiments. Each card is one paper in the index.

ICSE 2025 · preprint March 2024PrimeVul

68.26 → 3.09 F1

Existing sets had weak labels, heavy duplication, and splits that leaked. A 7B model’s score collapsed on the cleaned set. Larger closed models did not rescue the strictest setting. Paper · Code

2024Top score on the wrong exam

9 in 10 papers

Their survey of the prior five years: almost every paper defines the task as function-level yes or no. In almost all cases the function does not contain the answer. High scores were available from word counts. Paper

2025 · 13 models, 2,000 pairs, 99 CWEsCORRECT

0.7 F1, 0.8 precision

Li et al. say three beliefs — models are unreliable, they ignore patches, scale has plateaued — came from tests with too little context. The 0.7 / 0.8 figures are their example on key CWEs, once context is supplied. Most false positives were reasoning errors. Scale still helped, with diminishing returns and a recall tradeoff. Generalization stayed limited, and models overthought. Paper

2025 · on the MegaVul patchesMono

16.7% undecidable

Their LLM review of security patches corrected 31.0% of labeling errors and recovered 89% of inter-procedural vulnerabilities. 16.7% of CVEs had a patch that does not show the root cause. Paper · Code

2026 · five fine-tuned LLMsThe semantic trap

High score, wrong pairing

Vanilla supervised fine-tuning looked strong when vulnerable code was paired with unrelated normal code, and produced high false-positive rates against the real patch. Chain-of-thought supervision reduced the symptoms and cost recall. The authors say the models still misread control flow and hallucinate API behavior. Paper

2023 · 16 LLMsVulBench, the optimistic year

LLMs beat deep learning

The early measurement, before PrimeVul. Sixteen LLMs against six deep-learning detectors and static analyzers. Several LLMs beat the deep-learning models. The abstract does not claim they beat the static analyzers. Paper · Code

Three symptoms of the semantic trap

From the 2026 abstract. The names are theirs.

Pairing-sensitive

The score depends on what the vulnerable function is compared with. Unrelated code is an easy exam. The real patch is a hard one.

Gap-dictated

The decision tracks the textual distance between the bug and the fix, which is a surface feature.

Fragile to meaning-preserving edits

Change the code without changing the behavior and the judgment moves.

A related 2024 paper, LLM4Vuln, tries to separate “can the model reason” from retrieval, extra context, and clever prompts. Its benchmark covers 147 vulnerable and 147 non-vulnerable cases in Solidity, Java, and C/C++, run as 3,528 scenarios across six models. The abstract’s conclusion is that those add-ons have uneven effects.

03How the shelf filled up

Counts are unique titles in the curated tables. The archive lists one paper twice, so the files contain 91 rows and 90 papers. 2026 runs through this snapshot, 5 October.

  1. 2018–2022 · 7 papers

    The shelf starts before chat models

    VulDeePecker at NDSS 2018, μVulDeePecker, Devign at NeurIPS 2019, a 2020 deep-learning survey, “Are We There Yet?” in 2022, VulBERTa, and a transformer paper at ACSAC. These are neural detectors on code. The index keeps them so the LLM papers have a lineage. Nothing in the index is from 2021.

  2. 2023 · 4 papers

    The first head-to-head looks hopeful

    VulBench reports that several LLMs beat traditional deep-learning detectors. DiverseVul, a vulnerable-source dataset, is published at RAID the same year. The archive is still thin.

  3. 2024 · 43 papers

    The flood, and the doubt

    Nearly half the unique titles sit in this one year. Prompting, fine-tuning, retrieval, and analyzers with a model inside them: IRIS, LLift at OOPSLA, LLMDFA at NeurIPS. The same year holds the skeptical cluster. PrimeVul’s preprint is March 2024, “Top Score on the Wrong Exam” is August, and “LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?)” is at IEEE S&P, with code under secllmholmes.

  4. 2025 · 26 papers

    Venues, context, and a 7B reasoner

    Ten of the 2025 papers name a venue: three at ICSE, and one each at NDSS, ACL, Usenix, SP, NAACL, TOSEM, and ACL Findings. CORRECT and Mono attack the evaluation setup from opposite sides, missing context and dirty patches. VulnLLM-R trains a 7-billion-parameter reasoning model and wraps it in an agent. The authors report that the agent beats CodeQL and AFL++ on real projects and that it found previously unreported vulnerabilities. They do not say how many. An ICSE user study, Closing the Gap, looks at detection and repair inside an IDE. A systematic review counts 263 studies from January 2020 through November 2025.

  5. 2026 · 10 curated papers

    Agents, triage, and a partial year

    The live README opens with VulnGym, a benchmark of coding agents at repository scale, and with an ISSTA paper on agents that filter false positives. Multi-agent harnesses, project-scale studies, and the Semantic Trap paper are already on the curated list. A separate file, rewritten by a bot on the morning of this snapshot, holds 30 more arXiv hits and shares none of their ids.

04What the index looks like as an object

Code is linked for 39 papers

Eighteen of the 36 papers on the 2025–26 list have a repository. Twenty-one archive entries do. The daily arXiv file links none: the script’s Papers With Code lookup is commented out, and every code cell is the string “None”.

Ring is 39/90 of a circle, 156 degrees.

Papers per year

20181
20192
20201
20210
20223
20234
202443
202526
202610

Bar length is relative to 2024. The 2026 bar is the curated list only, and the year is not over. The daily window is counted separately below.

Where the named venues sit

Fifty-three of the 90 papers leave the venue cell blank. The other 37 use the index’s own labels.

ICSE9
IEEE, unnamed5
NeurIPS2
EMNLP2
Usenix2
ISSTA2
ACL Findings2
NDSS2

One paper each at ACL, NAACL, TOSEM, ASE, FSE, OOPSLA, LCTES, RAID, ACSAC, IEEE S&P, and SP. The index writes those last two separately. “IEEE, unnamed” means the cell says only “IEEE.” Bars are scaled to ICSE’s 9.

A map for reading

This is a way to place a paper, built from the titles. It is a reading aid. The cells are examples.

Model at the center
Agent at the center
Function or file

Function or file · model

Classify the snippetFine-tunes, prompts, mixture-of-experts. This is the task Risse et al. call the wrong exam, and the one the Semantic Trap paper stress-tests.

Function or file · agent

Triage a findingSifting the Noise, at ISSTA 2026, compares agents that filter false positives. AgenticSCR reviews code for vulnerabilities the authors call immature.
Whole repository

Repository · model

Bring in the programLLMxCPG feeds a code property graph. IRIS, LLift, and LLMDFA put a model inside static analysis. Vul-RAG retrieves vulnerability knowledge. CORRECT’s claim lives here.

Repository · agent

Let it work the repoVulnGym benchmarks coding agents. VulnLLM-R adds an agent scaffold to a 7B reasoner. JitVul, at ACL 2025, benchmarks agents on real repositories. The tools below sit in this cell.

05Six ways into the list

A short path through each theme. Glosses stay inside the title, or inside an abstract cited above. The browser at the bottom has every paper.

From Large to Mammoth

2025 · NDSS · live list

A comparative evaluation at NDSS. The index gives no code link. Read the paper for the ranking. This page does not restate one.

LLMxCPG

2025 · Usenix · live list

Code property graphs guide the model. The index’s answer to “the function is not enough.”

IRIS

2024 · archive

LLM-assisted static analysis. The model is a component of an analyzer.

LLift

2024 · OOPSLA · archive

An LLM integrated into static analysis, aimed at practical bug detection.

LLMDFA

2024 · NeurIPS · archive

Dataflow analysis with a large language model. A sibling paper, LLMSAN, sanitizes model output with dataflow.

QRS

2026 · live list

A neuro-symbolic triad that synthesizes rules for autonomous discovery. The 2025 shelf has a sibling: neuro-symbolic static analysis with LLM-written vulnerability patterns.

Vul-RAG

2024 · archive

Retrieval of vulnerability knowledge, rather than retrieval of similar code alone.

SV-TrustEval-C

2025 · SP · live list

Tests whether models can do structural and semantic reasoning on source code, which is the skill the other papers assume.

VulnLLM-R

2025 · live list

A 7B reasoning model plus an agent scaffold. Authors report it beating CodeQL and AFL++ on real projects, and finding previously unreported vulnerabilities.

VulnGym

2026 · live list

Benchmarks coding agents on repository-level detection. Code is under Tencent/VulnGym.

Sifting the Noise

2026 · ISSTA · live list

A comparison of LLM agents whose job is false-positive filtering. One of the few 2026 papers with a named venue.

MulVul

2026 · live list

Retrieval-augmented multi-agent detection, with prompts evolved across models.

CVE-Bench

2025 · NAACL · live list

Repair, not detection. It tests whether a software-engineering agent can fix real CVEs. The index files it on the same list.

R2Vul

2025 · live list

Reinforcement learning, plus distillation of structured reasoning.

VulInstruct

2025 · live list

Teaches root-cause reasoning from security specifications.

VULPO

2025 · live list

On-policy optimization, with context in the loop.

GPTScan

2024 · ICSE · archive

GPT combined with program analysis, aimed at logic bugs in smart contracts.

MOS

2025 · live list

Mixture-of-experts tuning, applied to smart-contract detection.

PrimeVul

2025 · ICSE

Built to fix label noise, duplication, and temporal leakage. The strict set is where scores fall apart.

Mono

2025 · live list

On MegaVul: 31.0% of labeling errors corrected, 89% of inter-procedural cases recovered, 16.7% of CVEs left with an undecidable patch.

DiverseVul

2023 · RAID · archive

A vulnerable-source dataset from the year before the evaluation crisis.

CleanVul

2024 · archive

Function-level labels from commits, using LLM heuristics to clean them.

SecVulEval

2025 · live list

A benchmark aimed at real-world C and C++ rather than a single historical dataset.

VulEval

2024 · archive

Repository-level evaluation, a year before the agent benchmarks.

Two surveys sit across all six paths. The Esslingen systematic review covers 263 studies from January 2020 through November 2025 and keeps a living repository. LLMs in Software Security names the open problems as cross-language detection, multimodal inputs, and repository-scale analysis. A 2024 review of detection and repair is filed twice in the archive.

06Tools and skills, beside the papers

The contributing guide asks for two extra tables. A project has to use LLMs or agents for detection or discovery. A skill has to let an agent do security auditing. Descriptions below follow the index.

Projects

Knostic

OpenAnt

LLM-powered multi-stage vulnerability discovery with adversarial verification.

GitHub

lintsinghua

DeepAudit

Multi-agent red-team platform with Docker sandbox exploit validation.

GitHub

larlarua

AutoCVE

Automated vulnerability detection and reporting with a multi-agent architecture.

GitHub

ASCIT31 · GPL-3.0

Darkmoon

Autonomous pentest platform and MCP host. Per-technology offensive sub-agents, Active Directory and Kubernetes coverage, 80+ orchestrated tools, and an evidence trail per finding.

GitHub

usestrix

strix

Open-source autonomous AI penetration testing tool, installed with pip.

GitHub

Vercel

deepsec

A security harness for deep codebase vulnerability scanning with coding agents.

GitHub

Agent skills

OpenAI

codex-security

Autonomous repository-level vulnerability scanning via Codex agents. The index links a page whose address calls the release a research preview.

OpenAI

Anthropic

defending-code-reference-harness

Reference skills for threat modeling, scanning, triage, and patching with Claude Code.

GitHub

Cloudflare

security-audit-skill

A six-phase security audit skill with parallel hunting agents and adversarial validation.

GitHub

07The curated list and the daily firehose

Two different documents live in the repo. The README and the archive are edited by people. arxiv.md is rewritten by a GitHub Action. On this snapshot the two do not overlap.

01 · Query

Four phrase pairs

The abstract must contain “large language models” and one of: vulnerability detection, bug detection, defect prediction, bug prediction.

02 · Schedule

Midnight UTC

A workflow checks out the repo, installs the arxiv library, and runs scripts/daily_arxiv.py. It can also be started by hand.

03 · File

Overwrite, keep 30

arxiv.md is replaced, not appended. Papers are sorted by update date. The code column is hardcoded empty.

04 · People

Promote a row

2025 and later go in the README. 2024 and earlier go in docs/papers_archive.md. Newest year first.

The workflow notes that it borrows from LLM4SE and was refactored onto the arxiv.py library. The script comments that Papers With Code has shut its API, which is why code links are not filled in.

The copy cloned for this page is commit 2110eab, 5 October 2026, 02:42 UTC, message “Github Action Automatic Update Arxiv Papers.” The file header says “Updated on 2026.10.05.” None of its 30 arXiv ids appear in the README or the archive. The phrase query is strict enough to miss a curated paper that says “LLMs” instead of “large language models,” and loose enough to admit neighbors.

In the window, not on the curated list

Titles only. These abstracts were not read for this page. They show what the bot is seeing that a person has not promoted yet.

All 30, as of this morning

Twelve titles are marked “adjacent.” The rule used here: the title is about unit tests, fuzzing, clone detection, hardware bugs, performance bugs, or developer-workflow agents. The other 18 sit closer to detection, discovery, datasets, surveys, or attacks on detectors. That split is a reading of titles.

    08Browse the 90

    Every unique title in the README and the archive. The duplicated 2024 review is listed once. Tags were assigned from the wording of the title so this list can be filtered. A paper often belongs to more than one tag.

    09What this page does not settle

    The scores above are copied from ten abstracts. They are the authors’ reports. A different split, prompt, or CWE slice would move them, which is part of what those papers are arguing.

    The Esslingen review covers 263 studies through November 2025. This index is a curated 90, with a tail back to VulDeePecker in 2018 and a 2026 front the review’s abstract does not include. Use both.

    The archive is “2024 and earlier,” and it is not a pure LLM list. It keeps the deep-learning lineage on purpose. A few entries are adjacent tasks, including Android malware (LAMD, on the live list) and multi-agent test generation.

    One title is pasted twice in docs/papers_archive.md: “Large Language Model for Vulnerability Detection and Repair: Literature Review and the Road Ahead.” It is counted once here.

    Venue strings are copied as written. 2026 is a partial year. The daily window moves every midnight UTC, so the 30 titles above are this morning’s window.

    10Sources