07The curated list and the daily firehose
Two different documents live in the repo. The README and the archive are edited by people. arxiv.md is rewritten by a GitHub Action. On this snapshot the two do not overlap.
01 · Query
Four phrase pairs
The abstract must contain “large language models” and one of: vulnerability detection, bug detection, defect prediction, bug prediction.
02 · Schedule
Midnight UTC
A workflow checks out the repo, installs the arxiv library, and runs scripts/daily_arxiv.py. It can also be started by hand.
03 · File
Overwrite, keep 30
arxiv.md is replaced, not appended. Papers are sorted by update date. The code column is hardcoded empty.
04 · People
Promote a row
2025 and later go in the README. 2024 and earlier go in docs/papers_archive.md. Newest year first.
The workflow notes that it borrows from LLM4SE and was refactored onto the arxiv.py library. The script comments that Papers With Code has shut its API, which is why code links are not filled in.
The copy cloned for this page is commit 2110eab, 5 October 2026, 02:42 UTC, message “Github Action Automatic Update Arxiv Papers.” The file header says “Updated on 2026.10.05.” None of its 30 arXiv ids appear in the README or the archive. The phrase query is strict enough to miss a curated paper that says “LLMs” instead of “large language models,” and loose enough to admit neighbors.
In the window, not on the curated list
Titles only. These abstracts were not read for this page. They show what the bot is seeing that a person has not promoted yet.
Updated 2026-08-04. Agents plus context, continuing the 2025 argument.
Updated 2026-07-15. Splits reasoning and exploration, at repository scale.
Updated 2026-07-27. A new theme for the shelf: attacking the detector.
Updated 2026-09-29. A dissection of retrieval, which Vul-RAG and others rely on.
Updated 2026-09-10. A dataset of vulnerabilities in code that a model wrote.
Updated 2026-07-04. Continues the neuro-symbolic thread of QRS and the 2025 static-analysis paper.
All 30, as of this morning
Twelve titles are marked “adjacent.” The rule used here: the title is about unit tests, fuzzing, clone detection, hardware bugs, performance bugs, or developer-workflow agents. The other 18 sit closer to detection, discovery, datasets, surveys, or attacks on detectors. That split is a reading of titles.