Program Graph Learning for Software Vulnerability Analysis: A Survey
Published:
Authors: Tong Yu† (Wuhan University), Junjie Wang† (Wuhan University), Ming Li (Zhejiang Normal University), Yuhang Hu, Junwei Hu, Xiantao Cai (Wuhan University), Wenbin Hu* (Wuhan University), Alessio Borgi (University of Cambridge & Sapienza University of Rome) — †equal contribution, *corresponding author
Journal: Transactions on Graph Intelligence and Network Applications (TGINA), published 9 September 2026 · open-access PDF

In plain words: a program is not a string — it is a web of statements that pass values to each other, call each other and gate each other with conditions. Draw that web as a graph and a security bug becomes a pattern in the graph: a tainted value flowing to a dangerous function, a check that should sit on a path and does not. This survey is a field guide to everything built on that observation, from the first graph neural networks over code to today’s LLM agents arguing about whether a function is exploitable.
Why code wants to be a graph
The stakes are not abstract. The survey opens with the numbers: the global average cost of a data breach reached USD 4.88 million in 2024, and Cybersecurity Ventures projected global cybercrime damages of USD 10.5 trillion annually by 2025. One analysis of 348 open-source smart contracts put their maximum total value locked above USD 11.2 billion — code whose bugs are directly convertible into money. And because modern software is assembled from open-source components, a single vulnerability propagates through dependency networks, multiplying both attack surface and remediation cost.
The classical arsenal — static analysis, dynamic analysis, symbolic execution, fuzzing — has well-known failure modes: high false-positive rates, incomplete path coverage, weak modelling of deep program semantics, and limited reasoning across procedure boundaries. The bugs that slip through are precisely the ones that hinge on semantic conditions: a value that is attacker-controllable three calls upstream, a sanitisation check present on one path and absent on its twin.
This is where the graph view earns its place. A token sequence flattens exactly the structure that matters; a program graph keeps it. Control-flow edges reveal execution constraints, data-flow edges reveal value propagation, call edges stitch functions together — the graph mirrors how the program actually behaves, not how it reads.

Four tasks, one formalism
The survey’s first contribution is organisational: every task in the field is a mapping out of the same typed graph. A program entity \(P\) — a function, file or project — becomes
with \(V\) the program elements, \(X\) their features, and one edge set \(E_r\) per relation type \(r\) — syntactic, control-flow, data-flow, call, control- and data-dependence. If you have met relational message passing on heterogeneous graphs, this is the same machinery pointed at code: one adjacency matrix \(A_r\) per relation, propagation along all of them.
The four tasks then line up as three mappings of increasing ambition:
Detection is the binary special case of classification (\(c=0\) benign, otherwise a CWE category). Localisation predicts which of the \(n\) physical source lines are implicated. Repair is the end of the ladder: generate edit operations \(\Delta\) — insert the missing check, replace the unsafe call, sanitise the tainted flow — whose application removes the vulnerability while preserving what the program was for. Read graph-theoretically, a repair is an edit that breaks the vulnerability-triggering dependence chain.
Four families of methods
The survey reviews the methods in four families, and the arc across them is the field’s recent history in miniature.
| Family | Representative methods | What it adds | What it costs |
|---|---|---|---|
| Graph-based | Devign, ReVeal, ReGVD, IVDetect | Explicit structure; GNNs over AST/CFG/DFG/CPG | Graph construction time; mostly intra-procedural |
| Multimodal | MGVD, GRACE, LLMxCPG, Vul-LMGNNs | Fuses graphs with text, slices, images, LLM context | Alignment overhead; redundant or noisy modalities |
| LLM-assisted | Vul-RAG, ProveRAG, VulTrial, Agent4Vul | External security knowledge; multi-step reasoning | Sensitive to retrieval quality; hallucination |
| Repair | SimFix, TBar, VulRepair, ChatRepair, InferFix | From flagging bugs to generating validated patches | Depends on templates, donors, tests, analysis quality |
Graph-based methods established the substrate. Devign fused AST, CFG, DFG and the token sequence into one composite graph; ReVeal moved to code property graphs and, importantly, to realistic imbalanced data; ReGVD went the other way, building cheap token co-occurrence graphs to dodge the construction cost of full program analysis. The trade running through the whole family: richer graphs carry richer semantics and cost more to build — a real constraint at repository scale.
Multimodal methods accept that no single view suffices. MGVD’s three-channel fusion of AST, PDG and text lifts F1 by roughly 10 points over single-view baselines. Two results stand out for where the field is heading. GRACE injects graph structure into LLM prompts alongside retrieved demonstrations and reports F1 gains of at least 28.65% — structure helping a language model, not competing with it. LLMxCPG uses the CPG to slice the program first, cutting the code the LLM must read by 67.84–90.93% while improving F1 by 15–40%: the graph as a relevance filter for an expensive reasoner.
LLM-assisted methods make the language model the reasoning engine and vary what evidence reaches it. Retrieval-augmented approaches (Vul-RAG, ProveRAG, VulInstruct) feed it vulnerability knowledge, historical cases and CWE descriptions — ProveRAG adding provenance tracking and self-critique to keep conclusions supported. The multi-agent line is the most entertaining to read and addresses a real failure mode, one-sided judgement: VulTrial literally stages a courtroom, with agents arguing for and against a vulnerability finding before a verdict; Agent4Vul plugs GNN-derived structural features into the agent loop.
Repair methods close the loop. The lineage runs from donor-code and template approaches (SimFix, TBar) through mined fix patterns, to neural generators — VulRepair’s CodeT5-style model reaches 44% perfect predictions on 8,482 real-world vulnerability fixes, repairing 745 of 1,706 real vulnerabilities — and on to LLM-era systems that fold in feedback: ChatRepair converses with failing tests, InferFix grounds prompts in static-analysis reports, VulDebugger lets an agent compare expected against actual runtime state. The direction of travel is consistent: away from unconstrained generation, towards evidence-grounded, validation-aware patching.
The toolbox: datasets, parsers, frameworks
A survey is also a shopping list, and Section 4 of the paper is the field’s most complete one.
Datasets split along a line that matters more than size: synthetic versus real. NIST’s SARD and the derived Juliet suite inject over a hundred CWE types into clean code — broad coverage, but easier than reality. Real-world sets mined from Git history (BigVul, CVEFixes, ReposVul) supply scale and patch context, with a caveat the survey is blunt about: they label code by fix-commit diffs, and not everything a fix touches is the vulnerability, nor is everything it leaves alone benign. That caveat returns with force in the challenges below.
Graph construction tools set the ceiling on everything downstream — a model cannot reason over structure the parser never extracted. Joern is the standard integrated CPG builder (AST + CFG + PDG + DFG across ten-plus languages); CodeQL turns program graphs into a queryable relational model, beloved of variant analysis; Tree-sitter and ANTLR handle fast AST construction; LLVM and Frama-C serve C/C++ with high-precision flow graphs; WALA and Soot cover the Java ecosystem; Slither is purpose-built for Solidity contracts.
Frameworks round out the pipeline: graph learning libraries, Transformer stacks, LLM inference engines and RAG infrastructure, composed into end-to-end vulnerability-analysis systems.
Five open problems — and the new one
The survey closes with five challenges, and they interlock: fixing one in isolation tends to founder on the others.
- Interpretability is correlational, not causal. Explainers highlight where the model attends — salient nodes, attributed subgraphs — but an exploitable bug is a causal story: attacker-controlled data reaching a sink because a check is missing on a path. The survey argues for program-semantics-level explanations: minimal vulnerability-evidence subgraphs, cross-validated against static analysis, symbolic execution and taint tracking. From “where the model looks” to “why the bug exists”.
- Cross-language generalisation collapses. Models trained on one language learn its surface statistics — syntax shapes, API idioms — rather than language-agnostic vulnerability semantics, and fall over on unseen languages. Unified graph schemas and multilingual pretraining are the proposed way out.
- Label noise is structural. The fix-commit labelling shortcut teaches models patch patterns and project style instead of vulnerability semantics. The remedy is finer-grained supervision: vulnerable lines, critical variables, source-to-sink paths, verifiable test cases.
- No standard benchmark. Inconsistent subsets, splits and filters make results incomparable, and evaluation fixates on function-level binary accuracy while localisation quality, explanation faithfulness and repair effectiveness go unmeasured.
- AI-generated code and agents change the threat model. This is the challenge that reads like a dispatch from this year rather than a literature review. When agents read repositories, invoke tools, install dependencies and submit patches, vulnerabilities stop being purely human coding mistakes: generation bias, prompt contamination, unsafe tool invocation, dependency pollution, configuration hijacking. Prompt templates, context files, tool declarations and MCP configurations become supply-chain attack surfaces.
Where this sits in the cybersecurity story
This post opens the Cybersecurity book, and a survey is the right doorway: one paper that shows the whole field through a single lens. The machinery comes from the GNN book — typed edges, relational message passing, graph-level readouts, the heterogeneous-graph toolkit — but the security domain supplies what benchmarks rarely do: relations with fixed, auditable semantics. A data-flow edge is not a learned similarity; it is a fact about the program, checkable by a compiler. That is why the structure survives the LLM era here more visibly than in most application domains: when your evidence must ultimately convince a security engineer, edges that mean something beat attention weights that merely light up.
It is also, candidly, a review — the contribution is the map, not a new model. But the map includes the first systematic framing I know of for graph-driven agentic software security, and that section alone is worth the read for anyone building with coding agents today.
✅ Key Takeaways
- Vulnerability analysis is now a graph learning problem over typed program graphs \(G_P=(V,\{E_r\},X,R)\), with four tasks — detection, classification, localisation, repair — as mappings of increasing ambition out of the same object.
- The graph representations are complementary, not competing: AST for syntax, CFG for execution order, DFG for value flow, PDG for dependence, CPG as the multi-relational overlay that became the field's standard substrate.
- Across four method families the role of the graph has shifted: from encoder substrate (Devign, ReVeal) to discipline for LLMs — slicing their input (LLMxCPG: up to 90% less code, higher F1), structuring their prompts (GRACE: ≥28.65% F1 gain), and grounding their claims.
- Repair closes the loop, from fix templates to neural generators (VulRepair: 44% perfect patches) to validation-aware LLM agents — the trend is evidence-grounded patching, not free generation.
- The ecosystem is mature enough to build on: SARD/Juliet and BigVul/CVEFixes/ReposVul for data (mind the fix-commit label noise), Joern and CodeQL for graph construction.
- Five interlocking open problems, the newest being agentic security: prompts, tool declarations and MCP configurations as supply-chain attack surfaces — with agent-behaviour graphs proposed as the analysis substrate.
References
- Yu, T.; Wang, J.; Li, M.; Hu, Y.; Hu, J.; Cai, X.; Hu, W.; Borgi, A. (2026). Program Graph Learning for Software Vulnerability Analysis: A Survey. Transactions on Graph Intelligence and Network Applications.
- Zhou, Y., Liu, S., Siow, J., Du, X., & Liu, Y. (2019). Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks. NeurIPS 2019.
- Chakraborty, S., Krishna, R., Ding, Y., & Ray, B. (2021). Deep Learning Based Vulnerability Detection: Are We There Yet? (ReVeal). IEEE TSE.
- Nguyen, V.-A., Nguyen, D. Q., Nguyen, V., Le, T., Tran, Q. H., & Phung, D. (2022). ReGVD: Revisiting Graph Neural Networks for Vulnerability Detection. ICSE 2022 Companion.
- Fan, J., Li, Y., Wang, S., & Nguyen, T. N. (2020). A C/C++ Code Vulnerability Dataset with Code Changes and CVE Summaries (BigVul). MSR 2020.
- Yamaguchi, F., Golde, N., Arp, D., & Rieck, K. (2014). Modeling and Discovering Vulnerabilities with Code Property Graphs (Joern). IEEE S&P 2014.
- Fu, M., Tantithamthavorn, C., Le, T., Nguyen, V., & Phung, D. (2022). VulRepair: A T5-Based Automated Software Vulnerability Repair. ESEC/FSE 2022.
- Xia, C. S., & Zhang, L. (2024). Automated Program Repair via Conversation (ChatRepair). ISSTA 2024.
