Program Graph Learning for Software Vulnerability Analysis: A Survey

Published in Transactions on Graph Intelligence and Network Applications (TGINA), Scilight, 2026, 2026

Abstract

Software vulnerabilities represent an enduring threat to modern cyberspace. Effective vulnerability detection increasingly relies on reasoning about complex program semantics, structural dependencies, and execution behaviours. In recent years, advances in graph representation learning and large language models have reframed vulnerability analysis as a graph learning problem over program entities. Building on and extending traditional methods — static analysis, dynamic analysis, symbolic execution, and fuzzing — this perspective enables more effective vulnerability detection, fine-grained localization, multi-class classification, and repair. The survey reviews recent progress in graph-driven vulnerability intelligence, structured around these tasks, together with representative graph-based, graph-based multimodal, LLM-assisted and repair methods, and the ecosystem of datasets, program graph construction tools, and computing frameworks that supports them.

Key Contributions

  • A unified graph-learning formalisation of the four core tasks: detection and classification as graph-level prediction over typed program graphs, localization as line-level prediction, and repair as graph-conditioned patch generation.
  • A structured review of four method families: graph-based (Devign, ReVeal, ReGVD and successors), graph-based multimodal (fusing graphs with text, slices, images and LLM-guided contexts), LLM-assisted (RAG, multi-agent, fine-tuning, tool-assisted) and vulnerability repair methods.
  • A practical map of the supporting ecosystem: synthetic and real-world datasets (SARD/Juliet, BigVul, CVEFixes, ReposVul), program graph construction tools (Joern, CodeQL, Tree-sitter, LLVM, WALA, Slither), and graph/Transformer/LLM/RAG computing frameworks.
  • Five persistent challenges: interpretability beyond correlation, cross-language generalization, data quality and label noise, benchmark standardization, and the new security risks of AI-generated code and agent toolchains (prompt contamination, unsafe tool invocation, MCP-style configuration attack surfaces).

Resources

Recommended citation: Yu, T.; Wang, J.; Li, M.; Hu, Y.; Hu, J.; Cai, X.; Hu, W.; Borgi, A. (2026). "Program Graph Learning for Software Vulnerability Analysis: A Survey." Transactions on Graph Intelligence and Network Applications. https://doi.org/10.53941/tgina.2026.100005.
Download Paper