In progress · University of Moratuwa
Understanding Security Vulnerabilities in LLM-Generated Web Application Code
A Red Teaming Approach to Characterising the Functional-Security Correctness Gap
The area. Large Language Models such as GPT-4o, Claude and DeepSeek-Coder are
now widely used to generate web application code. My research sits where AI meets application
security: it studies how safe that generated code really is once it becomes a complete,
deployable web application rather than an isolated snippet.
The gap I'm addressing. Code an LLM produces can pass every functional test and
still be insecure. Studies report that around 68.8% of generated snippets break
security rules, while the share of real-world output that is both functional and secure stays
below 6%. I call this divergence between "it works" and "it is safe" the
functional-security correctness gap, and my work characterises it through adaptive red
teaming so that AI-assisted development can be guided by evidence instead of assumption.