LMs: Understanding Code Syntax and Semantics for Code Analysis (2305.12138v4)

Published 20 May 2023 in cs.SE and cs.AI

Abstract: LLMs~(LLMs) demonstrate significant potential to revolutionize software engineering (SE) by exhibiting outstanding performance in SE tasks such as code and document generation. However, the high reliability and risk control requirements in software engineering raise concerns about the lack of interpretability of LLMs. To address this concern, we conducted a study to evaluate the capabilities of LLMs and their limitations for code analysis in SE. We break down the abilities needed for artificial intelligence~(AI) models to address SE tasks related to code analysis into three categories: 1) syntax understanding, 2) static behavior understanding, and 3) dynamic behavior understanding. Our investigation focused on the ability of LLMs to comprehend code syntax and semantic structures, which include abstract syntax trees (AST), control flow graphs (CFG), and call graphs (CG). We employed four state-of-the-art foundational models, GPT4, GPT3.5, StarCoder and CodeLlama-13b-instruct. We assessed the performance of LLMs on cross-language tasks involving C, Java, Python, and Solidity. Our findings revealed that while LLMs have a talent for understanding code syntax, they struggle with comprehending code semantics, particularly dynamic semantics. We conclude that LLMs possess capabilities similar to an Abstract Syntax Tree (AST) parser, demonstrating initial competencies in static code analysis. Furthermore, our study highlights that LLMs are susceptible to hallucinations when interpreting code semantic structures and fabricating nonexistent facts. These results indicate the need to explore methods to verify the correctness of LLM output to ensure its dependability in SE. More importantly, our study provides an initial answer to why the codes generated by LLM are usually syntax-correct but vulnerable.

PDF HTML Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

References (67)

Authors (10)

Wei Ma (106 papers)
Shangqing Liu (28 papers)
Wenhan Wang (22 papers)
Qiang Hu (149 papers)
Ye Liu (153 papers)
Cen Zhang (69 papers)
Liming Nie (9 papers)
Yang Liu (2253 papers)
Zhihao Lin (16 papers)
Li Li (655 papers)

Citations (5)

View on Semantic Scholar

Tweets

https://twitter.com/ComputerPapers/status/1755542879257186324

https://twitter.com/ComputerPapers/status/1757728825910112458

LMs: Understanding Code Syntax and Semantics for Code Analysis (2305.12138v4)

Related Papers

Tweets