Perplexed: Understanding When Large Language Models are Confused (2404.06634v1)

Published 9 Apr 2024 in cs.SE

Abstract: LLMs have become dominant in the NLP field causing a huge surge in progress in a short amount of time. However, their limitations are still a mystery and have primarily been explored through tailored datasets to analyze a specific human-level skill such as negation, name resolution, etc. In this paper, we introduce perplexed, a library for exploring where a particular LLM is perplexed. To show the flexibility and types of insights that can be gained by perplexed, we conducted a case study focused on LLMs for code generation using an additional tool we built to help with the analysis of code models called codetokenizer. Specifically, we explore success and failure cases at the token level of code LLMs under different scenarios pertaining to the type of coding structure the model is predicting, e.g., a variable name or operator, and how predicting of internal verses external method invocations impact performance. From this analysis, we found that our studied code LLMs had their worst performance on coding structures where the code was not syntactically correct. Additionally, we found the models to generally perform worse at predicting internal method invocations than external ones. We have open sourced both of these tools to allow the research community to better understand LLMs in general and LLMs for code generation.

PDF HTML Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

References (49)

Authors (2)

Nathan Cooper (35 papers)
Torsten Scholak (14 papers)

Tweets

https://twitter.com/ncooper57/status/1778859647325123008

https://twitter.com/ComputerPapers/status/1778335814960779694

Perplexed: Understanding When Large Language Models are Confused (2404.06634v1)

Related Papers

Tweets