Papers
Topics
Authors
Recent
Search
2000 character limit reached

CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text

Published 4 Mar 2024 in cs.AI and cs.PL | (2403.01784v1)

Abstract: LLMs such as ChatGPT are increasingly proficient in understanding and generating a mixture of code and text. Evaluation based on such $\textit{mixture}$ can lead to a more comprehensive understanding of the models' abilities in solving coding problems. However, in this context, current evaluation methods are either limited in task coverage or lack standardization. To address this issue, we propose using category theory as a framework for evaluation. Specifically, morphisms within a code category can represent code debugging and transformation, functors between two categories represent code translation, and functors between a code category and a natural language category represent code generation, explanation, and reproduction. We present an automatic evaluation framework called $\textbf{CatCode}$ ($\textbf{Cat}$egory $\textbf{Code}$) that can comprehensively assess the coding abilities of LLMs, including ChatGPT, Text-Davinci, and CodeGeeX.

Authors (3)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (25)
  1. Unified pre-training for program understanding and generation. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tür, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 2655–2668. Association for Computational Linguistics, 2021.
  2. Category theory for programming. CoRR, abs/2209.01259, 2022.
  3. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. arXiv preprint arXiv:1905.13319, 2019.
  4. Multi-lingual evaluation of code generation models. CoRR, abs/2210.14868, 2022.
  5. Tai-Danae Bradley. What is applied category theory? arXiv preprint arXiv:1809.05923, 2018.
  6. MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation. IEEE Transactions on Software Engineering, pages 1–17, 2023.
  7. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021.
  8. Codebert: A pre-trained model for programming and natural languages. arXiv preprint arXiv:2002.08155, 2020.
  9. Seven sketches in compositionality: An invitation to applied category theory. arXiv preprint arXiv:1803.05316, 2018.
  10. Graphcodebert: Pre-training code representations with data flow. arXiv preprint arXiv:2009.08366, 2020.
  11. Semantic robustness of models of source code. In IEEE International Conference on Software Analysis, Evolution and Reengineering, SANER 2022, Honolulu, HI, USA, March 15-18, 2022, pages 526–537. IEEE, 2022.
  12. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022.
  13. Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv preprint arXiv:2102.04664, 2021.
  14. Generating diverse code explanations using the GPT-3 large language model. In Jan Vahrenhold, Kathi Fisler, Matthias Hauswirth, and Diana Franklin, editors, ICER 2022: ACM Conference on International Computing Education Research, Lugano and Virtual Event Switzerland, August 7 - 11, 2022, Volume 2, pages 37–39. ACM, 2022.
  15. Training language models to follow instructions with human feedback. In NeurIPS, 2022.
  16. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA, pages 311–318. ACL, 2002.
  17. On the generalizability of neural program models with respect to semantic-preserving program transformations. Inf. Softw. Technol., 135:106552, 2021.
  18. Testing neural programs. CoRR, abs/1908.10711, 2019.
  19. Codebleu: a method for automatic evaluation of code synthesis. CoRR, abs/2009.10297, 2020.
  20. David I. Spivak. Category Theory for the Sciences. MIT Press, 2014.
  21. Intellicode compose: code generation using transformer. In Prem Devanbu, Myra B. Cohen, and Thomas Zimmermann, editors, ESEC/FSE ’20: 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Virtual Event, USA, November 8-13, 2020, pages 1433–1443. ACM, 2020.
  22. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 8696–8708. Association for Computational Linguistics, 2021.
  23. Adversarial examples for models of code. Proc. ACM Program. Lang., 4(OOPSLA):162:1–162:30, 2020.
  24. Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x, 2023.
  25. Multilingual code snippets training for program translation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11783–11790, 2022.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.