Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language Processing (2310.12664v1)
Abstract: The emergence of LLMs, such as ChatGPT, has revolutionized general natural language preprocessing (NLP) tasks. However, their expertise in the financial domain lacks a comprehensive evaluation. To assess the ability of LLMs to solve financial NLP tasks, we present FinLMEval, a framework for Financial LLM Evaluation, comprising nine datasets designed to evaluate the performance of LLMs. This study compares the performance of encoder-only LLMs and the decoder-only LLMs. Our findings reveal that while some decoder-only LLMs demonstrate notable performance across most financial tasks via zero-shot prompting, they generally lag behind the fine-tuned expert models, especially when dealing with proprietary datasets. We hope this study provides foundation evaluations for continuing efforts to build more advanced LLMs in the financial domain.