Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
110 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

FOLIO: Natural Language Reasoning with First-Order Logic (2209.00840v3)

Published 2 Sep 2022 in cs.CL

Abstract: LLMs have achieved remarkable performance on a variety of natural language understanding tasks. However, existing benchmarks are inadequate in measuring the complex logical reasoning capabilities of a model. We present FOLIO, a human-annotated, logically complex and diverse dataset for reasoning in natural language (NL), equipped with first-order logic (FOL) annotations. FOLIO consists of 1,430 examples (unique conclusions), each paired with one of 487 sets of premises used to deductively reason for the validity of each conclusion. The logical correctness of the premises and conclusions is ensured by their FOL annotations, which are automatically verified by an FOL inference engine. In addition to the main NL reasoning task, NL-FOL pairs in FOLIO constitute a new NL-FOL translation dataset. Our experiments on FOLIO systematically evaluate the FOL reasoning ability of supervised fine-tuning on medium-sized LLMs. For both NL reasoning and NL-FOL translation, we benchmark multiple state-of-the-art LLMs. Our results show that a subset of FOLIO presents a challenge for one of the most capable {LLM} publicly available, GPT-4.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (35)
  1. Simeng Han (20 papers)
  2. Hailey Schoelkopf (22 papers)
  3. Yilun Zhao (59 papers)
  4. Zhenting Qi (19 papers)
  5. Martin Riddell (4 papers)
  6. Luke Benson (2 papers)
  7. Lucy Sun (1 paper)
  8. Ekaterina Zubova (1 paper)
  9. Yujie Qiao (4 papers)
  10. Matthew Burtell (2 papers)
  11. David Peng (3 papers)
  12. Jonathan Fan (3 papers)
  13. Yixin Liu (108 papers)
  14. Brian Wong (5 papers)
  15. Malcolm Sailor (1 paper)
  16. Ansong Ni (17 papers)
  17. Linyong Nan (17 papers)
  18. Jungo Kasai (38 papers)
  19. Tao Yu (282 papers)
  20. Rui Zhang (1138 papers)
Citations (69)

Summary

We haven't generated a summary for this paper yet.

X Twitter Logo Streamline Icon: https://streamlinehq.com

Tweets