Formal logical reasoning capability of transformer-based LLMs in practice
Establish whether transformer-based large language models can perform formal logical reasoning in practice to solve tasks that require such reasoning, rather than relying primarily on probabilistic pattern-matching of training data.
References
While these works provide insights into the theoretical computational complexity of transformers, in practice, it remains unclear whether these LLMs can perform formal logical reasoning to solve tasks.
Across open-source models, following a specified negation semantics remains unsolved: the strongest scores $59$--$74\%$ across the four semantic viewpoints and the weakest $31$--$67\%$; all are order-sensitive on more than half of logically identical rule shufflings, while the two weaker models frequently overcommit on well-founded ``undefined.''