Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Anomaly Detection of Tabular Data Using LLMs (2406.16308v1)

Published 24 Jun 2024 in cs.LG, cs.AI, and cs.CL

Abstract: LLMs have shown their potential in long-context understanding and mathematical reasoning. In this paper, we study the problem of using LLMs to detect tabular anomalies and show that pre-trained LLMs are zero-shot batch-level anomaly detectors. That is, without extra distribution-specific model fitting, they can discover hidden outliers in a batch of data, demonstrating their ability to identify low-density data regions. For LLMs that are not well aligned with anomaly detection and frequently output factual errors, we apply simple yet effective data-generating processes to simulate synthetic batch-level anomaly detection datasets and propose an end-to-end fine-tuning strategy to bring out the potential of LLMs in detecting real anomalies. Experiments on a large anomaly detection benchmark (ODDS) showcase i) GPT-4 has on-par performance with the state-of-the-art transductive learning-based anomaly detection methods and ii) the efficacy of our synthetic dataset and fine-tuning strategy in aligning LLMs to this task.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Aodong Li (10 papers)
  2. Yunhan Zhao (13 papers)
  3. Chen Qiu (43 papers)
  4. Marius Kloft (65 papers)
  5. Padhraic Smyth (52 papers)
  6. Maja Rudolph (25 papers)
  7. Stephan Mandt (100 papers)
Citations (2)

Summary

We haven't generated a summary for this paper yet.