Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
110 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Visually Analyzing Contextualized Embeddings (2009.02554v1)

Published 5 Sep 2020 in cs.HC and cs.CL

Abstract: In this paper we introduce a method for visually analyzing contextualized embeddings produced by deep neural network-based LLMs. Our approach is inspired by linguistic probes for natural language processing, where tasks are designed to probe LLMs for linguistic structure, such as parts-of-speech and named entities. These approaches are largely confirmatory, however, only enabling a user to test for information known a priori. In this work, we eschew supervised probing tasks, and advocate for unsupervised probes, coupled with visual exploration techniques, to assess what is learned by LLMs. Specifically, we cluster contextualized embeddings produced from a large text corpus, and introduce a visualization design based on this clustering and textual structure - cluster co-occurrences, cluster spans, and cluster-word membership - to help elicit the functionality of, and relationship between, individual clusters. User feedback highlights the benefits of our design in discovering different types of linguistic structures.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (1)
  1. Matthew Berger (22 papers)
Citations (12)

Summary

We haven't generated a summary for this paper yet.