---
title: Bluesky Dataset Overview
url: https://www.emergentmind.com/topics/bluesky-dataset
type: topic
---

# Bluesky Dataset Overview

The Bluesky Dataset refers collectively to a set of publicly accessible, large-scale, and multi-faceted social network data resources generated from the decentralized microblogging platform Bluesky. Enabled by Bluesky’s open AT Protocol and public data policies, these datasets support computational social science, network analysis, and machine learning, offering structured, longitudinal, and granular representations of user behavior, social ties, content creation, and higher-order interactions in a decentralized context. The most significant, high-coverage Bluesky datasets include BlueTempNet, “A Blue Start”, Bluesky scholarly communication corpora, cross-platform comparative datasets, and population-scale interaction or content archives.

## 1. Dataset Architectures and Data Models

Bluesky datasets employ both standard graph and advanced multi-network/hypergraph representations, leveraging the platform’s feature set:

- **Pairwise directed networks**: The canonical representation is $G = (V, F)$, where $V$ is the set of users (nodes) and $F \subseteq V \times V$ is the set of directed “follow” relations. Edge attributes (e.g., timestamps for temporal networks, sign for follow/block, etc.) are integrated as needed [2505.11608, 2407.17451].
- **Higher-order interactions**: Bluesky enables group-mediated relations through “starter packs,” which are user-curated collections of accounts (and sometimes feeds). These are formalized as a hypergraph $H = (V, E)$, with $V$ as users/feeds and $E$ a set of hyperedges, each hyperedge corresponding to a starter pack [2505.11608].
- **Temporal multi-networks**: Datasets like BlueTempNet provide a temporal, multilayer graph,
  $$
  G = (V, E_1, E_2, E_3)
  $$
  where $E_1$ (follows), $E_2$ (blocks), and $E_3$ (user–feed actions) are time-stamped edge sets, analyzed jointly or per layer [2407.17451].
- **Content graphs and event logs**: Large-scale text and interaction corpora comprise per-post/message records structured by event type (post, reply, repost, like), including rich metadata, engagement counts, and sometimes full-thread context [2404.18984, 2501.12886, 2511.17241].
- **Structural and attribute-rich schemas**: Datasets standardize fields such as anonymized user IDs, content text, timestamps (ranging from minute- to millisecond-level granularity), metadata on feeds/communities, and derived features (e.g., sentiment scores, stance labels) [2404.18984, 2506.03443, 2501.11605, 2506.07606].

## 2. Data Acquisition, Anonymization, and Metadata

Bluesky datasets are built by directly querying official APIs or streaming Firehose endpoints:

- **Node, edge, and message extraction**: Utilizing endpoints such as `app.bsky.graph.getFollows`, `app.bsky.graph.getActorStarterPacks`, and app.bsky.feed.post/repost APIs, data collectors reconstruct network structure, group memberships, and complete content histories [2505.11608, 2402.03239, 2404.18984].
- **Temporal precision**: Millisecond-level precision is routinely available for edges (follows, blocks, join events), supporting inter-event time analysis and high-fidelity dynamic studies [2407.17451].
- **Anonymization**: Public dumps map all user and content IDs to irreversibly anonymized or randomized integers or UUIDs, and PII is removed in accordance with platform policy [2501.12886, 2404.18984, 2505.11608].
- **Comprehensive metadata**: Rich auxiliary files detail user activity, account type, number of followers, posts, starter-pack affiliations, feed creation or joining actions, language, and (when relevant) OpenAlex or NewsGuard article-level metadata [2507.18840, 2501.10557, 2407.17451].

| Dataset                | Users (n)       | Edges/Interactions (m)       | Groupings          | Content Volume    | Distinctive Features             |
|------------------------|-----------------|------------------------------|--------------------|-------------------|-----------------------------------|
| BlueTempNet [2407.17451] | 147k members    | 5.7M follows, 0.5M blocks    | 39,968 feeds       | N/A               | Temporal, multi-network, ms-level |
| A Blue Start [2505.11608]| 26.7M users     | 1.6B follows                 | 301k starter packs | N/A               | Hypergraphs for higher-order ties |
| "I'm in the Bluesky…" [2404.18984] | 4.1M users | 145M follows, 23M replies, 63M reposts | 11 feeds           | 237M posts        | Population-level, complete posts  |
| PolitiSky24 [2506.07606] | 8.5k users     | 16k (user-target stance pairs) | NA                | 18M posts         | User-level stance w/ rationale    |

## 3. Temporal, Structural, and Higher-Order Data Properties

The Bluesky datasets collectively cover a wide range of user interaction scenarios and address both pairwise and community/group-level network science:

- **Degree distributions**: All major interaction networks display heavy-tailed degree distributions, often characterized by fitted power-law exponents (e.g., $\alpha_{\text{out}} \approx 1.44$ for follows) [2405.17571].
- **Clustering and connected components**: High (normalized) clustering coefficients (10 to 200 times configuration-model baseline) are observed in follows, replies, and repost layers; giant strongly-connected components dominate, and smaller SCCs exist in both user and starter-pack hypergraphs [2505.11608, 2404.18984].
- **Higher-order metrics**: Starter-pack hypergraphs permit s-line-graph densities, hypergraph k-cores, user-pair co-occurrence rates across lists, and community entropy quantification [2505.11608].
- **Temporal coverage**: Datasets range from Feb 2023 (platform launch) to May 2025, with rolling updates and sub-daily event precision in some corpora, supporting fine-grained longitudinal research [2407.17451, 2404.18984].

## 4. Specialized Content, Behavior, and Population Subsets

Several datasets capture distinctive behavioral, topical, or population subsets within Bluesky:

- **Scholarly dissemination**: The scholarly communication dataset identifies 87,470 posts referencing DOIs, joined with OpenAlex for bibliometric content/discipline analysis, with derived originality scores based on post/title cosine similarity [2507.18840].
- **Political collections and stance**: PolitiSky24 and “Politics and polarization on Bluesky” provide both post-level and user-level stance detection for U.S. politics, including millions of stance-annotated posts, interaction graphs, and full pseudonymized posting histories [2506.07606, 2506.03443].
- **Cross-platform records**: MADOC integrates standardized Bluesky subsets with Reddit, Koo, Voat across 12 communities, designed for toxic behavior and moderation research [2501.12886].
- **Persona/thread-level behavior prediction**: The SocialSim challenge dataset includes 6.4M Bluesky conversation threads, 12 action classes (with rare-action focus), and 25 persona clusters [2511.17241].
- **News reliability**: The MurkySky dataset resolves news-sharing posts to NewsGuard-rated domains, labeling reliability and extracting hashtag/topic/audience segmentation networks, with observed unreliable content prevalence ≈2% [2501.10557].

## 5. Data Access, File Formats, and Tooling

Consistent with Bluesky’s open-data orientation, most core datasets are distributed under academic or open licenses (e.g., CC-BY 4.0) [2404.18984, 2505.11608, 2407.17451]. Key distribution and access conventions:

- **File formats**: CSV for edgelists/metadata; JSON and JSONL for full post/event records; GEXF for graph multilayers; Parquet for massive tabular content [2404.18984, 2407.17451, 2501.12886].
- **Dataset registries**: Zenodo (e.g., [https://doi.org/10.5281/zenodo.11082878](https://doi.org/10.5281/zenodo.11082878)), SOMAR ([https://socialmediaarchive.org/record/78](https://socialmediaarchive.org/record/78)), IEEE DataPort ([https://ieee-dataport.org/documents/bluetempnet-temporal-multi-network-dataset-social-interactions-bluesky-social](https://ieee-dataport.org/documents/bluetempnet-temporal-multi-network-dataset-social-interactions-bluesky-social)).
- **Code libraries**: Datasets such as MADOC provide Python (pyMADOC) and R (rMADOC) clients. Scripts for data retrieval and cleaning are routinely released alongside data [2501.12886, 2407.17451].
- **API coverage**: Most data can be reconstructed/updated from the public Bluesky AT Protocol APIs (PDS xRPC, AppView, Firehose), with no private key required [2402.03239, 2407.17451, 2404.18984].

## 6. Research Applications and Limitations

Bluesky datasets serve as testbeds for:

- **Dynamic network analysis**: Structural balance, signed link prediction, community detection, contagion spread, and higher-order diffusion on groups [2407.17451, 2505.11608, 2405.17571].
- **Content and engagement modeling**: Altmetrics, originality metrics, text-based stance classification using dense retrieval and LLMs, engagement analysis (likes, reposts, reply rates) [2507.18840, 2506.07606].
- **Comparative and migration studies**: Analysis of cross-platform user migration, peer contagion, and behavioral transplantation (e.g., Twitter→Bluesky) [2505.24801, 2501.12886].
- **Algorithmic curation and moderation**: Custom feed dynamics, user engagement with algorithmic rankings, and the impact of third-party labelers for decentralized moderation architectures [2405.17571, 2402.03239].
- **Limitations**: Datasets are hampered by incomplete coverage pre-custom feeds (pre-May 2023), lack of private actions, niche platform adoption (with consequential cultural/coverage biases), and evolving API/sampling schemas [2407.17451, 2505.11608, 2404.18984].

## 7. Significance and Prospects

The Bluesky Dataset family constitutes a unique resource for the network science, computational social science, and machine learning communities:

- **Open architecture**: Bluesky’s separation of identity, hosting, indexing, content feeds, and moderation exposes rich, composable data structures, facilitating novel research in decentralized systems [2402.03239, 2408.12449].
- **Granularity, scale, and diversity**: Millisecond-resolved multilayer graphs, higher-order hypergraph structures, population-scale activity streams, and integration with external knowledge bases make possible analyses previously infeasible on traditional, siloed platforms.
- **Use in benchmarking**: These datasets allow development and evaluation of new methods in signed-temporal network analysis, hypergraph learning, stance detection (user- and post-level), rare-action prediction, moderation, diffusion modeling, and migration dynamics.
- **Researcher access and legal/ethical compliance**: All major datasets are released under FAIR-compliant terms with open code, documented limitations, robust anonymization, and explicit adherence to Bluesky’s public data policy [2407.17451, 2404.18984, 2505.11608, 2501.12886].

The Bluesky Dataset ecosystem is thus central to current methodological advancement in network, content, and group-dynamics research on decentralized, open social platforms.

Source: https://www.emergentmind.com/topics/bluesky-dataset