Global AI Content Impact Dataset 2020-2025
- The paper provides sector-specific estimates on AI adoption and job loss using a 200-observation industry-country-year panel.
- Global AI Content Impact Dataset is an industry panel that captures AI adoption rates and job loss percentages without measuring textual or media content.
- The analysis employs a three-stage OLS framework, revealing heterogeneous effects, notably a significant inverse link between AI adoption and job loss in retail.
Searching arXiv for papers explicitly mentioning the Global AI Content Impact Dataset and closely related dataset-building work. arxiv_search(query="6\6 AI Content Impact Dataset6\6 OR 6\6 Content Impact Dataset6\6 OR 6\6 Impact of AI Adoption on Retail Across Countries and Industries6\6 OR 6\6 OR 6\6 over Preference: The Impact of AI-Generated Content on Online Content Ecology6\6 OR 6\6 Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit6\6 max_results=6 OR \6\6) The Global AI Content Impact Dataset is identified in the literature as the “Global AI Content Impact Dataset (6 OR \6\6 OR \6\6-6 OR \6\6 OR \6 OR \6)”, cited as a Kaggle dataset by Soundankar (6 OR \6\6 OR \6 OR \6), and used as the sole empirical foundation in a panel study of AI adoption and job loss across countries and industries. In that usage, the dataset is not a social-media corpus and is not documented as a text, image, or multimodal content archive; rather, it is treated as an industry-country-year panel centered on AI adoption rate (%) and job loss rate (%). The published description reports 6 OR \6\6\6^ observations across Australia, China, France, Japan, and the United Kingdom and ten industries, over 6 OR \6\6 OR \6\6^ through 6 OR \6\6 OR \6 OR \6^, with the main empirical contribution being sectorally heterogeneous estimates rather than a direct measurement of online content effects (&&&6\6&&&).
6 OR \6. Provenance and nominal scope
The dataset is named exactly as the Global AI Content Impact Dataset (6 OR \6\6 OR \6\6-6 OR \6\6 OR \6 OR \6) and is described as “publicly available on Kaggle.” In the paper that uses it, the observational unit is an industry-country-year cell, and the covered industries are manufacturing, finance, healthcare, education, marketing, media, retail, automotive, gaming, and legal services. The same paper states that the dataset includes 6 OR \6\6\6^ observations defined by industry, country, and year and spans 6 OR \6\6 OR \6\6^ through 6 OR \6\6 OR \6 OR \6^ (&&&6\6&&&).
The documented coverage creates an immediate internal tension. Five countries multiplied by ten industries and six years yields PRESERVED_PLACEHOLDER_6\6, whereas the reported sample size is 6 OR \6\6\6^. This suggests that the panel is unbalanced or only partially covered, although the paper does not explicitly label it that way. The text further notes that the paper provides no data appendix, panel-completeness table, or missingness matrix, so the exact source of the discrepancy cannot be resolved from the publication alone (&&&6\6&&&).
A second definitional issue concerns the phrase “content impact.” Despite that title, the documented use treats the dataset as a labor-market panel and “does not discuss any textual or media-content component.” A plausible implication is that the dataset’s name overstates or at least misstates the empirical object as it appears in the published analysis, because the measurable constructs are sectoral adoption and job loss rather than content production, circulation, exposure, or engagement (&&&6\6&&&).
6 OR \6. Observational structure and recorded variables
The key variables drawn from the dataset are AI adoption rate (%) and job loss rate (%). The former is the independent variable, intended to capture the extent of AI uptake within an industry-country-year observation; the latter is the dependent variable and is interpreted in the paper as “job loss due to AI” or “job loss rate.” The same source emphasizes that all variables retain original percentage scales without further normalization (&&&6\6&&&).
The paper’s preprocessing description is limited but explicit. It states that observations with missing values for AI adoption rate, job loss rate, or industry identifiers were removed. It then states that remaining missing entries were imputed using industry-median values, and that outliers beyond three standard deviations from the mean were excluded. No counts are reported for dropped observations, imputed values, or excluded outliers, and the paper does not clarify whether the final count of 6 OR \6\6\6^ refers to the raw or cleaned dataset (&&&6\6&&&).
The publication also reports sectoral descriptive means. Average AI adoption rates and job loss rates by industry are listed as follows: Automotive PRESERVED_PLACEHOLDER_6 OR \6; Education PRESERVED_PLACEHOLDER_6 OR \6; Finance PRESERVED_PLACEHOLDER_6 OR \6; Gaming PRESERVED_PLACEHOLDER_6 OR \6; Healthcare PRESERVED_PLACEHOLDER_6 OR \6; Legal $56.08, 28.23$; Manufacturing $57.01, 32.75$; Marketing $54.24, 19.58$; Media $47.26, 22.75$; Retail PRESERVED_PLACEHOLDER_6 OR \6\6. These summaries show that manufacturing has the highest average job loss, while retail has one of the lower average job loss rates and relatively modest average AI adoption (&&&6\6&&&).
What is not documented is equally consequential. The paper does not define the denominator for job loss rate; does not state whether AI adoption is measured from surveys, investments, firm reports, or another source; and does not specify coding conventions, source harmonization, or lower-level aggregation procedures. This suggests that the dataset is usable for regression exercises in its published form, but not yet fully interpretable as a standardized measurement instrument (&&&6\6&&&).
6 OR \6. Empirical use in sectoral labor-market analysis
The published empirical design is a three-stage OLS framework. The first stage estimates a full-sample association:
PRESERVED_PLACEHOLDER_6 OR \6 OR \6^
The second stage estimates separate regressions by industry:
PRESERVED_PLACEHOLDER_6 OR \6 OR \6^
The third stage introduces interaction terms for marketing and retail:
PRESERVED_PLACEHOLDER_6 OR \6 OR \6^
This structure is used to move from pooled association to sector-specific heterogeneity and then to differential marginal effects in the two industries judged “closest to significance” (&&&6\6&&&).
The interaction specification implies that the retail marginal slope is PRESERVED_PLACEHOLDER_6 OR \6 OR \6. Using the reported coefficients, the paper interprets the retail marginal effect as approximately
PRESERVED_PLACEHOLDER_6 OR \6 OR \6^
so that, in retail, a one-percentage-point increase in AI adoption is associated with an estimated 6\6.6 OR \6 OR \66^ percentage-point decrease in job loss. The interaction coefficient PRESERVED_PLACEHOLDER_6 OR \66^ is interpreted as an additional marginal effect specific to retail (&&&6\6&&&).
The methods section also states that all models report coefficients, standard errors, t-statistics, and p-values using unadjusted standard errors. The paper further mentions country fixed-effects specifications, subsample regressions by development status, and re-estimation after excluding outliers, but it does not print the fixed-effects equation, full robustness tables, or diagnostic statistics. A plausible implication is that the dataset was treated as suitable for associational inference, but not for a high-specification panel design with extensive covariate control (&&&6\6&&&).
6 OR \6. Reported findings and substantive interpretation
At the full-sample level, the paper reports essentially no linear association between AI adoption and job loss. The Pearson correlation is PRESERVED_PLACEHOLDER_6 OR \67, and the OLS coefficient is reported in the abstract as PRESERVED_PLACEHOLDER_6 OR \68 with PRESERVED_PLACEHOLDER_6 OR \69. The constant is PRESERVED_PLACEHOLDER_6 OR \6\6, with standard error PRESERVED_PLACEHOLDER_6 OR \6 OR \6, PRESERVED_PLACEHOLDER_6 OR \6 OR \6, and PRESERVED_PLACEHOLDER_6 OR \6 OR \6^ (&&&6\6&&&).
The industry-specific regressions are more differentiated. The reported slopes and p-values are: Marketing PRESERVED_PLACEHOLDER_6 OR \6 OR \6; Education PRESERVED_PLACEHOLDER_6 OR \6 OR \6; Automotive PRESERVED_PLACEHOLDER_6 OR \66; Healthcare PRESERVED_PLACEHOLDER_6 OR \67; Media PRESERVED_PLACEHOLDER_6 OR \68; Manufacturing PRESERVED_PLACEHOLDER_6 OR \69; Finance PRESERVED_PLACEHOLDER_6 OR \6\6; Gaming PRESERVED_PLACEHOLDER_6 OR \6 OR \6; Legal PRESERVED_PLACEHOLDER_6 OR \6 OR \6; Retail PRESERVED_PLACEHOLDER_6 OR \6 OR \6. Marketing and retail are thus the two sectors “closest to significance,” but with opposite signs: marketing positive, retail negative (&&&6\6&&&).
The interaction model is the source of the paper’s main substantive conclusion. The reported coefficients are: Constant PRESERVED_PLACEHOLDER_6 OR \6 OR \6; AI Adoption PRESERVED_PLACEHOLDER_6 OR \6 OR \6; Marketing PRESERVED_PLACEHOLDER_6 OR \66^ AI Adoption PRESERVED_PLACEHOLDER_6 OR \67; Retail PRESERVED_PLACEHOLDER_6 OR \68 AI Adoption PRESERVED_PLACEHOLDER_6 OR \69. On that basis, the paper concludes that higher AI adoption is linked to lower job loss in retail, whereas no significant aggregate effect appears in the pooled sample. The author interprets this as evidence that AI in retail may function as a productivity-enhancing, labor-enabling technology, citing intelligent replenishment systems and cashierless checkout as examples (&&&6\6&&&).
This suggests that the dataset’s principal analytical value lies in exposing sectoral heterogeneity rather than estimating a single global effect. It also suggests that the name “Global AI Content Impact Dataset” obscures the narrower empirical role documented in the paper: a compact comparative panel for sector- and country-level labor-market association studies (&&&6\6&&&).
6 OR \6. Documentation limits and evidence-quality questions
The paper exposes several limitations of the dataset, or at minimum of its documented use. The most serious is insufficient variable definition. The text does not specify how AI adoption rate (%) or job loss rate (%) are measured, how comparable they are across countries and industries, or what source instruments underlie them. The panel coverage is also unclear because the reported years and sample size do not align. The baseline regressions include no classical controls beyond the intercept and AI adoption, and the published tables omit confidence intervals, many standard errors, and most diagnostics (&&&6\6&&&).
Additional inferential constraints are stated or implied. The country-grouping exercise distinguishes “developed economies” as UK, France, Australia and gives China as the example of an emerging economy, but Japan is not classified in that sentence. The robustness discussion reports that qualitative results persist under country fixed effects and outlier exclusion, but the corresponding specifications are not fully displayed. The use of unadjusted standard errors is methodologically consequential because clustering or heteroskedasticity correction is not applied in the main tables (&&&6\6&&&).
A broader curation issue appears in the surrounding literature. A supposed empirical source on Pixiv and AI-generated content, “Understanding the Impact of AI Generated Content on Social Media: The Pixiv Case” (&&&6 OR \67&&&), is described in the technical summary as “an ACM LaTeX template/sample manuscript” rather than an actual Pixiv study. It contains no dataset relevant to Pixiv, AI-generated content, or platform impacts and is recommended for exclusion from substantive evidence synthesis. This episode is not evidence about the Global AI Content Impact Dataset itself, but it illustrates a directly relevant problem for AI-impact data infrastructures: document misidentification can generate false extractions unless ingestion pipelines explicitly screen out template artifacts and wrong-paper matches (&&&6 OR \67&&&).
Taken together, these issues imply that the dataset is best treated as a documented but weakly specified panel resource. It supports descriptive and associational analysis in the form published, but not precise cross-study harmonization without consultation of the underlying Kaggle source and its metadata (&&&6\6&&&).
6. Position within the broader AI-content-impact data landscape
The broader literature shows that the Global AI Content Impact Dataset (6 OR \6\6 OR \6\6-6 OR \6\6 OR \6 OR \6) occupies only one part of a much larger measurement landscape. In contrast to this industry-country-year labor-market panel, RedNote-Vibe is a five-year Xiaohongshu dataset with note title, text content, tags, publication timestamp, engagement metrics, and topic domain, explicitly designed to capture the temporal dynamics of AI-generated text on social media (&&&6 OR \6\6&&&). “Scale over Preference: The Impact of AI-Generated Content on Online Content Ecology” measures creator production, consumer preference, and recommendation exposure on a leading Chinese video-sharing platform, operationalizing impact as a restructuring of online content ecology rather than as labor-market association (&&&6 OR \6 OR \6&&&). “The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit” provides subreddit-day causal evidence on downstream engagement after Google AI Overviews and AI Mode, making interface design itself part of the impact construct (&&&6 OR \6 OR \6&&&).
Other resources emphasize representation, discourse, and human response. “Landscape of Generative AI in Global News” assembles 6 OR \6 OR \6,86 OR \67 English-language news articles from 76\6 OR \6^ outlets and studies topics, sentiment, and spatiotemporal variation in media framing (&&&6 OR \6 OR \6&&&). “Towards Leveraging News Media to Support Impact Assessment of AI Technologies” constructs 6 OR \67,689 description-impact pairs from AI-related news and uses them to generate negative impacts for impact assessment workflows (&&&6 OR \6 OR \6&&&). MhAIM contains 6 OR \6 OR \6 OR \6,6 OR \6 OR \6 OR \6^ online posts and introduces trustworthiness, impact, and openness as human-response metrics for multimodal AI content (&&&6 OR \6 OR \6&&&). These datasets are much closer to direct measurement of content perception, circulation, and reception than the industry-country-year structure documented for the Global AI Content Impact Dataset (&&&6 OR \6 OR \6&&&).
A further cluster of work addresses global representativeness and infrastructure. “Curating Grounded Synthetic Data with Global Perspectives for Equitable AI” builds a news-grounded synthetic dataset from 6 OR \6 OR \6^ languages and 6 OR \6 OR \6 OR \6^ countries (&&&6 OR \67&&&). DeepInnovationAI links 6 OR \6,6 OR \6 OR \6 OR \6,96 OR \69 AI-related papers and 6 OR \6,6 OR \6 OR \66,6 OR \6\6 OR \6^ AI-related patents, with a semantic transfer layer between research and industrial application (&&&6 OR \68&&&). BioMedJImpact derives journal-year AI engagement features from 6 OR \6.76 OR \6^ million PubMed Central articles across 6 OR \6,76 OR \6 OR \6^ journals (&&&6 OR \69&&&). “The Multilingual Divide and Its Impact on Global AI Safety” argues that multilingual coverage, local context, and language-specific safety evaluation are first-order design requirements for any genuinely global AI dataset (&&&6 OR \6\6&&&).
This comparison suggests that the Global AI Content Impact Dataset, as presently documented, is best understood not as a comprehensive “global AI content impact” observatory, but as a small comparative panel for AI adoption and job loss. Its title aligns only partially with current usage. A plausible implication is that a fuller system bearing that name would need to integrate at least four layers already visible in adjacent work: content provenance and detection, engagement and exposure, human judgments and behavioral response, and regional, linguistic, and policy metadata (&&&6 OR \6 OR \6&&&, &&&6 OR \6\6&&&, &&&6 OR \6\6&&&).