Data sources
This study utilizes three publicly accessible data categories spanning 2022–2024. Platform transparency reports from Meta (n = 12 quarterly reports), Google (n = 8), and Twitter (n = 7) provide violation volumes, content taxonomies, and proactive detection rates for K-means clustering (Phase 1) and Ridge regression variables (Phase 2).
Specifically, the dataset comprises quarterly transparency reports from: (1) Meta’s Facebook and Instagram platforms combined (12 quarterly reports covering Q1 2022 through Q4 2024), (2) Google’s YouTube platform (8 quarterly reports covering Q1 2023 through Q4 2024), and (3) Twitter/X platform (7 quarterly reports covering Q3 2022 through Q1 2024, with irregular publication schedules).
These three platforms were selected based on: (a) classification as Very Large Online Platforms (VLOPs) under EU Digital Services Act with monthly active users exceeding 45 million, (b) publicly available English-language quarterly transparency reports with consistent content moderation metrics, (c) comprehensive reporting of both proactive detection rates and violation volumes across multiple content categories, and (d) consistent quarterly publication throughout the study period. Other VLOPs were excluded due to incomplete quarterly reporting coverage or inconsistent content category taxonomies limiting cross-platform comparability.
Data were collected in May 2025 from official sources: Meta: “Community Standards Enforcement Report” (transparency.fb.com); Google: “YouTube Community Guidelines Enforcement” (transparencyreport.google.com); Twitter/X: “Twitter Transparency Report” (transparency.twitter.com); FBI IC3: “Internet Crime Report” (ic3.gov, 2022–2024); Europol: “Internet Organised Crime Threat Assessment” (europol.europa.eu, 2022–2024).
The final sample comprises 27 platform-quarter observations, distributed as 12 observations from Meta, 8 from Google, and 7 from Twitter.
Cybercrime databases (FBI IC3 and Europol IOCTA, n = 6 annual reports total) supply criminal typologies and financial loss data for risk calibration. Industry documentation from Trust and Safety Professional Association and publicly disclosed platform technology investments from earnings reports, supplemented by academic game-theoretic frameworks, inform parameter calibration (Phase 3). This platform-centric operational data approach mitigates traditional limitations, notably victim underreporting and temporal lags, inherent in retrospective datasets, offering a comparatively more timely basis for risk assessment than traditional post-incident analysis, though subject to the data limitations discussed below. Data Collection Procedure: Platform transparency reports were accessed from official transparency centers (Meta: transparency.fb.com; Google: transparencyreport.google.com; Twitter: transparency.twitter.com) in May 2025. Data extraction followed a systematic protocol capturing: total content actioned by category, proactive detection rates distinguishing AI-initiated versus human-initiated actions, appeal volumes and resolution outcomes, restoration rates, and temporal metadata. Extracted data were validated against platform-reported aggregate statistics to ensure consistency with official disclosures.
Several data limitations warrant acknowledgment. Platform transparency reports are self-reported disclosures subject to heterogeneous reporting standards across platforms, including differences in how violations are counted, how proactive detection is defined, and what content categories are disclosed. Strategic disclosure incentives tied to reputational and regulatory pressures may further influence the scope and framing of reported metrics. To mitigate these concerns, the study cross-validates against platform-reported aggregate statistics and employs relative measures such as proactive detection rates as proportions rather than absolute counts; nevertheless, platform-reported metrics should be interpreted as indicators of disclosed enforcement activity rather than comprehensive measures of actual moderation performance. Additionally, the final sample of 27 platform-quarter observations, while sufficient for the exploratory framework employed here, imposes constraints on statistical power for Phase 2 regression analysis, a limitation addressed further in Sect. 4.4 and reflected in the cautious interpretation of results in Sect. 6.
Variable operationalization
This section operationalizes data sources into analytical variables for the tripartite framework. Phase 1 clustering variables stratify cybercrime content into 12 categories derived from harmonized platform taxonomy classifications. These standardized content types comprise: (1) Financial fraud (including investment scams and deceptive advertising, mapped from Meta’s “financial scam” and Google’s “deceptive financial content”), (2) Phishing attacks (credential theft and account impersonation), (3) Malware distribution (viruses, trojans, and ransomware), (4) Hate speech (discriminatory content based on protected characteristics), (5) Misinformation (deliberately false news and unverified claims), (6) Spam (unsolicited commercial messages and fake engagement), (7) Human trafficking recruitment (exploitation and forced labor content), (8) Identity theft (personal data harvesting and synthetic identity fraud), (9) Child exploitation (child sexual abuse material and grooming behaviors), (10) Violent extremism (terrorism recruitment and radicalization content), (11) Harassment and bullying (targeted threats and intimidation), and (12) Intellectual property violations (copyright and trademark infringement). These categories represent the most prevalent violation types consistently reported across Meta, Google, and Twitter transparency reports during 2022–2024, supplemented by cybercrime typologies from FBI IC3 and Europol IOCTA databases.
The 12 content types are defined in relation to five variables to account for complex risk attributes, including: complexity of technological expertise (scaled from IC3/IOCTA Attack Sophistication Levels, ranging from level 1 = simple spam to level 5 = Advanced Persistent Threats), potential for propagation (log-transformed level of violation quantities from platform reporting data, reflecting content virality), financial weight of consequences (IC3 Average Loss per Crime Type, ranging from $240 for spam to $6,150 for Investment Fraud), legal gravity of consequences (maximum federal sentence in months, ranging from 3 months for lesser crimes to child exploitation of 144 months in prison), and detection ease (derived from 1 minus platform reporting of proactive detection rates, where higher values denote increased review by human operators).
Phase 2 regression variables include detection effectiveness (proactive violations per total violations) as the dependent variable, platform scale (log MAU) as an independent variable, AI capability index (number of detection features publicly disclosed via transparency reports) as an independent variable, and review efficiency (inverse appeal processing time) as an independent variable, using observations of 27 platform-quarters. Both the dependent variable and the AI capability index are derived from platform self-reported transparency data and are therefore subject to the reporting limitations discussed above.
Phase 3 game-theoretic parameters calibrate automation costs ($0.0075/item) and human costs ($0.45/item) based on cloud computing pricing for ML inference and global moderator compensation data, respectively, detection probabilities derived from proactive detection rates stratified by content complexity (AI: 0.45–0.85; human: 0.85–0.95), and breach losses ($500-$5,000 from FBI IC3 financial impact data), enabling Nash equilibrium computation for optimal resource allocation across the four content clusters identified in Phase 1.
Data processing and standardization
Data processing addresses cross-platform taxonomic heterogeneity and temporal alignment requirements. Because platform transparency reports reflect self-reported enforcement activity with potentially heterogeneous measurement definitions across platforms (e.g., what constitutes a “proactive” detection may differ between Meta’s automated classifiers and Twitter’s reporting criteria), the standardization procedures described below aim to maximize cross-platform comparability, though residual measurement inconsistencies cannot be fully eliminated. Content type standardization harmonizes platform-specific violation categories into 12 unified classifications through systematic mapping: Meta’s “financial scam” and Google’s “deceptive financial content” both map to the standardized “financial fraud” category. Temporal alignment converts all platform reports to quarterly granularity, with Twitter’s irregular publication schedules disaggregated to quarterly equivalents where necessary and Google’s variable reporting frequency retained as available observations. Missing value imputation applies linear interpolation for sporadic gaps and employs industry median substitution for unreported appeal processing times. Variable standardization is also conducted, applying Z-score normalization to all variables in Phase 1 clustering analysis, complemented by mean-centering of Phase 2 regression variables to aid in understanding coefficients. Data validation measures also include assessing temporal reproducibility, interquartile range outlier detection, as well as comparisons to platform-reported totals, ensuring strong analysis validity for all variables, although data availability is available in excess of 90% for key variables, with documented imputation methods accounting for remaining gaps.
Sample characteristics and descriptive overview
Our framework employs two data structures for different analytical purposes. Phase 1 clustering analyzes content-type-level data: we harmonized platform-specific violation categories (e.g., Meta’s “financial scam” and Google’s “deceptive financial content” both classified as “financial fraud”) into 12 standardized content types, then constructed a 12 × 5 feature matrix by characterizing each content type across five risk dimensions. These dimensions integrate platform-derived metrics (propagation potential from log-transformed violation volumes; detection difficulty calculated as 1 minus average proactive detection rate) with external risk indicators (technical complexity from IC3/IOCTA attack sophistication descriptions, financial severity from IC3 average loss data, and legal consequences from federal sentencing guidelines). All variables were Z-score normalized prior to K-means clustering.
Phase 2 regression uses the original 27 platform-quarter observations (Meta: 12, Google: 8, Twitter: 7) spanning 2022Q1 to 2024Q4 to examine platform operational factors. Detection effectiveness serves as the dependent variable, ranging from 39.2% to 87.3% (M = 66.8%, SD = 12.8%), with strong cross-platform heterogeneity evident in platform-specific averages varying from 58.9% (Twitter) to 71.2% (Google). Platform scale ranges from 237 million to 2.96 billion monthly active users (log-transformed for analysis). Cluster assignments from Phase 1 are incorporated as contextual variables in regression modeling.
Phase 3 game-theoretic calibration employs five core parameter types derived from preceding analyses: automation costs ($0.0075/item), human review costs ($0.45/item), complexity-stratified detection probabilities (AI: 0.45–0.85; human: 0.85–0.95), and breach loss values ($500-$5,000 from FBI IC3 data).
The sample size of 27 platform-quarter observations for Phase 2 regression analysis warrants careful consideration. With three predictors in the base model, this yields approximately nine observations per predictor, a ratio that constrains the stability of coefficient estimates in conventional regression settings. Ridge regression with L2 regularization partially addresses this by shrinking coefficients toward zero and reducing variance from overfitting, but does not fully compensate for the limited degrees of freedom available for detecting smaller effects. The extended interaction model introduces additional terms that further reduce effective degrees of freedom, increasing the risk of model instability. These constraints inform the analytical strategy: effect sizes and confidence intervals are reported alongside significance tests to provide a more complete picture of the strength and precision of observed associations, and results from the interaction model are interpreted with appropriate caution. Leave-one-out cross-validation results are reported in Sect. 6 to assess the robustness of base model findings. Additional stability assessments, including bootstrap resampling of coefficient estimates, are recommended for future validation with larger samples.
Click Here For The Original Source.
