In the first year, all participants established a unified vision and shared methodologies, crucial for developing technical tools and identifying ground truth data for AI training. A major challenge involved accessing social media data: changes to Twitter/X API policies and the discontinuation of Crowdtangle required diversifying data sources. A novel methodology for collecting Telegram data was devised, and YouTube access continued. Privacy compliance remained paramount, with all teams upholding privacy-by-design principles throughout.
From a scientific perspective, AI-driven methods were developed for textual, speech, audio, visual, and multimodal data analysis. A system architecture was designed to estimate the trustworthiness of claims or media items. Socio-behavioral and human-centered research — including focus groups, interviews, and ethnographic fieldwork — informed AI4TRUST model development and platform design, particularly around trust, explainability, and user requirements.
From a technical perspective, the AI4TRUST platform adopted a microservice-based architecture to integrate modules and support concurrent data processing, with ethical and privacy requirements embedded throughout development.
As regard the second reporting period (M15–M38), the main achievement were the following:
- Data Collection (WP2). A scalable cloud platform was deployed for continuous ingestion from YouTube, Telegram, Bluesky, and online news. Access to Twitter/X and Meta was denied, requiring strategic adaptation. By M38, the data lake exceeds 2.3 TB with over 1 billion entries: ~630k news articles, ~160k YouTube videos with 16M comments, 700M Bluesky messages, and 131M Telegram messages from 4.5k monitored channels. Ground truth datasets were collected, including the multilingual EuroVerdict dataset (1,642 verdicts in 8 languages) and a deepfake audio dataset (2,000+ segments).
- AI Analysis Methods (WP3). Advanced AI tools were delivered across all modalities:
* Textual: A disinformation signals tool was implemented that detects 39 signals across 6 manipulation tactics in 8 languages (weighted F1 ~56%). A multilingual verdict generation tool improved ROUGE-L from 0.238 to 0.312. A cross-lingual fact-checked claim retrieval tool supporting 25 language combinations was released.
* Audio: Speech-to-text models for all 8 languages were finalized. Two deepfake audio detection methods achieved state-of-the-art performance (EER 0.78% on ASVspoof 2021 LA).
*Visual: A faster reverse video search tool and a deepfake video API (90.5 AUC cross-dataset) were deployed. A sensational content detection method achieved 89.4% accuracy on 9,576 annotated images.
* Multimodal: The AuViRe method for temporal forgery localization achieved +30.5 AUC vs. state-of-the-art. The LAVAD video anomaly detection method achieved a 10x speed-up with maintained accuracy.
- Disinformation Warning System (WP3). Three DWS versions were developed. The final version aggregates outputs from six AI tools using an ensemble of LightGBM classifiers with isotonic regression calibration, providing interpretable triage categories and feature attribution for explainability.
- Trustworthy AI (WP4). Extensive qualitative fieldwork — expert interviews, co-design sessions, and post-piloting interviews — evaluated trust and explainability in practice. Social network analysis on Telegram, YouTube, and Bluesky produced novel findings on far-right network assortativity, re-information patterns, and coordinated inauthentic behaviour.
- Platform (WP5). The platform progressed toward an MVP following revision of 86 requirements across three evaluation rounds. The MVP features a redesigned Toolbox Dashboard, an enhanced Monitoring and Human Validation Dashboard, and a new Infodemic Observatory Dashboard. New data sources and a configurable interface were introduced.
- Piloting and Evaluation (WP6). Three evaluation rounds were completed, engaging 206 end users across media professionals, fact-checkers, policy experts, and students. Feedback drove iterative improvements. WP3 outcomes were published in 56 peer-reviewed scientific papers, and AI technologies were presented at the EBU Data Technology Seminar 2025.