MuLTa-Telegram: A Fine-Grained Italian and Polish Dataset for Hate Speech and Target Detection This paper introduces the MuLTa-Telegram dataset, a Multi- Lingual and multi-Target dataset specifically developed to detect hate speech on Telegram, an understudied yet influential platform in which … Read More
Scientific Publications
Generative AI and the Threat to Thinking
Generative AI and the Threat to Thinking Information security is concerned with maintaining the integrity of the information ecosystem. The proliferation of content created using generative artificial intelligence can overwhelm the ability of people to process information. Consideration of a … Read More
Activities and Needs of European Fact-checkers as a Basis for Designing Human-Centered AI Systems
Autonomation, Not Automation: Activities and Needs of European Fact-checkers as a Basis for Designing Human-Centered AI Systems To mitigate the negative effects of false information more effectively, the development of Artificial Intelligence (AI) systems to assist fact-checkers is needed. Nevertheless, … Read More
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts Recent LLMs are able to generate high-quality multilingual texts, indistinguishable for humans from authentic human-written ones. Research in machine-generated text detection is however mostly focused on the English language and … Read More
EuroVerdict: A multilingual dataset for verdict generation against misinformation
EuroVerdict: A multilingual dataset for verdict generation against misinformation Misinformation is a global issue that shapes public discourse, influencing opinions and decision-making across various domains. While automated fact-checking (AFC) has become essential in combating misinformation, most work in multilingual settings … Read More
Typographic Attacks in a Multi-Image Setting
Typographic Attacks in a Multi-Image Setting Large Vision-Language Models (LVLMs) are susceptible to typographic attacks, which are misclassifications caused by an attack text that is added to an image. In this paper, we introduce a multi-image setting for studying typographic … Read More
Fine-grained Fallacy Detection with Human Label Variation
Fine-grained Fallacy Detection with Human Label Variation We introduce FAINA, the first dataset for fallacy detection that embraces multiple plausible answers and natural disagreement. FAINA includes over 11K span-level annotations with overlaps across 20 fallacy types on social media posts … Read More
LLMs vs Established Text Augmentation Techniques for Classification
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs? The generative large language models (LLMs) are increasingly being used for data augmentation tasks, where text samples are LLM-paraphrased and then used for classifier fine-tuning. … Read More
ModaFact dataset with Event Factuality and Modality in Italian
Authors: Rovera Marco, Cristoforetti, Serena, Tonelli Sara ModaFact is a textual dataset annotated with Event Factuality and Modality in Italian. ModaFact’s goal is to model in a joint way factuality and modality values of event-denoting expressions in text. Original texts (sentences) … Read More
ModaFact: Multi-paradigm Evaluation for Joint Event Modality and Factuality Detection
Authors: Rovera Marco, Cristoforetti, Serena, Tonelli Sara Factuality and modality are two crucial aspects concerning events, since they convey the speaker’s commitment to a situation in discourse as well as how this event is supposed to occur in terms of … Read More