Automatic Combination of Sample Selection Strategies for Few-Shot Learning In few-shot learning, the selection of samples has a significant impact on the performance of the model. While effective sample selection strategies are well-established in supervised settings, research on large language … Read More
Scientific Publications
Authorship Attribution in Multilingual Machine-Generated Texts
Authorship Attribution in Multilingual Machine-Generated Texts As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult. While early efforts in MGT detection have focused on binary classification, the growing … Read More
Automated In-the-Wild Data Collection for Continual AI Generated Image Detection
Automated In-the-Wild Data Collection for Continual AI Generated Image Detection The rapid advancement of generative Artificial Intelligence (AI) has introduced significant challenges for reliable AI-generated image detection. Existing detectors often suffer from performance degradation under distribution shifts and when encountering … Read More
MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models
MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models Large language models (LLMs) are beginning to reshape how media professionals verify information, yet automated support for detecting check-worthy claims—a key step in the fact-checking process—remains limited. We … Read More
Beyond Speculation: Measuring the Growing Presence of LLM-Generated Texts in Multilingual Disinformation
Beyond Speculation: Measuring the Growing Presence of LLM-Generated Texts in Multilingual Disinformation Increased sophistication of large language models (LLMs) and the consequent quality of generated multilingual text raises concerns about potential disinformation misuse. While humans struggle to distinguish LLM-generated content … Read More
Bots into the Fediverse dataset
Bots into the Fediverse This dataset contains anonymized features for bot detection on Mastodon (Fediverse). It was created for the accompanying paper and consists of accounts labeled as bot or non-bot, collected from publicly accessible content via the Mastodon Application … Read More
Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approaches
Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approaches Retrieval of previously fact-checked claims is a well-established task, whose automation can assist professional fact-checkers in the initial steps of information verification. Previous works have mostly tackled the … Read More
Comparing Specialised Small and General Large Language Models on Text Classification
Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance When solving NLP tasks with limited labelled data, researchers typically either use a general large language model without further update, or use … Read More
Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking
Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking Natural Language Processing and Generation systems have recently shown the potential to complement and streamline the costly and timeconsuming job of professional fact-checkers. In this work, we lift several constraints of … Read More
A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of LLM
A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language Models In the age of social media and generative AI, the ability to automatically assess the credibility of online content has become increasingly critical, … Read More