Wednesday 09 April 2025
Researchers have created a new benchmark that allows them to evaluate and compare the effectiveness of attacks designed to evade machine-generated text detectors. The team’s findings highlight the importance of balancing attack effectiveness, text quality, and computational cost when developing these evading attacks.
Machine-generated texts (MGTs) are becoming increasingly common in our digital lives, with applications ranging from chatbots to academic papers. However, this trend has also led to concerns about the spread of misinformation and plagiarism. To combat these issues, machine-generated text detectors have been developed to identify MGTs. But, as researchers have discovered, attackers are now using evading attacks to circumvent these detectors.
The new benchmark, called TH-Bench, is designed to evaluate the effectiveness of these evading attacks against machine-generated text detectors. The team tested six state-of-the-art attacks on 13 different detectors across six datasets, spanning 19 domains and generated by 11 widely used large language models (LLMs). Their results show that no single attack excels across all three dimensions: evading effectiveness, text quality, and computational cost.
One of the key findings is that attacks that prioritize evading effectiveness often compromise text quality and computational cost. For example, an attack that modifies a sentence to evade detection may result in a text that is less coherent or fluent than the original. Similarly, an attack that uses more resources to modify the text may increase its computational cost.
The researchers also introduced two new techniques: Quality-Preserving Attack (QPA) and Attack Blending. QPA aims to preserve the quality of the generated text while still evading detection. This is achieved by incorporating constraints on text quality metrics, such as fluency, coherence, and semantic similarity. The team found that QPA significantly enhanced both semantic similarity metrics across all three attack types.
Attack Blending involves combining different attack methods to create a more effective and efficient approach. By alternating between two or more attacks, the resulting text is less likely to be detected by machine-generated text detectors. However, this approach requires careful selection of attack methods and optimization techniques to ensure that the combined attack achieves better results than using a single method.
The study highlights the importance of considering the trade-offs between evading effectiveness, text quality, and computational cost when developing evading attacks. By optimizing for one aspect, attackers may inadvertently compromise another critical dimension. The researchers suggest that future work should focus on refining these techniques to achieve better results and addressing the limitations of current approaches.
Cite this article: “Evading AI-Generated Text Detectors: A Comprehensive Evaluation of Attack Strategies and Quality-Preserving Attacks”, The Science Archive, 2025.
Machine-Generated Text, Evading Attacks, Text Detectors, Benchmarking, Language Models, Large Language Models, Quality-Preserved Attack, Attack Blending, Computational Cost, Misinformation Detection







