Loading…

CLEU ‐ A Cross‐language english‐urdu corpus and benchmark for text reuse experiments

Text reuse is becoming a serious issue in many fields and research shows that it is much harder to detect when it occurs across languages. The recent rise in multi‐lingual content on the Web has increased cross‐language text reuse to an unprecedented scale. Although researchers have proposed methods...

Full description

Saved in:

Bibliographic Details
Published in:	Journal of the American Society for Information Science and Technology 2019-07, Vol.70 (7), p.729-741
Main Authors:	Muneer, Iqra, Sharjeel, Muhammad, Iqbal, Muntaha, Nawab, Rao Muhammad Adeel, Rayson, Paul
Format:	Article
Language:	English
Subjects:	Benchmarks Corpus linguistics English language Language Languages Plagiarism Researchers Translation methods and strategies Urdu Urdu language
Citations:	Items that this one cites Items that cite this one
Online Access:	Get full text
Tags:	Add Tag No Tags, Be the first to tag this record!

Description
Summary:	Text reuse is becoming a serious issue in many fields and research shows that it is much harder to detect when it occurs across languages. The recent rise in multi‐lingual content on the Web has increased cross‐language text reuse to an unprecedented scale. Although researchers have proposed methods to detect it, one major drawback is the unavailability of large‐scale gold standard evaluation resources built on real cases. To overcome this problem, we propose a cross‐language sentence/passage level text reuse corpus for the English‐Urdu language pair. The Cross‐Language English‐Urdu Corpus (CLEU) has source text in English whereas the derived text is in Urdu. It contains in total 3,235 sentence/passage pairs manually tagged into three categories that is near copy, paraphrased copy, and independently written. Further, as a second contribution, we evaluate the Translation plus Mono‐lingual Analysis method using three sets of experiments on the proposed dataset to highlight its usefulness. Evaluation results (f1=0.732 binary, f1=0.552 ternary classification) indicate that it is harder to detect cross‐language real cases of text reuse, especially when the language pairs have unrelated scripts. The corpus is a useful benchmark resource for the future development and assessment of cross‐language text reuse detection systems for the English‐Urdu language pair.
ISSN:	2330-1635 2330-1643
DOI:	10.1002/asi.24074