الاثنين، 12 مايو 2014

الخميس، 13 يونيو 2013

DGT-Translation Memory

he DGT Translation Memory is currently available in 23 languages. The following table shows the coverage, expressed in the total number of translation units available for each language, separately for the DGT-TM releases 2007, 2011 and 2012. For the number of aligned translation units for each language pair and further statistics of the release DGT-TM-2011.
 
Language
Language code
Number of units in DGT - release 2007
Number of units in DGT - release 2011
Number of units in DGT - release 2012
Number of units in DGT - release 2013
English
EN
2 187 504
2 286 514
322 377
538949
Bulgarian
BG
708 658
454 812
272 595
378416
Czech
CS
890 025
1 985 152
283 826
478709
Danish
DA
433 871
1 997 649
279 746
472024
German
DE
532 668
1 922 568
284 072
472081
Greek
EL
371 039
1 901 490
285 483
462304
Spanish
ES
509 054
1 907 649
284 977
477829
Estonian
ET
1 047 503
1 867 786
280 549
461051
Finnish
FI
514 868
1 881 558
283 213
459927
French
FR
1 106 442
1 853 773
273 961
462431
Irish
GA
0
0
2 848
0
Hungarian
HU
1 159 975
1 869 246
284 282
480050
Italian
IT
542 873
1 926 532
281 503
474035
Lithuanian
LT
1 126 255
1 867 176
286 018
461375
Latvian
LV
1 120 835
1 859 781
284 641
461190
Maltese
MT
1 021 855
461 865
263 804
386677
Dutch
NL
502 557
1 914 628
281 683
469990
Polish
PL
1 052 136
1 879 469
282 551
454864
Portuguese
PT
945 203
1 922 585
284 310
471810
Romanian
RO
650 735
470 303
270 763
393398
Slovak
SK
1 065 399
1 894 676
285 422
480073
Slovene
SL
1 026 668
1 903 453
284 642
479147
Swedish
SV
555 362
1 934 964
283 589
478204
ALL
ALL
19,071,485
37,963,629
6,226,855
10,154,534
Size of DGT's Translation Memory expressed as the total number of translation units per language for each of the 23 official EU languages. The DGT Translation Memory only includes data in Irish since release DGT-TM-2012.

الجمعة، 31 مايو 2013

KALIMAT Multipurpose Arabic Corpus

KALIMAT a Multipurpose Arabic Corpus

We are pleased to announce the immediate availability of KALIMAT 1.0,

KALIMAT is an Arabic natural language resource that consists of:
1) 20,291 Arabic articles collected from the Omani newspaper Alwatan by (Abbas et al. 2011).
2) 20,291 Extractive Single-document system summaries.
3) 2,057 Extractive Multi-document system summaries.
4) 20,291 Named Entity Recognised articles.
5) 20,291 Part of Speech Tagged articles.
6) 20,291 Morphologically Analyse articles.

The data collection articles fall into six categories:
culture, economy, local-news, international-news, religion, and sports.

The process of creating KALIMAT was applied to the entire data collection (20,291 articles).

ARABIC LEARNER CORPUS v1 المدونة اللغوية لمتعلمي اللغة العربية

ARABIC LEARNER CORPUS v1

Contents
The first version of the Arabic learner corpus (ALC) comprises a collection of texts written by learners of Arabic in Saudi Arabia. The corpus covers two types of students, non-native Arabic speakers (NNAS) learning Arabic as a second language (ASL) for academic purpose (AAP), and native Arabic speaking students (NAS) learning to improve their written Arabic. Both groups are males in pre-university level.
The current version of ALC has been captured in November and December 2012, and it includes a total of 31272 words, 215 written texts (narrative and discussion) produced by 92 students from 24 nationalities and 26 different L1 backgrounds. 181 texts (84%) were written in class (timed essays), while 34 (16%) produced at home (untimed essays). Average length of the texts is 145 words. 95% of the texts were hand-written, so they had to be transcribed into a suitable computerised form. All identity information (e.g, names, contacts, dates of birth, etc.) have been removed from transcriptions.
Files format
Three types of non-annotated files have been generated: (1) with no header, (2) with metadata header in Arabic, (3) and in English. They are available in two formats, txt and XML. The metadata information enables researchers to identify characteristics of text and its producer in each transcription. The original hand-written sheets are also available after they have been scanned and saved into PDF-format files.
All corpus files were named in a method which indicates the basic characteristics of the text and its author (e.g. S038_T2_M_Pre_NNAS_W_C). They are in order, student identifier number, text number, author gender, level of study, nativeness, text mode, and place of text production.
Future work
AS a next stage, the entire corpus will be annotated for errors, and word-tagged with morphological tags to identify part of speech and certain grammatical sub-categories. Additionally, the correct form will be reconstructed by correcting the mistakes all.
Annotating errors will be performed using a detailed error-type tagset, which has been developed for Arabic learner corpora in general and to be used in the present corpus in particular (Alfaifi and Atwell, 2012). In future, further version will be issued including more materials (written and spoken), different genders (male and female), and different levels of study (pre-university and university).
References
Alfaifi, A. and Atwell, E. (2012). Arabic Learner Corpora (ALC): A Taxonomy of Coding Errors. In: the 8th International Computing Conference in Arabic (ICCA 2012), 26 - 28 December 2012, Cairo, Egypt.

للمزيد و للتحميل
http://www.comp.leeds.ac.uk/scayga/alc/index.html