Print

This page describes the parseme-tsv format. The training (annotated) and blind (non-annotated) corpora files provided to participants in the PARSEME shared task on on automatic detection of verbal MWEs use this format. See a sample blind corpus file and a sample training corpus file for illustration. test

The parseme-tsv format is an UTF-8 textual four-column format. Columns must be separated by single tabulations, not by blank spaces. Empty fields are marked by underscores1 ('_'):

Example of an annotated training corpus file:

1      Delegates _  _
2 are _  1:LVC
3 in _  1
4 little _  _ 
5 doubt _  1
6 that _  _ 
7 the _  _ 
8 shadow _  2:ID
9 cast _  2
10 over _  _ 
11 the _  _ 
12 city _  _ 
13 by _  _ 
14 the _  _ 
15 attacks _  _ 
16 will _  _ 
17 enhance _  _ 
18 the _  _ 
19 chances _  _ 
20 of _  _ 
21 agreement nsp _ 
22 . _  _ 
    
# sent_id = 2
# text = Don't talk the talk...
1-2 Don't _  _ 
1 Do _   1:ID
2 not _   1
3 talk _   1
4 the _   1
5 talk _   1
6 if _   1
7 you _   1
8-9 can't _  _ 
8 can _   1
9 not _   1
10 walk _   1
11 the _   1
12 walk nsp  1
13 . _  _ 
    
1 Questioning _  _ 
2 colonial _  _ 
3 boundaries _  _ 
4 would _  _ 
5 open _  1:ID
6 a _  _ 
7 dangerous _  _ 
8 Pandora nsp 1
9 ' nsp 1
10 s _  1
11 box nsp 1
12 . _   _
 

 

If one VMWE is embedded in another one, a new semicolon-separated code is added to column 4. The identifiers of both VMWEs must be distinct. In case of VMWE overlapping, such as in coordination, the same rule applies.

Example of an annotated training corpus file:

1      Once _  _           
2 again _  _ 
3 it _  _ 
4 was _  _ 
5 a _  _ 
6 senior _  _ 
7 BBC _  _ 
8 person _  _ 
9 who _  _ 
10 let _  1:ID;2:VCP
11 the _  1
12 cat _  1
13 out _  1;2
14 of _  1
15 the _  1
16 bag nsp  1
17 . _  _ 
    
1 They _  _ 
2 were _  _ 
3 letting _  1:VPC;2:VPC
4 us _  _ 
5 in _  1
6 and _  _ 
7 out _   2
8 for _  _ 
9 quite _  _ 
10 some _  _ 
11 time nsp _ 
12 . _  _ 
 

 

Example of a blind corpus file:

1  Delegates _  _
2 are _  _
3 in _  _
4 little _  _ 
5 doubt _  _
6 that _  _ 
7 the _  _ 
8 shadow _  _
9 cast _  _
10 over _  _ 
11 the _  _ 
12 city _  _ 
13 by _  _ 
14 the _  _ 
15 attacks _  _ 
16 will _  _ 
17 enhance _  _ 
18 the _  _ 
19 chances _  _ 
20 of _  _ 
21 agreement nsp _ 
22 . _  _ 
    
# sent_id = 2
# text = Don't talk the talk...
1-2 Don't _  _ 
1 Do _  _
2 not _  _
3 talk _  _
4 the _  _
5 talk _  _
6 if _  _
7 you _  _
8-9 can't _  _ 
8 can _  _
9 not _  _
10 walk _  _
11 the _  _
12 walk nsp _
13 . _  _ 
    
1 Questioning _  _ 
2 colonial _  _ 
3 boundaries _  _ 
4 would _  _ 
5 open _  _
6 a _  _ 
7 dangerous _  _ 
8 Pandora nsp _
9 ' nsp _
10 s _  _
11 box nsp _
12 . _  _
 

 

Deprecated versions of the format

While developing the annotation guidelines and methodology for the PARSEME shared task other variants of the format were used:

The final parseme-tsv format is slightly redefined with respect to these previous formats in that:


1 While the format requires empty fields to be marked by underscores ('_'), as in the CoNLL-U format, missing underscores are tolerated when used on input of the FLAT annotation platform.