Name: MultiCoNER Datasets
License: CC BY 4.0

Description

MultiCoNER 1 is a large multilingual dataset (11 languages) for Named Entity Recognition. It is designed to represent some of the contemporary challenges in NER, including low-context scenarios (short and uncased text), syntactically complex entities such as movie titles, and long-tail entity distributions. MultiCoNER 2 is a large multilingual dataset (12 languages) for fine grained Named Entity Recognition. Its fine-grained taxonomy contains 36 NE classes, representing real-world challenges for NER, where named entities, apart from the surface form, context represents a critical role in distinguishing between the different fine-grained types (e.g. Scientist vs. Athlete). Furthermore, the test data of MultiCoNER 2 contains noisy instances, where the noise has been applied to both context tokens as well as the entity tokens. The noise includes typing errors at character level based on keyboard layouts in the the different languages.

Update Frequency

N/A

License

CC BY 4.0

Documentation

https://multiconer.s3.us-west-2.amazonaws.com/readme.html

Managed By

See all datasets managed by Amazon.

Contact

besnikf@amazon.com

How to Cite

MultiCoNER Datasets was accessed on DATE from https://registry.opendata.aws/multiconer.

Usage Examples

Publications

Dynamic Gazetteer Integration in Multilingual Models for Cross-Lingual and Cross-Domain Named Entity Recognition by Besnik Fetahu, Anjie Fang, Oleg Rokhlenko and Shervin Malmasi
Gazetteer Enhanced Named Entity Recognition for Code-Mixed Web Queries by Besnik Fetahu, Anjie Fang, Oleg Rokhlenko and Shervin Malmasi
GEMNET: Effective Gated Gazetteer Representations for Recognizing Complex Entities in Low-context Input by Tao Meng, Anjie Fang, Oleg Rokhlenko and Shervin Malmasi
MultiCoNER: A Large-scale Multilingual Dataset for Complex Named Entity Recognition by Shervin Malmasi, Anjie Fang, Besnik Fetahu, Sudipta Kar, Oleg Rokhlenko

Resources on AWS

Description

MultiCoNER 1 Data files

Resource type

S3 Bucket

Amazon Resource Name (ARN)

arn:aws:s3:::multiconer/multiconer2022/

AWS Region

us-west-2

AWS CLI Access (No AWS account required)

aws s3 ls --no-sign-request s3://multiconer/multiconer2022/
Description

MultiCoNER 2 Data files

Resource type

S3 Bucket

Amazon Resource Name (ARN)

arn:aws:s3:::multiconer/multiconer2023/

AWS Region

us-west-2

AWS CLI Access (No AWS account required)

aws s3 ls --no-sign-request s3://multiconer/multiconer2023/