EC//EU

Project archive · 2012—2015

Open-source semantic text annotation

EntityClassifier

A multilingual system that found entities in free text, linked them to encyclopedic knowledge, and assigned meaningful classes at multiple levels of detail.

From document to linked knowledge

  1. 01SpotText spans
  2. 02LinkWikipedia · DBpedia
  3. 03ClassifyTargeted hypernyms
  4. 04EnrichDBpedia · YAGO2S
  5. 05PrioritizeEntity salience

01 / Method

A name was only the beginning.

Entityclassifier.eu combined named-entity recognition with semantic wikification. Its output described what a phrase referred to, what kind of thing it was, and how important it was to the document.

01 · Spot

Find candidate entities.

Lexico-syntactic patterns marked both named and common entities in free text, producing stable character spans for downstream processing.

European Commission → NamedEntity
research teams → CommonEntity
Brussels → NamedEntity

02 · Link

Resolve a phrase to an entity.

Wikipedia Search supplied candidate meanings. The most-frequent-sense strategy selected an article and mapped its URL to the corresponding DBpedia resource.

Brussels → Wikipedia
Wikipedia → dbpedia:Brussels

03 · Classify

Discover the “is-a” relationship.

Targeted Hypernym Discovery mined explanatory sentences in an entity’s Wikipedia article. Patterns such as “X is a Y” supplied a readable class.

“Marie Curie was a Polish and naturalised-French physicist and chemist...”

04 · Enrich

Join the knowledge graph.

Entities and hypernyms were cross-linked with DBpedia and enriched with DBpedia and YAGO2S ontology types. The Linked Hypernyms Dataset made these assignments reusable.

owl:Thingdbo:Persondbo:Scientistphysicist

05 · Prioritize

Estimate what mattered.

Later versions assigned entity salience so consumers could distinguish the document’s subject from incidental mentions.

Marie Curie0.92

Paris0.58

Sorbonne0.41

06 · Publish

Return interoperable annotations.

Results could be published through the RDF-based NLP Interchange Format, preserving document context, character offsets, entity links, and extracted types.

@prefix nif: <http://.../nif-core#> .

<#char=0,12>
  nif:anchorOf "Marie Curie" ;
  itsrdf:taIdentRef dbpedia:Marie_Curie .

02 / Languages

Three languages. One semantic layer.

ENEnglishDEGermanNLDutch

03 / Source

Four open components.

The archived Java code separated classification, service delivery, interface, and knowledge loading.

04 / Research record

Publications

The project connected information extraction, entity linking, hypernym discovery, linked data, and multilingual knowledge resources.

  1. 012013Milan Dojchinovski · Tomáš KliegrEntityclassifier.eu: Real-Time Classification of Entities in Text with WikipediaECML PKDD 2013 · LNCS 8190 · pp. 654–658
  2. 022012Milan Dojchinovski · Tomáš KliegrRecognizing, Classifying and Linking Entities with Wikipedia and DBpedia7th Workshop on Intelligent and Knowledge Oriented Technologies
  3. 032013M. Dojchinovski · I. Lašek · T. Kliegr · O. ZamazalWikipedia Search as Effective Entity Linking AlgorithmText Analysis Conference · TAC KBP 2013 · NIST
  4. 042014M. Dojchinovski · I. Lašek · O. Zamazal · T. KliegrEntityclassifier.eu and SemiTags: Entity Discovery, Linking and Classification with Wikipedia and DBpediaText Analysis Conference · TAC KBP 2014 · NIST
  5. 052014Tomáš Kliegr · Václav Zeman · Milan DojchinovskiLinked Hypernyms Dataset — Generation Framework and Use Cases3rd Workshop on Linked Data in Linguistics · LREC 2014

Archive note

The service is offline. The research remains linked.

Entityclassifier.eu is preserved as an open academic project by Milan Dojchinovski and Tomáš Kliegr. The GPL-3.0 source, datasets, and publications remain available.