
Transfer Learning Based Cross-lingual Knowledge Extraction for ...
The method is based on the ?Bag of Words? (BOW) representation of documents, where each document is modeled as a vector with a dimension for each term of the ... 
Répondre aux questions visuelles à propos d'entités nommées
Articles from the source language Wikipedia are translated into the target language in advance and then transformed into training data TDS. In next section, we ... 
Edinburgh Research Explorer - Transfer Learning Based Cross ...
recherche cross-modale apprise de manière cross-modale. les titres des articles Wikipédia, qui sont également susceptibles de contenir la nature. 
Lindicle D2.1 Cross-lingual Infobox Alignment
tomatic method, which primarily consists of word labeling and feature vector generation, to generate the training data set TD = {(x, g(x))} from these. 
Design and Implementation of Wiki Content Transformations and ...
Articles from the source language Wikipedia are translated into the target lan- guage in advance and then transformed into training data TDS. In ... 
Towards a Wiki Interchange Format (WIF) - CEUR-WS.org
MediaWiki syntax allows authors to append or prepend text directly to the link to the effect that the pre- or postfix will be rendered as part of the link. 
How-To Wiki - iGEM
Compactness: The single-page-WIF needs to encode all structural information of a wiki, e. g. nested lists, headlines, tables, nested paragraphs, emphasised or ... 
Cross-domain Text Classification using Wikipedia
Abstract?Traditional approaches to document classification requires labeled data in order to construct reliable and accurate classifiers. 
Design and Implementation of the Sweble Wikitext Parser
It presents the de- sign and implementation of a parser for Wikitext, the wiki markup language of MediaWiki. We use parsing expres- sion grammars where most ... 
WIKIR: A Python Toolkit for Building a Large-scale Wikipedia-based ...
Abstract. Wikipedia is one of the most visited websites in the world and is also a frequent subject of scientific research. 
Utilising Wikipedia for Text Mining Applications - SciSpace
Model that uses both local (exact matching of n- grams of characters) and distributed (word embeddings) representations to compute a relevance score (Mitra ... 
GitHub Wiki Design and Implementation
It introduces the most relevant definitions and the related work for the research fields of semantic relatedness, named entity recog- nition, word sense ... 
Identifying Featured Articles in Spanish Wikipedia - SEDICI
... character set ... Word processors or HTML. Markdown was created by John Gruber in 2004 and is the default mechanism for docu- menting ... 
Statistical Measure of Quality in Wikipedia
An n-gram in turn is a substring of n tokens of t, where a token can be a character, a word, or a part- of-speech (POS) tag. The Term Frequency ? Inverse ... 
TS Wikipedia Corpus - LDC Catalog
ABSTRACT. Wikipedia is commonly viewed as the main online encyclope- dia. Its content quality, however, has often been questioned. 
The Currency of Wiki Articles ? A Language Model-based Approach
Une banque de gènes est un dispositif de conservation ex situ de matériel génétique, qu'il s'agisse de plantes ou d'animaux. Dans le cas des plantes, ... 
Impact, Characteristics, and Detection of Wikipedia Hoaxes
Wikis are ubiquitous in organisational and private use and provide a wealth of textual data. Maintaining the currency of this textual data is important and ... 
Edit Categories and Editor Role Identification in Wikipedia
We study hoaxes in the context of Wikipedia, for which there are two good reasons: first, anyone can insert information into Wiki- pedia by creating and editing ... 
Who Did What: Editor Role Identification in Wikipedia
This results in: Information Insertion (I-I), Information Deletion (ID), and Information Modification (I-M), Template Insertion (T-. I), Template Deletion (T-D) ... 
La négociation des contributions dans les wikis publics : légitimation ...
Template Deletion (T-D), and Template Modification (T-M), ... Specifically, we used two of the multi-label classifier implemented in Mulan (Tsoumakas, Katakis, ...