<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-08-14T17:47:00Z</responseDate>
  <request identifier="1508" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:1508</identifier>
        <datestamp>2024-03-06T11:08:01Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Bootstrapping an interactive information extraction system for FlyBase curation</dc:title>
          <dc:creator>Briscoe, Ted</dc:creator>
          <dc:creator>Gasperin, Caroline</dc:creator>
          <dc:creator>Lewin, Ian</dc:creator>
          <dc:creator>Vlachos, Andreas</dc:creator>
          <dc:subject>Biomedical Text Mining</dc:subject>
          <dc:subject>Interactive Information Extraction</dc:subject>
          <dc:subject>Natural Language Processing</dc:subject>
          <dc:description>We describe an adaptive information extraction (IE) system designed to&#13;
aid the curation of papers about fruit fly genomics for incorporation&#13;
into FlyBase. FlyBase employs a team of about eight curators who fill&#13;
in prespecified IE templetes (called proformas) for each gene and&#13;
allele discussed in a given paper with curatable information&#13;
associated with it. The normal approach to curation is to load the PDF&#13;
of the paper into a tool such as Acroread and to use the `Find'&#13;
function to search for repeated mentions of an entity of interest. The&#13;
relevant information is then typed into the appropriate template&#13;
fields. Templates are then checked for consistency and automatically&#13;
integrated into the database.&#13;
&#13;
We have developed PaperBrowser, a tool designed to make it easier for&#13;
curators to locate relevant information. The tool takes the PDF&#13;
version of the paper as input and rerenders it as SciXML, a standard&#13;
developed at Cambridge for representing the logical structure of&#13;
scientific articles in a fashion amenable to text mining. The basic&#13;
SciXML is augmented by a gene name recogniser and anaphora&#13;
resolution module so that PaperBrowser is able to highlight gene names&#13;
in the paper and to provide a navigation bar which allows the curator&#13;
to jump to specific mentions of a given gene in the various sections&#13;
of the paper. Alternatively, the curator can select a specific gene&#13;
mention and the browser will highlight all the noun phrases which are&#13;
anaphorically linked to that gene mention. These anaphoric links can&#13;
either be coreferential, or associative to the gene's products or&#13;
components, such as proteins or RNA.&#13;
&#13;
User-based evaluation of PaperBrowser in comparison to the use of&#13;
Acroread, with FlyBase curators undertaking the task of finding the&#13;
set of genes and alleles for which templates should be constructed,&#13;
has demonstrated that curation is 20\% faster at no cost to accuracy&#13;
when using PaperBrowser. PaperBrowser uses a conditional random field&#13;
model to perform gene name recognition bootstrapped from training data&#13;
derived automatically via information in FlyBase. The anaphora&#13;
resolution algorithm is unsupervised but uses information from the&#13;
Sequence Ontology augmented with lexemes from UMLS to identify noun&#13;
phrases referring to gene products and components. The PDF extraction&#13;
tool uses a commercial OCR package augmented with a seed-based machine&#13;
learning technique to learn the mapping from font and format&#13;
information to the logical structure of the paper. Papers describing&#13;
the complete processing pipeline, intrinsic evaluation of the&#13;
individual components and user-based experiments, along with test&#13;
datasets are available from the FlySlip Project website</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Ted Briscoe and Caroline Gasperin and Ian Lewin and Andreas Vlachos</dc:contributor>
          <dc:date>2008</dc:date>
          <dc:relation>Is Part Of Dagstuhl Seminar Proceedings, Volume 8131, Ontologies and Text Mining for Life Sciences : Current Status and Future Perspectives (2008)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/DagSemProc.08131.3</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-15086</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/DagSemProc.08131.3</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
