Showing posts with label semantic web. Show all posts
Showing posts with label semantic web. Show all posts

Monday, July 4, 2011

Sabbatical 2011-2012: Formalizing Scientific Discourse

Objective: The goal of my research program is to enable biologists to compose and evaluate scientific hypotheses using a diverse set of informational sources (ontology, database, text, equations, and web services). The purpose of my 2011-2012 sabbatical is to develop expertise in formalizing scientific discourse, with a particular focus on formalizing textual descriptions and mathematical equations such that they interoperate with knowledge represented in databases and structured documents. In particular, I am interested in using high quality facts derived from text and dynamic computation from formalized equations to answer questions and provide evidence for scientific hypotheses.

Background: Advancing knowledge in the life sciences involves experimentally testing hypotheses and interpreting the results based on prior scientific work. In generating a valid hypothesis, biologists face the overwhelming challenge of collecting, evaluating and integrating large and increasing amounts of different kinds of information about organisms, cells, genes and proteins from thousands of articles, hundreds of databases and dozens of tools. A biologist’s ability to efficiently construct and evaluate a hypothesis over current knowledge requires that i) knowledge, data and hypotheses are formally represented so they may be reasoned about, ii) adequate software exists to manage and query formal knowledge, and iii) data can be obtained by searching relevant databases and invoking the right analytical tools. The inability to efficiently discover relevant information can negatively impact scientific research directions and proposed activities. Methods for facilitating the construction and evaluation of hypotheses against the current state of knowledge could translate into greater scientific insight and increased productivity. Innovative approaches for knowledge discovery could be applied to data on the emerging Semantic Web and be transformative on a global scale.


Proposed Activities
1. Text to triples: The purpose of this leg of the sabbatical is to gain an understanding of the current state of the art in natural language processing and develop skill in producing high quality triples from parsing scientific text.
Time Frame: July 2011-September 2011
Location: European Bioinformatics Institute, Hinxton, UK.
Host: Dr. Rebholz-Schuhmann

2. Formalizing equations: The purpose of this leg of the sabbatical is to investigate the ontology of equations, represent scientific equations using Semantic Web technologies (principally the Rule Interchange Format), and implement semantic web services that serve to compute over formalized scientific equations.
Time Frame: October 2011-December 2011
Location: Universidad de Concepcion, Concepcion, Chile.
Host: Dr. Leo Ferres

3. Formalizing Research Hypotheses: The purpose of this leg of the sabbatical is to explore the formalization of hypotheses concerning disease. Specifically, I will extract meaningful facts from AlzForum and integrate these with resources from the National Centre for BioOntology (NCBO) and Bio2RDF, our large scale Semantic Web project.

Time Frame: January 2012-March 2012
Location: Stanford University, Palo Alto, California, USA.
Host: Dr. Mark Musen

4. Integrated Framework for Knowledge Discovery: The last leg of the sabbatical will be focused towards maximizing interoperability between text, equations, ontologies and database-derived facts. I will use SADI, our platform semantic web services framework, towards achieving this objective.
Time Frame: April 2012-June 2012
Location: India, Thailand, Singapore


Scientific Value and Broader Beneficial Impacts 
The development and application of efficient strategies for knowledge discovery is a major goal in bioinformatics. My research into new strategies for the representation and evaluation of scientific hypotheses using ontologies, scientific text, data and bioinformatic services will create a novel platform that will significantly contribute to scientific productivity and ultimately improve our understanding of biology. The proposed sabbatical will provide me with new skills that will be used to train a future training of young scientists in the areas of formal knowledge representation, text mining and the Semantic Web. Ultimately, it is expected that the sabbatical will cultivate new partnerships with leading scientists and open new doors to work with industry and government agencies.

Friday, May 22, 2009

Critique of OBO Foundry Principles

The OBO Foundry aims to create a suite of orthogonal interoperable reference ontologies in the biomedical domain. They have outlined their principles here: 

http://www.obofoundry.org/wiki/index.php/OBO_Foundry_Principles

In reading some of these I found that they poorly expressed true principles in ontology design. I provide here a brief critique on some of the contentious points:

"3. The ontologies possesses a unique identifier space within the OBO Foundry. The identifier uniquely and persistently identifies a definition, which itself unambiguous identifies some type of biological entity. The identifier is for the definition: it is NOT the name and it is NOT an identifier for the name.

There are systems that use alphanumeric id's - eg MetaCyc. This should be dis-encouraged, especially as these have semantic content."

This mixes up a number of issues. An identifier is a symbol for an entity, which should guarantee uniqueness in the lexical space, unlike human readable names which are not required to be unique. So it doesn’t matter whether the identifier is numeric, alphanumeric or alphabetic and thus the latter part of this principle, referring to alphanumeric MetaCyc ids, is pure nonsense. It is the description of the entity that *matters*, and that the textual description is arguably unchanging (What does OBOF say about when a description changes by even one word? Should a new identifier be crafted? How does one assess whether the previous identifier is in fact compatible with the new one? Should one be directed to use the new identifier – is it possible that the semantics are *fundamentally* different? These are far more important questions to address)

"6. The ontology must be orthogonal to other ontologies already lodged within OBO. For each domain, there should be convergence upon a single reference ontology that is recommended for use by those who wish to become involved with the Foundry initiative"

This is a contestable claim. Given that there is no universal agreement on many biological terms, any given ontology will not necessarily capture the semantics of what one wants to express. Anyone familiar with the word "gene" can easily demonstrate this as a case in point.

 

"7. The ontologies include textual definitions for all terms."

Textual descriptions aren't really useful unless they succinctly capture the essence of the entity in question.  For instance, definitions in the (OWL version) BFO are incomprehensible to many people (certainly to my undergrad students). In many other cases the textual descriptions can be shown to be either overly vague or constraining in unrealistic ways.

 

Einstein said "Make everything as simple as possible, but not simpler" - a good mantra in crafting term descriptions is "Be as accurate as possible, while not adding superfluous information or imposing unnecessary constraints.

 

"9. The ontology is well documented."

Be more specific - What does "well documented" mean?


"10. The ontology has a plurality of independent users."

This is another unreasonable demand. The defining characteristic is that for every ontology, there exists requirements (possibly in the form of use cases) that the ontology can be demonstrated to satisfy.  Ultimately, an ontology should have demonstrated utility. Paraphrasing Salinger - If you build it, (and have shown it to be useful) they will come.

 

"11. The ontology will be developed collaboratively with other OBO Foundry members."

A long standing myth is that ontologies need to be developed collaboratively - but in fact, we have found that such an approach is in fact wholly unproductive. What is productive is collecting use cases, undertaking focused development, and conducting a peer review and refinement process in which the needs of the community can be publicly solicited and addressed. This kind of procedure is in place at the W3C, and results in high quality standards. The OBO Foundry should consider setting up such a facility, with open calls for review across all relevant mailing lists, including quality assessment, additions/removals etc - particularly before things get published as a so-called "standard"