<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="EN">
<front>
<journal-meta><journal-id journal-id-type="nlm-ta">PLoS ONE</journal-id><journal-id journal-id-type="publisher-id">plos</journal-id><journal-id journal-id-type="pmc">plosone</journal-id><!--===== Grouping journal title elements =====--><journal-title-group><journal-title>PLoS ONE</journal-title></journal-title-group><issn pub-type="epub">1932-6203</issn><publisher>
<publisher-name>Public Library of Science</publisher-name>
<publisher-loc>San Francisco, USA</publisher-loc></publisher></journal-meta>
<article-meta><article-id pub-id-type="publisher-id">09-PONE-RA-12952</article-id><article-id pub-id-type="doi">10.1371/journal.pone.0008070</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="Discipline"><subject>Computational Biology</subject><subject>Computational Biology/Systems Biology</subject><subject>Mathematics/Algorithms</subject></subj-group></article-categories><title-group><article-title>Effective Identification of Conserved Pathways in Biological Networks Using Hidden Markov Models</article-title><alt-title alt-title-type="running-head">Pathway Alignment with HMMs</alt-title></title-group><contrib-group>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Qian</surname><given-names>Xiaoning</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Yoon</surname><given-names>Byung-Jun</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib>
</contrib-group><aff id="aff1"><label>1</label><addr-line>Department of Computer Science and Engineering, University of South Florida, Tampa, Florida, United States of America</addr-line>       </aff><aff id="aff2"><label>2</label><addr-line>Department of Electrical and Computer Engineering, Texas A&amp;M University, College Station, Texas, United States of America</addr-line>       </aff><contrib-group>
<contrib contrib-type="editor" xlink:type="simple"><name name-style="western"><surname>Di Bernardo</surname><given-names>Diego</given-names></name>
<role>Editor</role>
<xref ref-type="aff" rid="edit1"/></contrib>
</contrib-group><aff id="edit1">Fondazione Telethon, Italy</aff><author-notes>
<corresp id="cor1">* E-mail: <email xlink:type="simple">bjyoon@ece.tamu.edu</email></corresp>
<fn fn-type="con"><p>Conceived and designed the experiments: XQ BJY. Performed the experiments: XQ. Analyzed the data: XQ BJY. Wrote the paper: XQ BJY.</p></fn>
<fn fn-type="conflict"><p>The authors have declared that no competing interests exist.</p></fn></author-notes><pub-date pub-type="collection"><year>2009</year></pub-date><pub-date pub-type="epub"><day>7</day><month>12</month><year>2009</year></pub-date><volume>4</volume><issue>12</issue><elocation-id>e8070</elocation-id><history>
<date date-type="received"><day>17</day><month>9</month><year>2009</year></date>
<date date-type="accepted"><day>29</day><month>10</month><year>2009</year></date>
</history><!--===== Grouping copyright info into permissions =====--><permissions><copyright-year>2009</copyright-year><copyright-holder>Qian, Yoon</copyright-holder><license><license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</license-p></license></permissions><abstract><sec>
<title>Background</title>
<p>The advent of various high-throughput experimental techniques for measuring molecular interactions has enabled the systematic study of biological interactions on a global scale. Since biological processes are carried out by elaborate collaborations of numerous molecules that give rise to a complex network of molecular interactions, comparative analysis of these biological networks can bring important insights into the functional organization and regulatory mechanisms of biological systems.</p>
</sec><sec>
<title>Methodology/Principal Findings</title>
<p>In this paper, we present an effective framework for identifying common interaction patterns in the biological networks of different organisms based on hidden Markov models (HMMs). Given two or more networks, our method efficiently finds the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e001" xlink:type="simple"/></inline-formula> matching paths in the respective networks, where the matching paths may contain a flexible number of consecutive insertions and deletions.</p>
</sec><sec>
<title>Conclusions/Significance</title>
<p>Based on several protein-protein interaction (PPI) networks obtained from the Database of Interacting Proteins (DIP) and other public databases, we demonstrate that our method is able to detect biologically significant pathways that are conserved across different organisms. Our algorithm has a polynomial complexity that grows linearly with the size of the aligned paths. This enables the search for very long paths with more than 10 nodes within a few minutes on a desktop computer. The software program that implements this algorithm is available upon request from the authors.</p>
</sec></abstract><funding-group><funding-statement>This work was supported in part by the National Cancer Institute under 2 R25CA090301-06. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</funding-statement></funding-group><counts><page-count count="8"/></counts></article-meta>
</front>
<body><sec id="s1">
<title>Introduction</title>
<p>Recent advances in high-throughput experimental techniques for measuring molecular interactions <xref ref-type="bibr" rid="pone.0008070-Ito1">[1]</xref>–<xref ref-type="bibr" rid="pone.0008070-Krogan1">[4]</xref> have enabled the systematic study of biological interactions on a global scale for an increasing number of organisms <xref ref-type="bibr" rid="pone.0008070-vonMering1">[5]</xref>. Genome-scale interaction networks provide invaluable resources for investigating the functional organization of cells and understanding their regulatory mechanisms. Biological networks can be conveniently represented as graphs, in which the nodes represent the basic entities in a given network and the edges indicate the interactions between them. Network alignment provides an effective means for comparing the networks of different organisms by aligning these graphs and finding their common substructures. This can facilitate the discovery of conserved functional modules and ultimately help us study their functions and the detailed molecular mechanisms that contribute to these functions. For this reason, there have been growing efforts to develop efficient network alignment algorithms that can effectively detect conserved interaction patterns in various biological networks, including protein-protein interaction (PPI) networks <xref ref-type="bibr" rid="pone.0008070-Kelley1">[6]</xref>–<xref ref-type="bibr" rid="pone.0008070-Zaslavskiy1">[20]</xref>, metabolic networks <xref ref-type="bibr" rid="pone.0008070-Koyutrk1">[7]</xref>, <xref ref-type="bibr" rid="pone.0008070-Yang1">[12]</xref>, <xref ref-type="bibr" rid="pone.0008070-Pinter1">[21]</xref>, gene regulatory networks <xref ref-type="bibr" rid="pone.0008070-Akutsu1">[22]</xref>, and signal transduction networks <xref ref-type="bibr" rid="pone.0008070-Steffen1">[23]</xref>. It has been demonstrated that network alignment algorithms can detect many known biological pathways and also make statistically significant predictions of novel pathways.</p>
<p>Network alignment can be broadly divided into two categories, namely, <italic>global alignment</italic>, which tries to find the best coherent mapping between nodes in different networks that covers all nodes; and <italic>local alignment</italic>, which simply tries to detect significant common substructures in the given networks. Typically, the global network alignment problem is formulated as a graph matching problem whose goal is to find the optimal alignment that maximizes a global objective function that simultaneously measures the similarity between the constituent nodes and also between their interaction patterns. This optimization problem can be solved by a number of techniques, such as integer programming <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>, spectral clustering <xref ref-type="bibr" rid="pone.0008070-Singh1">[16]</xref>, <xref ref-type="bibr" rid="pone.0008070-Liao1">[17]</xref>, and message passing <xref ref-type="bibr" rid="pone.0008070-Zaslavskiy1">[20]</xref>. To cope with the high complexity of the global alignment problem, many algorithms incorporate heuristic techniques, such as greedy extension of high scoring subnetwork alignments and progressive construction of multiple network alignments <xref ref-type="bibr" rid="pone.0008070-Flannick1">[9]</xref>, <xref ref-type="bibr" rid="pone.0008070-Kalaev1">[15]</xref>, <xref ref-type="bibr" rid="pone.0008070-Liao1">[17]</xref>, <xref ref-type="bibr" rid="pone.0008070-Tian1">[19]</xref>.</p>
<p>There are also many local network alignment algorithms, where examples include PathBLAST <xref ref-type="bibr" rid="pone.0008070-Kelley1">[6]</xref>, NetworkBLAST <xref ref-type="bibr" rid="pone.0008070-Scott1">[10]</xref>, QPath <xref ref-type="bibr" rid="pone.0008070-Shlomi1">[11]</xref>, PathMatch and GraphMatch <xref ref-type="bibr" rid="pone.0008070-Yang1">[12]</xref>, just to name a few. These algorithms can effectively find conserved substructures with relatively small sizes, but many of them suffer from high computational complexity that makes it difficult to find larger substructures. Furthermore, many algorithms have limited flexibility of handling node insertions and deletions and/or rely on randomized heuristics that may not necessarily yield optimal results. In <xref ref-type="bibr" rid="pone.0008070-Qian1">[18]</xref>, we introduced an effective framework for local network alignment based on hidden Markov models (HMMs) that can effectively overcome many of these issues. The HMM framework can naturally integrate both the “node similarity” (typically estimated by sequence similarity) and the “interaction reliability” into the scoring scheme for comparing aligned paths, and it can deal with a large class of path isomorphism. Based on the HMM-based framework, we devised an efficient algorithm that can find the optimal homologous pathway for a given query pathway in a PPI network, whose complexity is linear with respect to the network size and the query length, making it applicable to search for long pathways. It was demonstrated that the algorithm can accurately detect homologous pathways that are biologically significant. However, the algorithm in <xref ref-type="bibr" rid="pone.0008070-Qian1">[18]</xref> was mainly developed for <italic>querying</italic> pathways in a target network, hence it cannot be directly used for local alignment of general networks.</p>
<p>In this paper, we extend the HMM-based framework proposed in <xref ref-type="bibr" rid="pone.0008070-Qian1">[18]</xref> to make it applicable for local alignment of general biological networks. Especially, we focus on the problem of identifying similar pathways that are conserved across two or more biological networks. Based on HMMs, we propose a general probabilistic framework for scoring pathway alignments and present an efficient search algorithm that can find the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e002" xlink:type="simple"/></inline-formula> alignments of homologous pathways with the highest scores. The algorithm has polynomial complexity which increases linearly with the length of the aligned pathways as well as the number of interactions in each network. The aligned pathways in a predicted alignment may contain flexible number of consecutive insertions and/or deletions. By combining the high-scoring pathway alignments that overlap with another, we can also detect conserved subnetworks with a general structure. Note that the algorithm can be also used for network querying, by designating one network as the query and another network as the target network.</p>
</sec><sec id="s2" sec-type="methods">
<title>Methods</title>
<p>In this section, we present an algorithm for solving the local network alignment problem based on HMMs. For simplicity, we first focus on the problem of aligning two networks, which can be formally defined as follows: Given two biological networks <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e003" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e004" xlink:type="simple"/></inline-formula> and a specified length <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e005" xlink:type="simple"/></inline-formula>, find the most similar pair <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e006" xlink:type="simple"/></inline-formula> of linear paths, where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e007" xlink:type="simple"/></inline-formula> belongs to the network <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e008" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e009" xlink:type="simple"/></inline-formula> belongs to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e010" xlink:type="simple"/></inline-formula>, and each of them have <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e011" xlink:type="simple"/></inline-formula> nodes. As we show later, the pairwise network alignment algorithm can be easily extended for aligning multiple networks in a straightforward manner.</p>
<sec id="s2a">
<title>Pairwise Network Alignment</title>
<p>Let <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e012" xlink:type="simple"/></inline-formula> be a graph representing a biological network. We assume that <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e013" xlink:type="simple"/></inline-formula> has a set <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e014" xlink:type="simple"/></inline-formula> of <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e015" xlink:type="simple"/></inline-formula> nodes, representing the entities in the network, and a set <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e016" xlink:type="simple"/></inline-formula> of <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e017" xlink:type="simple"/></inline-formula> edges, where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e018" xlink:type="simple"/></inline-formula> represents the interaction (binding or regulation) between <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e019" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e020" xlink:type="simple"/></inline-formula>. When the network <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e021" xlink:type="simple"/></inline-formula> is undirected, we assume that both <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e022" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e023" xlink:type="simple"/></inline-formula> are present in the set <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e024" xlink:type="simple"/></inline-formula> for simplicity. For example, when <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e025" xlink:type="simple"/></inline-formula> represents a PPI network, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e026" xlink:type="simple"/></inline-formula> corresponds to a protein, and the edge between <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e027" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e028" xlink:type="simple"/></inline-formula> indicates that these proteins can bind to each other. For a pair <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e029" xlink:type="simple"/></inline-formula> of interacting nodes such that <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e030" xlink:type="simple"/></inline-formula>, we define their interaction reliability as <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e031" xlink:type="simple"/></inline-formula>. Similarly, let <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e032" xlink:type="simple"/></inline-formula> be another graph with <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e033" xlink:type="simple"/></inline-formula> nodes and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e034" xlink:type="simple"/></inline-formula> edges, representing a different biological network. We denote the interaction reliability between two nodes <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e035" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e036" xlink:type="simple"/></inline-formula> in the graph <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e037" xlink:type="simple"/></inline-formula> as <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e038" xlink:type="simple"/></inline-formula>. Finally, we denote the similarity between two nodes <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e039" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e040" xlink:type="simple"/></inline-formula> in the respective networks as <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e041" xlink:type="simple"/></inline-formula>, which may be derived using the sequence similarity between two biological entities represented by two nodes as in our experiments.</p>
<p>Our goal is to find the best matching pair of paths <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e042" xlink:type="simple"/></inline-formula> (<inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e043" xlink:type="simple"/></inline-formula>) and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e044" xlink:type="simple"/></inline-formula> (<inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e045" xlink:type="simple"/></inline-formula>) in the respective networks that maximizes a predefined pathway alignment score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e046" xlink:type="simple"/></inline-formula>. In order to obtain meaningful results, the alignment score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e047" xlink:type="simple"/></inline-formula> should sensibly integrate the similarity score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e048" xlink:type="simple"/></inline-formula> between aligned nodes <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e049" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e050" xlink:type="simple"/></inline-formula> (<inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e051" xlink:type="simple"/></inline-formula>), the interaction reliability scores <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e052" xlink:type="simple"/></inline-formula> between <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e053" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e054" xlink:type="simple"/></inline-formula> (<inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e055" xlink:type="simple"/></inline-formula>) and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e056" xlink:type="simple"/></inline-formula> between <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e057" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e058" xlink:type="simple"/></inline-formula> (<inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e059" xlink:type="simple"/></inline-formula>), and the penalty for any gaps in the alignment.</p>
<p><xref ref-type="fig" rid="pone-0008070-g001">Figure 1C</xref> illustrates an example of an alignment between two similar paths <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e060" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e061" xlink:type="simple"/></inline-formula>, where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e062" xlink:type="simple"/></inline-formula> belongs to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e063" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e064" xlink:type="simple"/></inline-formula> belongs to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e065" xlink:type="simple"/></inline-formula> as shown in <xref ref-type="fig" rid="pone-0008070-g001">Fig. 1A</xref>. The dashed lines in <xref ref-type="fig" rid="pone-0008070-g001">Fig. 1A</xref> that connect two nodes <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e066" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e067" xlink:type="simple"/></inline-formula> indicate that there exist significant similarities between the connected nodes. In the example shown in <xref ref-type="fig" rid="pone-0008070-g001">Figure 1C</xref>, the optimal alignment that maximizes the alignment score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e068" xlink:type="simple"/></inline-formula> has two gaps at <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e069" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e070" xlink:type="simple"/></inline-formula>. Note that “insertions” and “deletions” are relative terms, and an insertion in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e071" xlink:type="simple"/></inline-formula> (e.g., <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e072" xlink:type="simple"/></inline-formula>) can be viewed as a deletion in the aligned path <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e073" xlink:type="simple"/></inline-formula>, and similarly, an insertion in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e074" xlink:type="simple"/></inline-formula> (e.g., <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e075" xlink:type="simple"/></inline-formula>) can be viewed as a deletion in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e076" xlink:type="simple"/></inline-formula>.</p>
<fig id="pone-0008070-g001" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0008070.g001</object-id><label>Figure 1</label><caption>
<title>Network representation and alignment.</title>
<p>(A) Example of two undirected biological networks <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e077" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e078" xlink:type="simple"/></inline-formula>. (B) A virtual path <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e079" xlink:type="simple"/></inline-formula> that corresponds to the alignment of best matching paths. (C) The top-scoring alignment between two similar paths <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e080" xlink:type="simple"/></inline-formula> (in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e081" xlink:type="simple"/></inline-formula>) and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e082" xlink:type="simple"/></inline-formula> (in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e083" xlink:type="simple"/></inline-formula>).</p>
</caption><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.g001" xlink:type="simple"/></fig></sec><sec id="s2b">
<title>Network Representation by HMM</title>
<p>To define the alignment score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e084" xlink:type="simple"/></inline-formula>, we adopt the hidden Markov model (HMM) formalism. We begin by constructing two HMMs based on the network graphs <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e085" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e086" xlink:type="simple"/></inline-formula>. Let us first focus on the construction of HMM for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e087" xlink:type="simple"/></inline-formula>. Each node <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e088" xlink:type="simple"/></inline-formula> in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e089" xlink:type="simple"/></inline-formula> corresponds to a hidden state in the HMM. For convenience, we represent this hidden state using the same notation <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e090" xlink:type="simple"/></inline-formula>. For each edge <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e091" xlink:type="simple"/></inline-formula> in the graph <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e092" xlink:type="simple"/></inline-formula>, we add an edge from state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e093" xlink:type="simple"/></inline-formula> to state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e094" xlink:type="simple"/></inline-formula> in the HMM. The resulting HMM has an identical structure as the network graph <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e095" xlink:type="simple"/></inline-formula>. The HMM for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e096" xlink:type="simple"/></inline-formula> can be constructed in a similar way. <xref ref-type="fig" rid="pone-0008070-g002">Figure 2A</xref> illustrates the HMMs that correspond to the network graphs shown in <xref ref-type="fig" rid="pone-0008070-g001">Fig. 1A</xref>. In order to find the best matching pairs of paths in the given networks, we define the concept of a “virtual” path <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e097" xlink:type="simple"/></inline-formula> that contains <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e098" xlink:type="simple"/></inline-formula> nodes, as shown in <xref ref-type="fig" rid="pone-0008070-g001">Fig. 1B</xref>. A node <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e099" xlink:type="simple"/></inline-formula> in the virtual path can be viewed as a symbol that is emitted by a pair of hidden states <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e100" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e101" xlink:type="simple"/></inline-formula> in the respective HMMs. From this point of view, the two HMMs can be regarded as generative models that <italic>jointly</italic> produce (or “emit”) the virtual path <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e102" xlink:type="simple"/></inline-formula>, and the underlying state sequence for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e103" xlink:type="simple"/></inline-formula> will be a pair of state sequences <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e104" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e105" xlink:type="simple"/></inline-formula> in the respective HMMs. Therefore, the concept of a virtual path can naturally couple a path in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e106" xlink:type="simple"/></inline-formula> with another in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e107" xlink:type="simple"/></inline-formula>, providing a convenient framework for identifying conserved pathways in the original biological networks.</p>
<fig id="pone-0008070-g002" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0008070.g002</object-id><label>Figure 2</label><caption>
<title>Hidden Markov models for network alignment.</title>
<p>(A) Ungapped hidden Markov models (HMMs) for finding the best matching pair of paths. The dots next to the hidden states represent all possible symbols corresponding to virtual nodes in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e108" xlink:type="simple"/></inline-formula> that can be emitted. (B) Modified HMMs that allow insertions and deletions. For simplicity, changes to the HMMs are shown only for the nodes <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e109" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e110" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e111" xlink:type="simple"/></inline-formula> in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e112" xlink:type="simple"/></inline-formula>; <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e113" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e114" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e115" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e116" xlink:type="simple"/></inline-formula> in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e117" xlink:type="simple"/></inline-formula>.</p>
</caption><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.g002" xlink:type="simple"/></fig>
<p>The described HMM-based network representation allows us to naturally integrate the interaction reliability scores and the node similarity scores into an effective probabilistic framework. We first define two mappings <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e118" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e119" xlink:type="simple"/></inline-formula>, which convert the interaction reliability scores <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e120" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e121" xlink:type="simple"/></inline-formula> between two nodes in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e122" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e123" xlink:type="simple"/></inline-formula> to the following transition probabilities<disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e124" xlink:type="simple"/><label>(1)</label></disp-formula><disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e125" xlink:type="simple"/><label>(2)</label></disp-formula>between the corresponding hidden states in the constructed HMMs. The mapping <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e126" xlink:type="simple"/></inline-formula> is defined so that (i) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e127" xlink:type="simple"/></inline-formula> for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e128" xlink:type="simple"/></inline-formula>, (ii) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e129" xlink:type="simple"/></inline-formula> for all <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e130" xlink:type="simple"/></inline-formula>, and (iii) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e131" xlink:type="simple"/></inline-formula> for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e132" xlink:type="simple"/></inline-formula>. Similarly, the mapping <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e133" xlink:type="simple"/></inline-formula> follows the same constraints: (i) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e134" xlink:type="simple"/></inline-formula> for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e135" xlink:type="simple"/></inline-formula>, (ii) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e136" xlink:type="simple"/></inline-formula> for all <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e137" xlink:type="simple"/></inline-formula>, and (iii) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e138" xlink:type="simple"/></inline-formula> for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e139" xlink:type="simple"/></inline-formula>. To specify the emission probability of a virtual symbol <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e140" xlink:type="simple"/></inline-formula> at a pair of hidden states in the two HMMs, we define another mapping <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e141" xlink:type="simple"/></inline-formula> that converts the node similarity score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e142" xlink:type="simple"/></inline-formula> to the following “pairing” probability<disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e143" xlink:type="simple"/><label>(3)</label></disp-formula>where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e144" xlink:type="simple"/></inline-formula> is the pair of underlying hidden states for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e145" xlink:type="simple"/></inline-formula>. The mapping <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e146" xlink:type="simple"/></inline-formula> is defined so that (i) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e147" xlink:type="simple"/></inline-formula> for all possible pairs of <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e148" xlink:type="simple"/></inline-formula>, and (ii) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e149" xlink:type="simple"/></inline-formula> for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e150" xlink:type="simple"/></inline-formula>.</p>
</sec><sec id="s2c">
<title>Ungapped Alignment</title>
<p>Based on the HMM framework, the problem of finding the best matching pair of paths is transformed into the problem of finding the optimal pair of state sequences in the two HMMs that jointly maximize the observation probability of the virtual path <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e151" xlink:type="simple"/></inline-formula>. In an ungapped pathway alignment, the underlying state pair <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e152" xlink:type="simple"/></inline-formula> of a virtual symbol <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e153" xlink:type="simple"/></inline-formula> directly corresponds to a pair of aligned nodes in the original networks. We can find the optimal pair of paths in polynomial time by using a dynamic programming algorithm defined in the following, which is conceptually similar to the Viterbi algorithm. We first define <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e154" xlink:type="simple"/></inline-formula> as the log-probability of the most probable pair of paths for a subsequence <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e155" xlink:type="simple"/></inline-formula> of length <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e156" xlink:type="simple"/></inline-formula>, where the underlying states for the virtual symbol <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e157" xlink:type="simple"/></inline-formula> are <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e158" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e159" xlink:type="simple"/></inline-formula>. The log-probability <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e160" xlink:type="simple"/></inline-formula> can be recursively computed as follows:<disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e161" xlink:type="simple"/><label>(4)</label></disp-formula></p>
<p>We repeat the above iterations until <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e162" xlink:type="simple"/></inline-formula>. At the end of the iterations, the maximum log-probability of the virtual path <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e163" xlink:type="simple"/></inline-formula> is given by:<disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e164" xlink:type="simple"/><label>(5)</label></disp-formula>where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e165" xlink:type="simple"/></inline-formula> is the optimal pair of state sequences that correspond to the best matching paths in the original biological networks. Once we have computed <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e166" xlink:type="simple"/></inline-formula>, it is straightforward to find <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e167" xlink:type="simple"/></inline-formula> by tracing the recursive equations that led to the maximum log-probability <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e168" xlink:type="simple"/></inline-formula>. Although the above algorithm only finds the top-scoring pair of paths, we can easily extend it to find the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e169" xlink:type="simple"/></inline-formula> pairs simply by replacing the <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e170" xlink:type="simple"/></inline-formula> operator by an operator that finds the <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e171" xlink:type="simple"/></inline-formula> largest scores.</p>
<p>The computational complexity of the above algorithm is <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e172" xlink:type="simple"/></inline-formula> for finding the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e173" xlink:type="simple"/></inline-formula> pairs of matching paths, where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e174" xlink:type="simple"/></inline-formula> is the length of the aligned paths that we want to find, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e175" xlink:type="simple"/></inline-formula> is the number of edges in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e176" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e177" xlink:type="simple"/></inline-formula> is the number of edges in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e178" xlink:type="simple"/></inline-formula>. Note that the complexity is linear with respect to all the parameters <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e179" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e180" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e181" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e182" xlink:type="simple"/></inline-formula>.</p>
<p>The log-probability <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e183" xlink:type="simple"/></inline-formula> can serve as a good alignment score for the paths <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e184" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e185" xlink:type="simple"/></inline-formula> that effectively combines node similarity and interaction reliability. In principle, we can also use non-stochastic emission (pairing) scores <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e186" xlink:type="simple"/></inline-formula> and transition scores <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e187" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e188" xlink:type="simple"/></inline-formula> in the recursive equation (4), in place of the log-probabilities <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e189" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e190" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e191" xlink:type="simple"/></inline-formula>, respectively. This will yield a non-stochastic pathway alignment score instead of an observation probability.</p>
<p>As we can see, the concept of the “virtual” path provides an intuitive way of coupling states in two different HMMs. In fact, by taking a closer look at the recursive equation (4), the proposed alignment algorithm can also be viewed as a Markovian walk on a product graph, whose nodes consist of all possible pairs of hidden states in the respective HMMs and the edges between these nodes are determined by the connectivity (or transition probability) between the corresponding states in the HMMs. The algorithm searches for the optimal path (or the top-<inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e192" xlink:type="simple"/></inline-formula> paths) in the product graph that yields the highest score based on the parameters of the given HMMs.</p>
</sec><sec id="s2d">
<title>Alignment with Gaps</title>
<p>To accommodate gaps in the aligned paths <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e193" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e194" xlink:type="simple"/></inline-formula>, we modify the previous HMMs as follows. First, we add an accompanying state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e195" xlink:type="simple"/></inline-formula> for every state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e196" xlink:type="simple"/></inline-formula> in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e197" xlink:type="simple"/></inline-formula>, and similarly, we add an accompanying state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e198" xlink:type="simple"/></inline-formula> for every state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e199" xlink:type="simple"/></inline-formula> in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e200" xlink:type="simple"/></inline-formula>. Next, we add an outgoing edge from each state to the corresponding accompanying state. In addition to this, we also add outgoing edges from the accompanying state to all the neighboring states of the original state. To be more precise, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e201" xlink:type="simple"/></inline-formula> will have an outgoing edge to every <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e202" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e203" xlink:type="simple"/></inline-formula> will have an outgoing edge to every <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e204" xlink:type="simple"/></inline-formula>. By varying the transition probabilities <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e205" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e206" xlink:type="simple"/></inline-formula>, we can control the probabilities of having insertions and/or deletions, and thereby control the “gap penalties” in a pathway alignment. We adjust the outgoing transition probability from <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e207" xlink:type="simple"/></inline-formula> so that <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e208" xlink:type="simple"/></inline-formula>; and for the outgoing transition probability from <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e209" xlink:type="simple"/></inline-formula> so that <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e210" xlink:type="simple"/></inline-formula>. We can also control the probabilities of having <italic>consecutive</italic> insertions or deletions by adjusting the probabilities <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e211" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e212" xlink:type="simple"/></inline-formula> for making self-transitions at either <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e213" xlink:type="simple"/></inline-formula> or <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e214" xlink:type="simple"/></inline-formula>. The outgoing transition probabilities <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e215" xlink:type="simple"/></inline-formula> from an accompanying state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e216" xlink:type="simple"/></inline-formula> are chosen so that they are proportional to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e217" xlink:type="simple"/></inline-formula> and satisfy <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e218" xlink:type="simple"/></inline-formula>. The transition probabilities in <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e219" xlink:type="simple"/></inline-formula> can be chosen in a similar manner. The structures of the modified HMMs are depicted in <xref ref-type="fig" rid="pone-0008070-g002">Fig. 2B</xref>. Note that, in a gapped alignment, the matching paths (or state sequences) <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e220" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e221" xlink:type="simple"/></inline-formula> will still contain <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e222" xlink:type="simple"/></inline-formula> nodes each, and the only difference from an ungapped alignment is that the paths may now contain one or more accompanying nodes which represent gaps. The proposed framework does not impose any restriction on the number of gaps and their locations in the pathway alignment.</p>
<p>In order to find the optimal pair of paths (and their alignment) that maximize the pathway alignment score, we can apply the same dynamic programming algorithm described in the previous section. The retrieved paths can contain any of the hidden states <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e223" xlink:type="simple"/></inline-formula> <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e224" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e225" xlink:type="simple"/></inline-formula> <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e226" xlink:type="simple"/></inline-formula> in the modified HMMs, where we define <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e227" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e228" xlink:type="simple"/></inline-formula> for notational convenience. The optimal paths <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e229" xlink:type="simple"/></inline-formula> is the best matching pair of paths from two networks, and they may now contain insertions and/or deletions. As before, if we want to find the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e230" xlink:type="simple"/></inline-formula> pairs instead of a single top-scoring pair, we can simply replace the <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e231" xlink:type="simple"/></inline-formula> operator by an operator that finds the <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e232" xlink:type="simple"/></inline-formula> largest scores. Note that the computational complexity of the algorithm is <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e233" xlink:type="simple"/></inline-formula>, which is still linear with respect to all the parameters.</p>
</sec><sec id="s2e">
<title>Extension to Multiple Networks</title>
<p>It is straightforward to extend the described pairwise network alignment algorithm for aligning multiple networks. Without loss of generality, we only consider the extension to the alignment of three networks. Given three network graphs <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e234" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e235" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e236" xlink:type="simple"/></inline-formula>, we construct the corresponding HMMs based on their structures. We again use the concept of virtual paths, and now we assume that a virtual path <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e237" xlink:type="simple"/></inline-formula> is jointly emitted by these three HMMs. The emission of a virtual symbol <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e238" xlink:type="simple"/></inline-formula> is now governed by a pairing probability <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e239" xlink:type="simple"/></inline-formula> of three hidden states <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e240" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e241" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e242" xlink:type="simple"/></inline-formula> that belong to the HMMs that correspond to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e243" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e244" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e245" xlink:type="simple"/></inline-formula>, respectively. We can find the best matching paths based on the following recursive equation:<disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e246" xlink:type="simple"/><label>(6)</label></disp-formula>where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e247" xlink:type="simple"/></inline-formula> is assumed for simplicity. We repeat the above iterations until we reach <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e248" xlink:type="simple"/></inline-formula> and compute the maximum log-probability as follows:<disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e249" xlink:type="simple"/><label>(7)</label></disp-formula>where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e250" xlink:type="simple"/></inline-formula> corresponds to the set of best matching paths in the three networks.</p>
</sec><sec id="s2f">
<title>Implementation of the Alignment Algorithm</title>
<p>It should be noted that although we fix the length of the virtual path to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e251" xlink:type="simple"/></inline-formula>, we can in fact find any top-scoring alignment with a shorter length <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e252" xlink:type="simple"/></inline-formula>, since we store all the alignment scores for shorter alignments while running the dynamic programming algorithm. The recursive equations in (4) and (6) do not restrict multiple occurrence of the same node in the final pathway alignment. However, when it is desirable to avoid such multiple occurrence, we can easily incorporate a “look-back” step into each iteration in order to prevent adding a node that is already included in the (intermediate) alignment. As this requires tracing the intermediate optimal (or top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e253" xlink:type="simple"/></inline-formula>) alignment, the computational complexity of the recursive equations (4) and (6) with a “look-back” step will be increased in proportion to the length of the intermediate alignment.</p>
<p>In order to obtain more general subnetwork alignments, not just alignments of linear paths, we can combine the overlapping paths among the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e254" xlink:type="simple"/></inline-formula> retrieved pairs of paths. The edges that are already contained in the constructed subnetwork alignment (which correspond to the conserved molecular interactions in the biological networks) are then removed from the HMMs, and we run the dynamic programming algorithm again to find another subnetwork alignment that does not overlap with the retrieved subnetworks. By repeating this “search and peel-off” process, we can effectively find diverse subnetwork regions that are conserved in the given networks.</p>
<p>The memory complexity of the proposed algorithm is <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e255" xlink:type="simple"/></inline-formula> for finding the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e256" xlink:type="simple"/></inline-formula> pathway alignments for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e257" xlink:type="simple"/></inline-formula> networks. Although the required amount of memory increases only linearly with respect to each parameter, it can still make the algorithm infeasible when we want to align multiple number of large networks. To overcome this problem, we may assign non-zero pairing probabilities <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e258" xlink:type="simple"/></inline-formula> to a set of nodes (in the respective networks) only if every pair in this set has considerable node similarity that exceeds a certain threshold. Assuming that there are <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e259" xlink:type="simple"/></inline-formula> sets of nodes that satisfy this condition, we only need to consider these <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e260" xlink:type="simple"/></inline-formula> possible node alignments, in which case the overall memory complexity reduces to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e261" xlink:type="simple"/></inline-formula>. Since <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e262" xlink:type="simple"/></inline-formula> is often much smaller than <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e263" xlink:type="simple"/></inline-formula>, this scheme can save significant amount of memory, thereby making the algorithm feasible.</p>
</sec></sec><sec id="s3">
<title>Results</title>
<p>To demonstrate the effectiveness of the HMM-based network alignment algorithm, we carried out the following experiments. First, we used our algorithm to align two pairs of small synthetic networks that were used to validate the network alignment algorithm proposed in <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>. Second, we used the proposed algorithm for finding putative pathways in the fruit fly PPI network that look similar to known human pathways. Finally, we applied the algorithm for aligning microbial PPI networks to assess its ability to find conserved functional modules.</p>
<sec id="s3a">
<title>Aligning Synthetic Networks</title>
<p>To illustrate the potential capability of aligning different types of molecular networks, we first tested our algorithm using two small synthetic examples, which include a pair of undirected networks and another pair of directed networks. These examples were obtained from the tutorial files in the PathBLAST plugin of software Cytopscape (version 1.1, <ext-link ext-link-type="uri" xlink:href="http://www.cytoscape.org/plugins1.php" xlink:type="simple">http://www.cytoscape.org/plugins1.php</ext-link>) and they were used for the validation of a network alignment algorithm called MNAligner <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>.</p>
<sec id="s3a1">
<title>HMM parameterization</title>
<p>For aligning the synthetic networks, we parameterized the HMMs as follows. We set the transition scores <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e264" xlink:type="simple"/></inline-formula> directly based on the “adjacent matrices” given in <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>, which contain the interaction scores between two nodes in the respective networks. Every interaction score takes a value between 0 and 1, hence we can view it as the “interaction probability”. We took the logarithm of this interaction probability as the transition score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e265" xlink:type="simple"/></inline-formula>. When there is no interaction between two nodes, we have <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e266" xlink:type="simple"/></inline-formula>. This keeps the HMM from making a direct transition from a state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e267" xlink:type="simple"/></inline-formula> to a non-relevant state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e268" xlink:type="simple"/></inline-formula>, thereby preventing the inclusion of irrelevant protein interactions that do not have any biological support in the network. Similarly, we obtained the emission scores <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e269" xlink:type="simple"/></inline-formula> by taking the logarithm of the similarity scores between nodes given by the “similarity matrices” in <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>. The adjacent matrices and the similarity scores for the two examples can be found in the <xref ref-type="supplementary-material" rid="pone.0008070.s001">Supporting Information S1</xref>.</p>
</sec><sec id="s3a2">
<title>Example 1: Aligning undirected networks</title>
<p>We first used our algorithm for aligning a pair of undirected networks. To compare the alignment results with the results obtained by MNAligner <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>, we looked for the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e270" xlink:type="simple"/></inline-formula> alignments without gaps, where the length of the virtual path was set to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e271" xlink:type="simple"/></inline-formula>. By incorporating “look-back” steps into our dynamic programming algorithm, we restricted the multiple occurrence of the same node pair in the obtained pathway alignment. The top-scoring pathway alignment obtained from our algorithm was <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e272" xlink:type="simple"/></inline-formula>, which is identical to the optimal alignment identified by both PathBLAST <xref ref-type="bibr" rid="pone.0008070-Kelley1">[6]</xref> and MNAligner <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>. Unlike PathBLAST, the proposed HMM-based algorithm and the MNAligner both keep the natural order of the nodes in the original networks. We also noticed that the paths <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e273" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e274" xlink:type="simple"/></inline-formula> can be aligned with several other potential similar paths in the corresponding networks from the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e275" xlink:type="simple"/></inline-formula> aligned results. After removing the interactions included in the top-scoring alignment, we searched for the next top-scoring alignment. This returned the alignment <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e276" xlink:type="simple"/></inline-formula>, which was also ranked as the second best alignment by MNAligner <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>. Repeating the experiment after removing this alignment returned <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e277" xlink:type="simple"/></inline-formula> as the third best alignment. This is different from the alignment <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e278" xlink:type="simple"/></inline-formula> that was found by MNAligner, which got a lower score in our experiment. We noted that the alignment <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e279" xlink:type="simple"/></inline-formula> is not as significant as the three alignments that we found, as <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e280" xlink:type="simple"/></inline-formula> can be aligned with many other paths with the same alignment score. By repeating the above experiments and combining the pathway alignment results, we obtained the global network alignment illustrated in <xref ref-type="fig" rid="pone-0008070-g003">Fig. 3A</xref>, where a bold line represents that the corresponding edges in the respective networks are matched, whereas a thin line indicates a mismatch. These results show that the HMM-based method can effectively identify the top matching paths in different undirected networks, and it yields better results with higher alignment scores integrating both node similarity and interaction probability compared to PathBLAST and MNAligner for this purpose.</p>
<fig id="pone-0008070-g003" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0008070.g003</object-id><label>Figure 3</label><caption>
<title>The alignment results for synthetic networks.</title>
<p>(A) Undirected networks; (B) Directed networks.</p>
</caption><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.g003" xlink:type="simple"/></fig></sec><sec id="s3a3">
<title>Example 2: Aligning directed networks</title>
<p>Without any modification, our algorithm can also be used for aligning directed networks. We demonstrate this by using the second example that contains a pair of small directed networks. In this experiment, we set the length of the virtual path to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e281" xlink:type="simple"/></inline-formula>, which is the length of the longest path in these two networks. As there are fewer legitimate paths in these networks, we only looked for the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e282" xlink:type="simple"/></inline-formula> aligned pairs of paths. The obtained pathway alignments were combined to get the global network alignment shown in <xref ref-type="fig" rid="pone-0008070-g003">Fig. 3B</xref>. The alignment results were similar to those obtained by MNAligner <xref ref-type="bibr" rid="pone.0008070-Li1">[24]</xref>, except that we found fewer aligned nodes and edges. This is natural since there exist only a few similar pairs of nodes in the given networks (see <xref ref-type="supplementary-material" rid="pone.0008070.s001">Supporting Information S1</xref>) and as our algorithm focuses on finding the best local alignments instead of a global alignment. Note that, unlike PathBLAST, which finds path alignments based on several heuristics, the proposed algorithm can find the mathematically optimal path alignment for the given networks.</p>
</sec></sec><sec id="s3b">
<title>Aligning Annotated Pathways with PPI Networks</title>
<sec id="s3b1">
<title>HMM parameterization</title>
<p>The proposed algorithm can also be used for identifying putative pathways in a new biological network, which look similar to known pathways. To demonstrate this, we used our algorithm to search for human signaling pathways in the fruit fly PPI network. In order to compare the search results with those of the network querying algorithm in <xref ref-type="bibr" rid="pone.0008070-Qian1">[18]</xref>, the HMMs were parameterized according to the non-stochastic scoring scheme in <xref ref-type="bibr" rid="pone.0008070-Qian1">[18]</xref> as we describe in the following. The transition score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e283" xlink:type="simple"/></inline-formula> was set to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e284" xlink:type="simple"/></inline-formula> in the presence of interaction between the proteins that correspond to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e285" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e286" xlink:type="simple"/></inline-formula>, and it was set to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e287" xlink:type="simple"/></inline-formula> in the absence of any interaction. To allow gaps in alignments, the transition score from a state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e288" xlink:type="simple"/></inline-formula> to its accompanying state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e289" xlink:type="simple"/></inline-formula> was set to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e290" xlink:type="simple"/></inline-formula>, and we set the self-transition score at <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e291" xlink:type="simple"/></inline-formula> to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e292" xlink:type="simple"/></inline-formula> to allow consecutive gaps. Furthermore, the score for making a transition from <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e293" xlink:type="simple"/></inline-formula> to a regular state <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e294" xlink:type="simple"/></inline-formula> was set to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e295" xlink:type="simple"/></inline-formula> for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e296" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e297" xlink:type="simple"/></inline-formula> for <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e298" xlink:type="simple"/></inline-formula>. The emission score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e299" xlink:type="simple"/></inline-formula> for two proteins <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e300" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e301" xlink:type="simple"/></inline-formula> in different networks (where the query network is simply a linear path in this case) was computed based on their sequence similarity. For each protein pair <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e302" xlink:type="simple"/></inline-formula>, we computed its E-value using the PRSS routine in the FASTA package <xref ref-type="bibr" rid="pone.0008070-Pearson1">[25]</xref>, <xref ref-type="bibr" rid="pone.0008070-Pearson2">[26]</xref>, which is known to yield more accurate E-values compared to BLASTP <xref ref-type="bibr" rid="pone.0008070-Pagni1">[27]</xref>. We regarded a protein pair <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e303" xlink:type="simple"/></inline-formula> as a “match” if its E-value <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e304" xlink:type="simple"/></inline-formula> was below a threshold <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e305" xlink:type="simple"/></inline-formula>. Otherwise, we regarded the pair as a “mismatch”, which implies that the proteins do not bear significant similarity. Based on this criterion, we set the emission score <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e306" xlink:type="simple"/></inline-formula> as follows:<disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e307" xlink:type="simple"/><label>(8)</label></disp-formula>The value <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e308" xlink:type="simple"/></inline-formula> can be viewed as the mismatch penalty, and is selected so that <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e309" xlink:type="simple"/></inline-formula>. We set the insertion and deletion penalty also to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e310" xlink:type="simple"/></inline-formula>. Finally, since two accompanying states cannot be paired with each other, we set <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e311" xlink:type="simple"/></inline-formula>.</p>
</sec><sec id="s3b2">
<title>Querying human pathways in the fruit fly PPI network</title>
<p>We first obtained the PPI network of <italic>Drosophila melanogaster</italic> from the Database of Interacting Proteins (DIP) <xref ref-type="bibr" rid="pone.0008070-Xenarios1">[28]</xref> and constructed the “target HMM”. Then we constructed a “query HMM” for the human hedgehog signaling pathway and another query HMM based on the human MAP kinase pathway. When constructing the query HMMs, we regarded each signaling pathway as a “directed network” with a linear structure, instead of a “sequence of proteins” as in <xref ref-type="bibr" rid="pone.0008070-Qian1">[18]</xref>. The similarity threshold was set to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e312" xlink:type="simple"/></inline-formula> and the gap penalty was set to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e313" xlink:type="simple"/></inline-formula>, as in <xref ref-type="bibr" rid="pone.0008070-Qian1">[18]</xref>. The constructed query HMMs were then used to search for matching paths in the target HMM. Despite the generality and the different implementation of the proposed algorithm, the top pathways retrieved by the proposed algorithm agree with the predictions in <xref ref-type="bibr" rid="pone.0008070-Qian1">[18]</xref>, which is the direct consequence of the mathematical optimality of both methods. For the human hedgehog signaling pathway lhh–Ptch–Smo–Stk36–Gli, the top-scoring pathway in the <italic>D. melanogaster</italic> network agreed well with the putative <italic>D. melanogaster</italic> hedgehog signaling pathway reported in the KEGG database <xref ref-type="bibr" rid="pone.0008070-Kanehisa1">[29]</xref>. In fact, the best aligned path in the fruit fly network contained shh–ptc–Smo–fu–ci, which is identical to the core portion of the putative fly hedgehog signaling pathway (<ext-link ext-link-type="uri" xlink:href="http://www.genome.jp/dbget-bin/get_pathway?org_name=dme&amp;mapno=04340" xlink:type="simple">http://www.genome.jp/dbget-bin/get_pathway?org_name=dme&amp;mapno=04340</ext-link>) in the KEGG database <xref ref-type="bibr" rid="pone.0008070-Kanehisa1">[29]</xref>. The query result of the human MAP kinase pathway Egfr–drk–Sos–Ras85D–ph1–Mekk1–ERKA was also biologically significant, and the seven proteins in the retrieved pathway matched exactly with the proteins in the putative fruit fly MAP kinase pathway (<ext-link ext-link-type="uri" xlink:href="http://www.genome.jp/dbget-bin/get_pathway?org_name=map&amp;mapno=04010" xlink:type="simple">http://www.genome.jp/dbget-bin/get_pathway?org_name=map&amp;mapno=04010</ext-link>) reported in KEGG. These results compare favorably to the results obtained by one of the state-of-the-art algorithms <xref ref-type="bibr" rid="pone.0008070-Shlomi1">[11]</xref>, where they found two identical proteins in the putative fly hedgehog signaling pathway and five proteins in the putative fly MAPK pathway.</p>
</sec></sec><sec id="s3c">
<title>Aligning Microbial PPI Networks</title>
<p>In order to validate the accuracy of our algorithm for predicting functional modules that are conserved in different organisms, we performed additional experiments using three microbial PPI networks obtained from <xref ref-type="bibr" rid="pone.0008070-Flannick1">[9]</xref>. In our experiments, we performed a pairwise alignment between the <italic>E. coli</italic> and the <italic>C. crescentus</italic> networks as well as a pairwise alignment between the <italic>E. coli</italic> and the <italic>S. typhimurium</italic> networks. We assessed the accuracy of our algorithm based on the consistency of the <italic>KEGG ortholog (KO) group</italic> annotations <xref ref-type="bibr" rid="pone.0008070-Kanehisa1">[29]</xref> of the aligned proteins. In order to measure the consistency of KO group annotations, we computed the specificity of the predictions based on a similar methodology that was used in <xref ref-type="bibr" rid="pone.0008070-Flannick2">[14]</xref>. To compute this measure, we first remove all the aligned protein pairs that do not have complete KO annotations, and then compute the total number of annotated protein pairs. An annotated protein pair is regarded as being <italic>correct</italic> if both proteins have the same KO group annotations, and <italic>incorrect</italic> if the annotations do not agree. The specificity is defined as the ratio of the number of “correct” protein pairs among all annotated protein pairs.</p>
<p>For this experiment, the parameters of the HMMs have been chosen as follows. First, the transition scores <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e314" xlink:type="simple"/></inline-formula> have been obtained by taking the logarithm of the protein interaction probabilities in the microbial networks, which had been assigned by the SRINI algorithm <xref ref-type="bibr" rid="pone.0008070-Srinivasan1">[30]</xref>. The emission scores <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e315" xlink:type="simple"/></inline-formula> have been computed based on the sequence similarity between the proteins <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e316" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e317" xlink:type="simple"/></inline-formula>, as in the previous section, where the protein similarities have been estimated based on the BLASTP hit scores between protein pairs provided in <xref ref-type="bibr" rid="pone.0008070-Flannick1">[9]</xref>.</p>
<p>Based on the constructed HMMs, we used our algorithm to find the top-scoring pathway alignment with gaps. At each iteration, we looked for the top aligned pair of paths, stored the alignment, and removed the interactions included in the alignment from the respective networks for the next iteration. By repeating this iteration, we found 200 high-scoring path alignments. This experiment has been repeated with varying virtual path length: <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e318" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e319" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e320" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e321" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e322" xlink:type="simple"/></inline-formula>. In all our experiments, we disallowed multiple occurrence of identical protein pairs and set the gap/mismatch penalty to <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e323" xlink:type="simple"/></inline-formula>. For each experiment, we computed the cumulative specificity for the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e324" xlink:type="simple"/></inline-formula> alignments, which is given by<disp-formula><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.e325" xlink:type="simple"/><label>(9)</label></disp-formula>where <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e326" xlink:type="simple"/></inline-formula> is the total number of correctly aligned protein pairs in the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e327" xlink:type="simple"/></inline-formula> alignments, and <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e328" xlink:type="simple"/></inline-formula> is the total number of annotated protein pairs also in the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e329" xlink:type="simple"/></inline-formula> alignments. The result from the pairwise alignment of the <italic>E. coli</italic> and the <italic>C. crescentus</italic> networks is shown in <xref ref-type="fig" rid="pone-0008070-g004">Fig. 4A</xref>, and the result from the alignment of the <italic>E. coli</italic> and the <italic>S. typhimurium</italic> networks is shown in <xref ref-type="fig" rid="pone-0008070-g004">Fig. 4B</xref>. As we can see in both <xref ref-type="fig" rid="pone-0008070-g004">Fig. 4A</xref> and <xref ref-type="fig" rid="pone-0008070-g004">Fig. 4B</xref>, the cumulative specificity <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e330" xlink:type="simple"/></inline-formula> generally decreases when we increase the alignment length <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e331" xlink:type="simple"/></inline-formula>. This is expected since the algorithm tends to recruit more protein pairs in the alignment if we increase <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e332" xlink:type="simple"/></inline-formula>. Furthermore, <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e333" xlink:type="simple"/></inline-formula> generally decreases if we increase <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e334" xlink:type="simple"/></inline-formula>. This is natural, since alignments with lower scores correspond to less conserved pathways with larger variations. Although it is difficult to directly compare our results with those reported in <xref ref-type="bibr" rid="pone.0008070-Flannick2">[14]</xref>, it is still worth to note that the cumulative specificity (for the top 200 alignments) of the proposed HMM-based algorithm is higher than the specificity of the alignment algorithm Græmlin 2.0 <xref ref-type="bibr" rid="pone.0008070-Flannick2">[14]</xref>, for both pairwise network alignments. These results clearly indicate that our HMM-based algorithm can produce accurate network alignments that are biologically meaningful.</p>
<fig id="pone-0008070-g004" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0008070.g004</object-id><label>Figure 4</label><caption>
<title>Functional specificity for microbial network alignment.</title>
<p>The cumulative specificity of the top <inline-formula><inline-graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0008070.e335" xlink:type="simple"/></inline-formula> aligned pathways obtained from (A) the pairwise alignment between <italic>E. coli</italic> and <italic>C. crescentus</italic> networks; and (B) the pairwise alignment between <italic>E. coli</italic> and <italic>S. typhimurium</italic> networks.</p>
</caption><graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.g004" xlink:type="simple"/></fig>
<p>Further analysis of the predicted alignments led to a number of interesting observations. For example, the alignment of <italic>E. coli</italic> and <italic>C. crescentus</italic> networks and the alignment of <italic>E. coli</italic> and <italic>S. typhimurium</italic> networks both detected conserved DNA replication modules. The module contained components of the primosome (dnaA, gyrA, gyrB), subunits of topoisomerase IV (parC, parE), and a subunit of DNA polymerase III (dnaN). These protein families are all known to be involved in DNA replication. We also found other interesting conserved modules, which include both large and small subunits of ribosomal protein complexes (rplA, rplB, rplC, rplE, rplK, rplP; and rpsA, rpsB, rpsC, rpsE, rpsG, rpsK); DNA-directed RNA polymerase complex containing rpoA, rpoB, rpoC, and other subunits; the citrate cycle (TCA cycle) containing 2-oxoglutarate dehydrogenase E1 component (sucA, sucB) and succinyl-CoA synthetase (sucC, sucD); NADH dehydrogenase I (nuoA, nuoB, nuoC, nuoF, nuoH, nuoI, nuoL, nuoM), which is a part of the oxidative phosphorylation pathway; nitrate reductase 1 (with narG, narH, narI, and narJ); and a portion of the bacterial secretion system (with secA, secD, secY).</p>
</sec></sec><sec id="s4">
<title>Discussion</title>
<p>In this paper, we proposed an HMM-based network alignment algorithm that can be used for finding conserved pathways in two or more biological networks. The HMM framework and the proposed alignment algorithm has a number of important advantages compared to other existing local network alignment algorithms. First of all, despite its generality, the proposed algorithm is very simple and efficient. In fact, the alignment algorithm based on the proposed HMM framework is a variant of the Viterbi algorithm. As a result, it has a very low polynomial computational complexity, which grows only linearly with respect to the length of the identified pathways and the number of edges in each network. This makes it possible to find conserved pathways with more than 10 nodes in networks with thousands of nodes and tens of thousands of interactions within a few minutes on a personal computer. Furthermore, the HMM-based framework can handle a large class of path isomorphism, which allows us to find pathway alignments with any number of gaps (node insertions and deletions) at arbitrary locations. In addition to this, the proposed framework is very flexible in choosing the scoring scheme for pathway alignments, where different penalties can be used for mismatches, insertions and deletions. We can also assign different penalties for gap opening and gap extension, which can be convenient when comparing networks that are remotely related to each other. Another important advantage of the proposed framework is that it allows us to use an efficient dynamic programming algorithm for finding the mathematically optimal alignment. Considering that many available algorithms rely on heuristics that cannot guarantee the optimality of the obtained solutions, this is certainly a significant merit of the HMM-based approach. Although the mathematical optimality does not guarantee the biological significance of the obtained solution, it can certainly lead to more accurate predictions if combined with a realistic scoring scheme for assessing pathway homology. As demonstrated in our experiments, the proposed algorithm yields accurate and biologically meaningful results both for querying known pathways in the network of another organism and also for finding conserved functional modules in the networks of different organisms. Finally, the HMM-based framework presented in this paper can be extended for aligning multiple networks. While many current multiple network alignment algorithms adopt a progressive approach for comparing multiple networks <xref ref-type="bibr" rid="pone.0008070-Flannick1">[9]</xref>, <xref ref-type="bibr" rid="pone.0008070-Flannick2">[14]</xref>–<xref ref-type="bibr" rid="pone.0008070-Liao1">[17]</xref>, our HMM-based framework provides a potential way to simultaneously align multiple networks to find the optimal set of conserved pathways with maximum alignment score.</p>
<p>For future research, we plan to evaluate the performance of our HMM-based algorithm more extensively by investigating the consistency of the predicted alignments based on other available functional annotations, including the gene ontology (GO) annotations <xref ref-type="bibr" rid="pone.0008070-Ashburner1">[31]</xref>. It would be also beneficial to develop a more elaborate scoring scheme that integrates additional information, such as the GO annotations and the KO group annotations, to obtain more reliable alignment results. Finally, we are currently working on simultaneous multiple network alignment based on the HMM framework, where the goal is to construct a scalable multiple alignment algorithm that yields network alignments with higher fidelity.</p>
</sec><sec id="s5">
<title>Supporting Information</title>
<supplementary-material id="pone.0008070.s001" mimetype="application/pdf" position="float" xlink:href="info:doi/10.1371/journal.pone.0008070.s001" xlink:type="simple"><label>Supporting Information S1</label><caption>
<p>(0.06 MB PDF)</p>
</caption></supplementary-material></sec></body>
<back>
<ack>
<p>The authors would also like to thank Maxim Kalaev, Wenhong Tian, as well as Jason Flannick for sharing the datasets and for the helpful communication.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="pone.0008070-Ito1"><label>1</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Ito</surname><given-names>T</given-names></name>
<name name-style="western"><surname>Chiba</surname><given-names>T</given-names></name>
<name name-style="western"><surname>Ozawa</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Yoshida</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Hattori</surname><given-names>M</given-names></name>
<etal/></person-group>             <year>2001</year>             <article-title>A comprehensive two-hybrid analysis to explore the yeast protein interactome.</article-title>             <source>Proc Natl Acad Sci USA</source>             <volume>98</volume>             <fpage>4569</fpage>             <lpage>4574</lpage>          </element-citation></ref>
<ref id="pone.0008070-Mann1"><label>2</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Mann</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Hendrickson</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Pandey</surname><given-names>A</given-names></name>
</person-group>             <year>2001</year>             <article-title>Analysis of proteins and proteomes by mass spectrometry.</article-title>             <source>Annu Rev Biochem</source>             <volume>70</volume>             <fpage>437</fpage>             <lpage>473</lpage>          </element-citation></ref>
<ref id="pone.0008070-Uetz1"><label>3</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Uetz</surname><given-names>P</given-names></name>
<name name-style="western"><surname>Rajagopala</surname><given-names>S</given-names></name>
<name name-style="western"><surname>Dong</surname><given-names>Y</given-names></name>
<name name-style="western"><surname>Haas</surname><given-names>J</given-names></name>
</person-group>             <year>2004</year>             <article-title>From orfeomes to protein interaction maps in viruses.</article-title>             <source>Genome Res</source>             <volume>14</volume>             <fpage>2029</fpage>             <lpage>2033</lpage>          </element-citation></ref>
<ref id="pone.0008070-Krogan1"><label>4</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Krogan</surname><given-names>N</given-names></name>
<etal/></person-group>             <year>2006</year>             <article-title>Global landscape of protein complexes in the yeast saccharomyces cerevisiae.</article-title>             <source>Nature</source>             <volume>440</volume>             <fpage>4412</fpage>             <lpage>4415</lpage>          </element-citation></ref>
<ref id="pone.0008070-vonMering1"><label>5</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>von Mering</surname><given-names>C</given-names></name>
<name name-style="western"><surname>Krause</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Snel</surname><given-names>B</given-names></name>
<name name-style="western"><surname>Cornell</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Oliver</surname><given-names>S</given-names></name>
<etal/></person-group>             <year>2002</year>             <article-title>Comparative assessment of large-scale data sets of protein-protein interactions.</article-title>             <source>Nature</source>             <volume>417</volume>             <fpage>399</fpage>             <lpage>403</lpage>          </element-citation></ref>
<ref id="pone.0008070-Kelley1"><label>6</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Kelley</surname><given-names>B</given-names></name>
<name name-style="western"><surname>Sharan</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Karp</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Sittler</surname><given-names>T</given-names></name>
<name name-style="western"><surname>Root</surname><given-names>D</given-names></name>
<etal/></person-group>             <year>2003</year>             <article-title>Conserved pathways within bacteria and yeast as revealed by global protein network alignment.</article-title>             <source>Proc Natl Acad Sci USA</source>             <volume>100</volume>             <fpage>11394</fpage>             <lpage>11399</lpage>          </element-citation></ref>
<ref id="pone.0008070-Koyutrk1"><label>7</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Koyutürk</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Grama</surname><given-names>A</given-names></name>
<name name-style="western"><surname>Szpankowski</surname><given-names>W</given-names></name>
</person-group>             <year>2004</year>             <article-title>An efficient algorithm for detecting frequent subgraphs in biological networks.</article-title>             <source>Bioinformatics</source>             <volume>20</volume>             <fpage>SI200</fpage>             <lpage>207</lpage>          </element-citation></ref>
<ref id="pone.0008070-Sharan1"><label>8</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Sharan</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Suthram</surname><given-names>S</given-names></name>
<name name-style="western"><surname>Kelley</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Kuhn</surname><given-names>T</given-names></name>
<name name-style="western"><surname>McCuine</surname><given-names>S</given-names></name>
<etal/></person-group>             <year>2005</year>             <article-title>Conserved patterns of protein interaction in multiple species.</article-title>             <source>Proc Natl Acad Sci USA</source>             <volume>102</volume>             <fpage>1974</fpage>             <lpage>1979</lpage>          </element-citation></ref>
<ref id="pone.0008070-Flannick1"><label>9</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Flannick</surname><given-names>J</given-names></name>
<name name-style="western"><surname>Novak</surname><given-names>A</given-names></name>
<name name-style="western"><surname>Srinivasan</surname><given-names>B</given-names></name>
<name name-style="western"><surname>McAdams</surname><given-names>H</given-names></name>
<name name-style="western"><surname>Batzoglou</surname><given-names>S</given-names></name>
</person-group>             <year>2006</year>             <article-title>Græmlin: general and robust alignment of multiple large interaction networks.</article-title>             <source>Genome Res</source>             <volume>16</volume>             <fpage>1169</fpage>             <lpage>1181</lpage>          </element-citation></ref>
<ref id="pone.0008070-Scott1"><label>10</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Scott</surname><given-names>J</given-names></name>
<name name-style="western"><surname>Ideker</surname><given-names>T</given-names></name>
<name name-style="western"><surname>Karp</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Sharan</surname><given-names>R</given-names></name>
</person-group>             <year>2006</year>             <article-title>Efficient algorithms for detecting signaling pathways in protein interaction networks.</article-title>             <source>J Comput Biol</source>             <volume>13</volume>             <fpage>133</fpage>             <lpage>144</lpage>          </element-citation></ref>
<ref id="pone.0008070-Shlomi1"><label>11</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Shlomi</surname><given-names>T</given-names></name>
<name name-style="western"><surname>Segal</surname><given-names>D</given-names></name>
<name name-style="western"><surname>Ruppin</surname><given-names>E</given-names></name>
<name name-style="western"><surname>Sharan</surname><given-names>R</given-names></name>
</person-group>             <year>2006</year>             <article-title>QPath: a method for querying pathways in a protein-protein interaction network.</article-title>             <source>BMC Bioinformatics</source>             <volume>7</volume>          </element-citation></ref>
<ref id="pone.0008070-Yang1"><label>12</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Q</given-names></name>
<name name-style="western"><surname>Sze</surname><given-names>S</given-names></name>
</person-group>             <year>2007</year>             <article-title>Path matching and graph matching in biological networks.</article-title>             <source>J Comput Biol</source>             <volume>14</volume>             <fpage>56</fpage>             <lpage>67</lpage>          </element-citation></ref>
<ref id="pone.0008070-Dost1"><label>13</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Dost</surname><given-names>B</given-names></name>
<name name-style="western"><surname>Shlomi</surname><given-names>T</given-names></name>
<name name-style="western"><surname>Gupta</surname><given-names>N</given-names></name>
<name name-style="western"><surname>Ruppin</surname><given-names>E</given-names></name>
<name name-style="western"><surname>Bafna</surname><given-names>V</given-names></name>
<etal/></person-group>             <year>2008</year>             <article-title>QNet: a tool for querying protein interaction networks.</article-title>             <source>J Comput Biol</source>             <volume>15</volume>             <fpage>913</fpage>             <lpage>925</lpage>          </element-citation></ref>
<ref id="pone.0008070-Flannick2"><label>14</label><element-citation publication-type="other" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Flannick</surname><given-names>J</given-names></name>
<name name-style="western"><surname>Novak</surname><given-names>A</given-names></name>
<name name-style="western"><surname>Dol</surname><given-names>C</given-names></name>
<name name-style="western"><surname>Srinivasan</surname><given-names>B</given-names></name>
<name name-style="western"><surname>Batzoglou</surname><given-names>S</given-names></name>
</person-group>             <year>2008</year>             <article-title>Automatic parameter learning for multiple network alignment.</article-title>             <comment>In: Proc of the 10th Annu Int Conf Res Comput Mol Bio (RECOMB 2008)</comment>          </element-citation></ref>
<ref id="pone.0008070-Kalaev1"><label>15</label><element-citation publication-type="other" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Kalaev</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Bafna</surname><given-names>V</given-names></name>
<name name-style="western"><surname>Sharan</surname><given-names>R</given-names></name>
</person-group>             <year>2008</year>             <article-title>Fast and accurate alignment of multiple protein networks.</article-title>             <comment>In: Proc of the 10th Annu Int Conf Res Comput Mol Bio (RECOMB 2008)</comment>          </element-citation></ref>
<ref id="pone.0008070-Singh1"><label>16</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Singh</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Xu</surname><given-names>J</given-names></name>
<name name-style="western"><surname>Berger</surname><given-names>B</given-names></name>
</person-group>             <year>2008</year>             <article-title>Global alignment of multiple protein interaction networks with application to functional orthology detection.</article-title>             <source>Proc Natl Acad Sci USA</source>             <volume>105</volume>             <fpage>12763</fpage>             <lpage>12768</lpage>          </element-citation></ref>
<ref id="pone.0008070-Liao1"><label>17</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Liao</surname><given-names>C</given-names></name>
<name name-style="western"><surname>Lu</surname><given-names>K</given-names></name>
<name name-style="western"><surname>Baym</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Singh</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Berger</surname><given-names>B</given-names></name>
</person-group>             <year>2009</year>             <article-title>IsoRankN: spectral methods for global alignment of multiple protein networks.</article-title>             <source>Bioinformatics</source>             <volume>25</volume>             <fpage>253</fpage>             <lpage>258</lpage>          </element-citation></ref>
<ref id="pone.0008070-Qian1"><label>18</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Qian</surname><given-names>X</given-names></name>
<name name-style="western"><surname>Sze</surname><given-names>S</given-names></name>
<name name-style="western"><surname>Yoon</surname><given-names>B</given-names></name>
</person-group>             <year>2009</year>             <article-title>Querying pathways in protein interaction networks based on hidden markov models.</article-title>             <source>J Comput Biol</source>             <volume>16</volume>             <fpage>145</fpage>             <lpage>157</lpage>          </element-citation></ref>
<ref id="pone.0008070-Tian1"><label>19</label><element-citation publication-type="other" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Tian</surname><given-names>W</given-names></name>
<name name-style="western"><surname>Samatova</surname><given-names>N</given-names></name>
</person-group>             <year>2009</year>             <article-title>Pairwise alignment of interaction networks by fast identification of maximal conserved patterns.</article-title>             <fpage>99</fpage>             <lpage>110</lpage>             <comment>In: Pac Symp Biocomput. volume 14</comment>          </element-citation></ref>
<ref id="pone.0008070-Zaslavskiy1"><label>20</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Zaslavskiy</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Bach</surname><given-names>F</given-names></name>
<name name-style="western"><surname>Vert</surname><given-names>J</given-names></name>
</person-group>             <year>2009</year>             <article-title>Global alignment of protein-protein interaction entworks by graph matching methods.</article-title>             <source>Bioinformatics</source>             <volume>25</volume>             <fpage>259</fpage>             <lpage>267</lpage>          </element-citation></ref>
<ref id="pone.0008070-Pinter1"><label>21</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Pinter</surname><given-names>R</given-names></name>
<name name-style="western"><surname>Rokhlenko</surname><given-names>O</given-names></name>
<name name-style="western"><surname>Yeger-Lotem</surname><given-names>E</given-names></name>
<name name-style="western"><surname>Ziv-Ukelson</surname><given-names>M</given-names></name>
</person-group>             <year>2005</year>             <article-title>Alignment of metabolic pathways.</article-title>             <source>Bioinformatics</source>             <volume>21</volume>             <fpage>3401</fpage>             <lpage>3408</lpage>          </element-citation></ref>
<ref id="pone.0008070-Akutsu1"><label>22</label><element-citation publication-type="other" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Akutsu</surname><given-names>T</given-names></name>
<name name-style="western"><surname>Kuhara</surname><given-names>S</given-names></name>
<name name-style="western"><surname>Maruyama</surname><given-names>O</given-names></name>
<name name-style="western"><surname>Miyano</surname><given-names>S</given-names></name>
</person-group>             <year>1998</year>             <article-title>Identification of gene regulatory networks by strategic gene disruptions and gene overexpressions.</article-title>             <fpage>695</fpage>             <lpage>706</lpage>             <comment>In: Proc. 9th Annu. ACM-SIAM Symp. Discrete Alg</comment>          </element-citation></ref>
<ref id="pone.0008070-Steffen1"><label>23</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Steffen</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Petti</surname><given-names>A</given-names></name>
<name name-style="western"><surname>Aach</surname><given-names>J</given-names></name>
<name name-style="western"><surname>D'haeseleer</surname><given-names>P</given-names></name>
<name name-style="western"><surname>Church</surname><given-names>G</given-names></name>
</person-group>             <year>2002</year>             <article-title>Automated modelling of signal transduction networks.</article-title>             <source>BMC Bioinformatics</source>             <volume>3</volume>             <fpage>34</fpage>          </element-citation></ref>
<ref id="pone.0008070-Li1"><label>24</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Li</surname><given-names>Z</given-names></name>
<name name-style="western"><surname>Zhang</surname><given-names>S</given-names></name>
<name name-style="western"><surname>Wang</surname><given-names>Y</given-names></name>
<name name-style="western"><surname>Zhang</surname><given-names>X</given-names></name>
<name name-style="western"><surname>Chen</surname><given-names>L</given-names></name>
</person-group>             <year>2007</year>             <article-title>Alignment of molecular networks by integer quadratic programming.</article-title>             <source>Bioinformatics</source>             <volume>23</volume>             <fpage>1631</fpage>             <lpage>1639</lpage>          </element-citation></ref>
<ref id="pone.0008070-Pearson1"><label>25</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Pearson</surname><given-names>W</given-names></name>
<name name-style="western"><surname>Lipman</surname><given-names>D</given-names></name>
</person-group>             <year>1988</year>             <article-title>Improved tools for biological sequence comparison.</article-title>             <source>Proc Natl Acad Sci USA</source>             <volume>85</volume>             <fpage>2444</fpage>             <lpage>2448</lpage>          </element-citation></ref>
<ref id="pone.0008070-Pearson2"><label>26</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Pearson</surname><given-names>W</given-names></name>
</person-group>             <year>1996</year>             <article-title>Effective protein sequence comparison.</article-title>             <source>Methods Enzymol</source>             <volume>266</volume>             <fpage>227</fpage>             <lpage>258</lpage>          </element-citation></ref>
<ref id="pone.0008070-Pagni1"><label>27</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Pagni</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Jongeneel</surname><given-names>C</given-names></name>
</person-group>             <year>2001</year>             <article-title>Making sense of score statistics for sequence alignments.</article-title>             <source>Brief Bioinform</source>             <volume>2</volume>             <fpage>51</fpage>             <lpage>67</lpage>          </element-citation></ref>
<ref id="pone.0008070-Xenarios1"><label>28</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Xenarios</surname><given-names>I</given-names></name>
<name name-style="western"><surname>Salwinski</surname><given-names>L</given-names></name>
<name name-style="western"><surname>Duan</surname><given-names>X</given-names></name>
<name name-style="western"><surname>Higney</surname><given-names>P</given-names></name>
<name name-style="western"><surname>Kim</surname><given-names>S</given-names></name>
<etal/></person-group>             <year>2002</year>             <article-title>DIP, the Database of Interacting Proteins: a research tool for studying cellular networks of protein interactions.</article-title>             <source>Nucleic Acids Res</source>             <volume>30</volume>             <fpage>303</fpage>             <lpage>305</lpage>          </element-citation></ref>
<ref id="pone.0008070-Kanehisa1"><label>29</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Kanehisa</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Goto</surname><given-names>S</given-names></name>
</person-group>             <year>2000</year>             <article-title>KEGG: Kyoto encyclopedia of genes and genomes.</article-title>             <source>Nucleic Acids Res</source>             <volume>28</volume>             <fpage>27</fpage>             <lpage>30</lpage>          </element-citation></ref>
<ref id="pone.0008070-Srinivasan1"><label>30</label><element-citation publication-type="other" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Srinivasan</surname><given-names>B</given-names></name>
<name name-style="western"><surname>Novak</surname><given-names>A</given-names></name>
<name name-style="western"><surname>Flannick</surname><given-names>J</given-names></name>
<name name-style="western"><surname>Batzoglou</surname><given-names>S</given-names></name>
<name name-style="western"><surname>McAdams</surname><given-names>H</given-names></name>
</person-group>             <year>2006</year>             <article-title>Integrated protein interaction networks for 11 microbes.</article-title>             <comment>In: Proc of the 10th Annu Int Conf Res Comput Mol Bio (RECOMB 2006)</comment>          </element-citation></ref>
<ref id="pone.0008070-Ashburner1"><label>31</label><element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author">
<name name-style="western"><surname>Ashburner</surname><given-names>M</given-names></name>
<name name-style="western"><surname>Ball</surname><given-names>CA</given-names></name>
<name name-style="western"><surname>Blake</surname><given-names>JA</given-names></name>
<name name-style="western"><surname>Botstein</surname><given-names>D</given-names></name>
<name name-style="western"><surname>Butler</surname><given-names>H</given-names></name>
<etal/></person-group>             <year>2000</year>             <article-title>Gene ontology: tool for the unification of biology. the gene ontology consortium.</article-title>             <source>Nat Genet</source>             <volume>25</volume>             <fpage>25</fpage>             <lpage>29</lpage>          </element-citation></ref>
</ref-list>

</back>
</article>