<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">plos</journal-id>
<journal-id journal-id-type="nlm-ta">PLoS Comput Biol</journal-id>
<journal-id journal-id-type="pmc">ploscomp</journal-id><journal-title-group>
<journal-title>PLoS Computational Biology</journal-title></journal-title-group>
<issn pub-type="ppub">1553-734X</issn>
<issn pub-type="epub">1553-7358</issn>
<publisher>
<publisher-name>Public Library of Science</publisher-name>
<publisher-loc>San Francisco, USA</publisher-loc></publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">PCOMPBIOL-D-13-02211</article-id>
<article-id pub-id-type="doi">10.1371/journal.pcbi.1003818</article-id>
<article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Biology and life sciences</subject><subj-group><subject>Evolutionary biology</subject><subj-group><subject>Evolutionary processes</subject><subj-group><subject>Evolutionary adaptation</subject><subject>Natural selection</subject></subj-group></subj-group></subj-group><subj-group><subject>Genetics</subject><subj-group><subject>Mutation</subject></subj-group></subj-group></subj-group></article-categories>
<title-group>
<article-title>The Time Scale of Evolutionary Innovation</article-title>
<alt-title alt-title-type="running-head">The Time Scale of Evolutionary Innovation</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Chatterjee</surname><given-names>Krishnendu</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Pavlogiannis</surname><given-names>Andreas</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Adlam</surname><given-names>Ben</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Nowak</surname><given-names>Martin A.</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib>
</contrib-group>
<aff id="aff1"><label>1</label><addr-line>IST Austria, Klosterneuburg, Austria</addr-line></aff>
<aff id="aff2"><label>2</label><addr-line>Program for Evolutionary Dynamics, Department of Organismic and Evolutionary Biology, Department of Mathematics, Harvard University, Cambridge, Massachusetts, United States of America</addr-line></aff>
<contrib-group>
<contrib contrib-type="editor" xlink:type="simple"><name name-style="western"><surname>Beerenwinkel</surname><given-names>Niko</given-names></name>
<role>Editor</role>
<xref ref-type="aff" rid="edit1"/></contrib>
</contrib-group>
<aff id="edit1"><addr-line>ETH Zurich, Switzerland</addr-line></aff>
<author-notes>
<corresp id="cor1">* E-mail: <email xlink:type="simple">Krishnendu.Chatterjee@ist.ac.at</email></corresp>
<fn fn-type="conflict"><p>The authors have declared that no competing interests exist.</p></fn>
<fn fn-type="con"><p>Conceived and designed the experiments: KC AP BA MAN. Analyzed the data: KC AP BA MAN. Wrote the paper: KC AP BA MAN.</p></fn>
</author-notes>
<pub-date pub-type="collection"><month>9</month><year>2014</year></pub-date>
<pub-date pub-type="epub"><day>11</day><month>9</month><year>2014</year></pub-date>
<volume>10</volume>
<issue>9</issue>
<elocation-id>e1003818</elocation-id>
<history>
<date date-type="received"><day>13</day><month>12</month><year>2013</year></date>
<date date-type="accepted"><day>21</day><month>7</month><year>2014</year></date>
</history>
<permissions>
<copyright-year>2014</copyright-year>
<copyright-holder>Chatterjee et al</copyright-holder><license xlink:type="simple"><license-p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/" xlink:type="simple">Creative Commons Attribution License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</license-p></license></permissions>
<abstract>
<p>A fundamental question in biology is the following: what is the time scale that is needed for evolutionary innovations? There are many results that characterize single steps in terms of the fixation time of new mutants arising in populations of certain size and structure. But here we ask a different question, which is concerned with the much longer time scale of evolutionary trajectories: how long does it take for a population exploring a fitness landscape to find target sequences that encode new biological functions? Our key variable is the length, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e001" xlink:type="simple"/></inline-formula> of the genetic sequence that undergoes adaptation. In computer science there is a crucial distinction between problems that require algorithms which take polynomial or exponential time. The latter are considered to be intractable. Here we develop a theoretical approach that allows us to estimate the time of evolution as function of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e002" xlink:type="simple"/></inline-formula> We show that adaptation on many fitness landscapes takes time that is exponential in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e003" xlink:type="simple"/></inline-formula> even if there are broad selection gradients and many targets uniformly distributed in sequence space. These negative results lead us to search for specific mechanisms that allow evolution to work on polynomial time scales. We study a regeneration process and show that it enables evolution to work in polynomial time.</p>
</abstract>
<abstract abstract-type="summary"><title>Author Summary</title>
<p>Evolutionary adaptation can be described as a biased, stochastic walk of a population of sequences in a high dimensional sequence space. The population explores a fitness landscape. The mutation-selection process biases the population towards regions of higher fitness. In this paper we estimate the time scale that is needed for evolutionary innovation. Our key parameter is the length of the genetic sequence that needs to be adapted. We show that a variety of evolutionary processes take exponential time in sequence length. We propose a specific process, which we call ‘regeneration processes’, and show that it allows evolution to work on polynomial time scales. In this view, evolution can solve a problem efficiently if it has solved a similar problem already.</p>
</abstract>
<funding-group><funding-statement>Austrian Science Fund (FWF) Grant No P23499-N23, FWF NFN Grant No S11407-N23 (RiSE), ERC Start grant (279307: Graph Games), and Microsoft Faculty Fellows award. Support from the John Templeton foundation is gratefully acknowledged. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</funding-statement></funding-group><counts><page-count count="7"/></counts></article-meta>
</front>
<body><sec id="s1">
<title>Introduction</title>
<p>Our planet came into existence 4.6 billion years ago. There is clear chemical evidence for life on earth 3.5 billion years ago <xref ref-type="bibr" rid="pcbi.1003818-Allwood1">[1]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Schopf1">[2]</xref>. The evolutionary process generated procaria, eucaria and complex multi-cellular organisms. Throughout the history of life, evolution had to discover sequences of biological polymers that perform specific, complicated functions. The average length of bacterial genes is about 1000 nucleotides, that of human genes about 3000 nucleotides. The longest known bacterial gene contains more than <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e004" xlink:type="simple"/></inline-formula> nucleotides, the longest human gene more than <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e005" xlink:type="simple"/></inline-formula>. A basic question is what is the time scale required by evolution to discover the sequences that perform desired functions. While many results exist for the fixation time of individual mutants <xref ref-type="bibr" rid="pcbi.1003818-Kimura1">[3]</xref>–<xref ref-type="bibr" rid="pcbi.1003818-Ohta1">[15]</xref>, here we ask how the time scale of evolution depends on the length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e006" xlink:type="simple"/></inline-formula> of the sequence that needs to be adapted. We consider the crucial distinction of polynomial versus exponential time <xref ref-type="bibr" rid="pcbi.1003818-Papadimitriou1">[16]</xref>–<xref ref-type="bibr" rid="pcbi.1003818-Valiant1">[18]</xref>. A time scale that grows exponentially in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e007" xlink:type="simple"/></inline-formula> is infeasible for long sequences.</p>
<p>Evolutionary dynamics operates in sequence space, which can be imagined as a discrete multi-dimensional lattice that arises when all sequences of a given length are arranged such that nearest neighbors differ by one point mutation <xref ref-type="bibr" rid="pcbi.1003818-MaynardSmith1">[19]</xref>. For constant selection, each point in sequence space is associated with a non-negative fitness value (reproductive rate). The resulting fitness landscape is a high dimensional mountain range. Populations explore fitness landscapes searching for elevated regions, ridges, and peaks <xref ref-type="bibr" rid="pcbi.1003818-Fontana1">[20]</xref>–<xref ref-type="bibr" rid="pcbi.1003818-Worden1">[27]</xref>.</p>
<p>A question that has been extensively studied is how long does it take for existing biological functions to improve under natural selection. This problem leads to the study of adaptive walks on fitness landscapes <xref ref-type="bibr" rid="pcbi.1003818-Ohta1">[15]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Fontana1">[20]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Fontana2">[21]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Crow1">[28]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Gillespie1">[29]</xref>. In this paper we ask a different question: how long does it take for evolution to discover a new function? More specifically, our aim is to estimate the expected discovery time of new biological functions: how long does it take for a population of reproducing organisms to discover a biological function that is not present at the beginning of the search. We will discuss two approximations for rugged fitness landscapes. We also discuss the significance of clustered peaks.</p>
<p>We consider an alphabet of size four, as is the case for DNA and RNA, and a nucleotide sequence of length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e008" xlink:type="simple"/></inline-formula>. We consider a population of size <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e009" xlink:type="simple"/></inline-formula>, which reproduces asexually. The mutation rate, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e010" xlink:type="simple"/></inline-formula>, is small: individual mutations are introduced and evaluated by natural selection and random drift one at a time. The probability that the evolutionary process moves from a sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e011" xlink:type="simple"/></inline-formula> to a sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e012" xlink:type="simple"/></inline-formula>, which is at Hamming distance one from <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e013" xlink:type="simple"/></inline-formula>, is given by <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e014" xlink:type="simple"/></inline-formula>, where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e015" xlink:type="simple"/></inline-formula> is the fixation probability of sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e016" xlink:type="simple"/></inline-formula> in a population consisting of sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e017" xlink:type="simple"/></inline-formula>. In the special case of a flat fitness landscape, we have <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e018" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e019" xlink:type="simple"/></inline-formula>. Thus we have an evolutionary random walk, where each step is a jump to a neighboring sequence of Hamming distance one.</p>
</sec><sec id="s2">
<title>Results</title>
<p>Consider a high-dimensional sequence space. A particular biological function can be instantiated by some of the sequences. Each sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e020" xlink:type="simple"/></inline-formula> has a fitness value <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e021" xlink:type="simple"/></inline-formula>, which measures the ability of the sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e022" xlink:type="simple"/></inline-formula> to encode the desired function. Biological fitness landscapes are typically expected to have many peaks <xref ref-type="bibr" rid="pcbi.1003818-Gillespie1">[29]</xref>–<xref ref-type="bibr" rid="pcbi.1003818-Orr2">[31]</xref>. They can be highly rugged due to epistatic effects of mutations <xref ref-type="bibr" rid="pcbi.1003818-Weinreich1">[32]</xref>–<xref ref-type="bibr" rid="pcbi.1003818-Woods1">[34]</xref>. They can also contain large regions or networks of neutrality <xref ref-type="bibr" rid="pcbi.1003818-Fontana1">[20]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Fontana2">[21]</xref>. Empirical studies of short RNA sequences have revealed that the underlying fitness landscape has low peak density <xref ref-type="bibr" rid="pcbi.1003818-Jimenez1">[35]</xref>: around <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e023" xlink:type="simple"/></inline-formula> peaks in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e024" xlink:type="simple"/></inline-formula> sequences.</p>
<p>For the purpose of estimating the expected discovery time we can approximate the fitness landscape with a binary step function over the sequence space. We discuss two different approximations (<xref ref-type="fig" rid="pcbi-1003818-g001">Figure 1</xref>). For the first approximation, we consider the scenario where fitness values below some threshold, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e025" xlink:type="simple"/></inline-formula>, have negligible contribution; those sequences do not instantiate the desired function (either not at all or only below the minimum level that could be detected by natural selection). We approximate the rugged fitness landscape as follows: if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e026" xlink:type="simple"/></inline-formula> then <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e027" xlink:type="simple"/></inline-formula>; if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e028" xlink:type="simple"/></inline-formula> then <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e029" xlink:type="simple"/></inline-formula>. The set of sequences with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e030" xlink:type="simple"/></inline-formula> constitutes the target set, and the remaining fitness landscape is neutral.</p>
<fig id="pcbi-1003818-g001" position="float"><object-id pub-id-type="doi">10.1371/journal.pcbi.1003818.g001</object-id><label>Figure 1</label><caption>
<title>Approximations of a highly rugged fitness landscape by broad peaks and neutral regions.</title>
<p>The figures depict examples of highly rugged fitness landscapes where the sequence space has been projected in one dimension. (A) Sequences with fitness below some level <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e031" xlink:type="simple"/></inline-formula> are functionally very different to the desired function, and selection cannot act upon them. All other sequences are considered as targets. The fitness landscape is approximated by a step function: if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e032" xlink:type="simple"/></inline-formula>, then <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e033" xlink:type="simple"/></inline-formula>, otherwise <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e034" xlink:type="simple"/></inline-formula>. (B) Local maxima below the desired fitness threshold <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e035" xlink:type="simple"/></inline-formula> are known to slow down the evolutionary random walk towards sequences that attain fitness at least <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e036" xlink:type="simple"/></inline-formula>. We approximate the fitness landscape by broad peaks and neutral regions by increasing the fitness of every sequence that belongs in a mountain range with fitness below <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e037" xlink:type="simple"/></inline-formula> to the maximal local maxima <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e038" xlink:type="simple"/></inline-formula> below <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e039" xlink:type="simple"/></inline-formula>. Note that the target set starts from the upslope of a mountain range whose peak exceeds <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e040" xlink:type="simple"/></inline-formula>.</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pcbi.1003818.g001" position="float" xlink:type="simple"/></fig>
<p>The second approximation works as follows. Consider the evolutionary process exploring a rugged fitness landscape where the goal is to attain a fitness level <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e041" xlink:type="simple"/></inline-formula>. Local maxima below <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e042" xlink:type="simple"/></inline-formula> slow down the evolutionary process to attain <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e043" xlink:type="simple"/></inline-formula>, because the evolutionary walk might get stuck in those local maxima. In order to derive lower bounds for the expected discovery time, the rugged fitness landscape can be approximated as follows. Let <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e044" xlink:type="simple"/></inline-formula> be the fitness value of the highest local maximum below <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e045" xlink:type="simple"/></inline-formula>. Then for every sequence in a mountain range with a local maximum below <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e046" xlink:type="simple"/></inline-formula> we assign the fitness value <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e047" xlink:type="simple"/></inline-formula>. The mountain ranges with local maxima above <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e048" xlink:type="simple"/></inline-formula> are the target sequences. Note that the target set includes sequences that start at the upslope of mountain ranges with peaks above <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e049" xlink:type="simple"/></inline-formula>. Thus, again we obtain a fitness landscape with clustered targets and neutral region, where the neutral region consists of all sequences whose fitness values have been assigned to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e050" xlink:type="simple"/></inline-formula>. The two approximations are illustrated in <xref ref-type="fig" rid="pcbi-1003818-g001">Figure 1</xref>. For <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e051" xlink:type="simple"/></inline-formula> the second approximation generates larger target areas than the first approximation and is therefore more lenient.</p>
<p>Our key results for estimating the discovery time can now be formulated for binary fitness landscapes, but they apply to any type of rugged landscape using one of the two approximations. We note that our methods can also be applied for certain non-binary fitness landscapes, and an example of a fitness landscape with a large gradient arising from multiplicative fitness effects is discussed in Sections 6 and 7 of <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>.</p>
<p>We now present our main results in the following order. We first estimate the discovery time of a single search aiming to find a single broad peak. Then we study multiple simultaneous searches for a single broad peak. Finally, we consider multiple broad peaks that are uniformly randomly distributed in sequence space.</p>
<p>We first study a broad peak of target sequences described as follows: consider a specific sequence; any sequence within a certain Hamming distance of that sequence belongs to the target set. Specifically, we consider that the evolutionary process has succeeded, if the population discovers a sequence that differs from the specific sequence in no more than a fraction <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e052" xlink:type="simple"/></inline-formula> of positions. We refer to the specific sequence as the target center and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e053" xlink:type="simple"/></inline-formula> as the width (or radius) of the peak. For example, if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e054" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e055" xlink:type="simple"/></inline-formula>, then the target center is surrounded by a cloud of approximately <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e056" xlink:type="simple"/></inline-formula> sequences. For a single broad peak with width <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e057" xlink:type="simple"/></inline-formula>, the target set contains at least <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e058" xlink:type="simple"/></inline-formula> sequences, which is an exponential function of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e059" xlink:type="simple"/></inline-formula>. The fitness landscape outside the broad peak is flat. We refer this binary fitness landscape as a broad peak landscape. The population needs to discover any one of the target sequences in the broad peak, starting from some sequence that is not in the broad peak. We establish the following result.</p>
<sec id="s2a">
<title/>
<sec id="s2a1">
<title>Theorem 1</title>
<p><italic>Consider a single search exploring a broad peak landscape with width </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e060" xlink:type="simple"/></inline-formula><italic> and mutation rate </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e061" xlink:type="simple"/></inline-formula><italic>. The following assertions hold</italic>:</p>
<list list-type="bullet"><list-item>
<p><italic>if </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e062" xlink:type="simple"/></inline-formula><italic>, then there exists </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e063" xlink:type="simple"/></inline-formula><italic> such that for all sequence spaces of sequence length </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e064" xlink:type="simple"/></inline-formula><italic>, the expected discovery time is at least </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e065" xlink:type="simple"/></inline-formula><italic>;</italic></p>
</list-item><list-item>
<p><italic>if </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e066" xlink:type="simple"/></inline-formula><italic>, then for all sequence spaces of sequence length </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e067" xlink:type="simple"/></inline-formula><italic>, the expected discovery time is at most </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e068" xlink:type="simple"/></inline-formula><italic>.</italic></p>
</list-item></list>
<p>Our result can be interpreted as follows (see Theorem S2 and Corollary S2 in <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>): (i) If <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e069" xlink:type="simple"/></inline-formula>, then the expected discovery time is exponential in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e070" xlink:type="simple"/></inline-formula>; and (ii) if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e071" xlink:type="simple"/></inline-formula>, then the expected discovery time is polynomial in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e072" xlink:type="simple"/></inline-formula>. Thus, we have derived a <italic>strong dichotomy</italic> result which shows a sharp transition from polynomial to exponential time depending on whether a specific condition on <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e073" xlink:type="simple"/></inline-formula> does or does not hold.</p>
<p>For the four letter alphabet most random sequences have Hamming distance <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e074" xlink:type="simple"/></inline-formula> from the target center. If the population is further away than this Hamming distance, then random drift will bring it closer. If the population is closer than this Hamming distance, then random drift will push it further away. This argument constitutes the intuitive reason that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e075" xlink:type="simple"/></inline-formula> is the critical threshold. If the peak has a width of less than <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e076" xlink:type="simple"/></inline-formula>, then we prove that the expected discovery time by random drift is exponential in the sequence length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e077" xlink:type="simple"/></inline-formula> (see <xref ref-type="fig" rid="pcbi-1003818-g002">Figure 2</xref>). This result holds for any population size, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e078" xlink:type="simple"/></inline-formula>, as long as <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e079" xlink:type="simple"/></inline-formula>, which is certainly the case for realistic values of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e080" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e081" xlink:type="simple"/></inline-formula>. In the <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref> we also present a more general result, where along with a single broad peak, instead of a flat landscape outside the peak we consider a multiplicative fitness landscape and establish a sharp dichotomy result that generalizes Theorem 1 (see Corollary S2 in <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>).</p>
<fig id="pcbi-1003818-g002" position="float"><object-id pub-id-type="doi">10.1371/journal.pcbi.1003818.g002</object-id><label>Figure 2</label><caption>
<title>Broad peak with different fitness landscapes.</title>
<p>For the broad peak there is a specific sequence, and all sequences that are within Hamming distance <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e082" xlink:type="simple"/></inline-formula> are part of the target set. The fitness landscape is flat outside the broad peak. (A) If the width of the broad peak is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e083" xlink:type="simple"/></inline-formula>, then the expected discovery time is exponential in sequence length, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e084" xlink:type="simple"/></inline-formula>. (B) If the width of the broad peak is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e085" xlink:type="simple"/></inline-formula>, then the expected discovery time is polynomial in sequence length, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e086" xlink:type="simple"/></inline-formula>. (C) Numerical calculations for broad peak fitness landscapes. We observe exponential expected discovery time for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e087" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e088" xlink:type="simple"/></inline-formula>, whereas polynomial expected discovery time for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e089" xlink:type="simple"/></inline-formula>.</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pcbi.1003818.g002" position="float" xlink:type="simple"/></fig></sec><sec id="s2a2">
<title>Remark 1</title>
<p><italic>We highlight two important aspects of our results.</italic></p>
<list list-type="order"><list-item>
<p><italic>First, when we establish exponential lower bounds for the expected discovery time, then these lower bounds hold even if the starting sequence is only a few steps away from the target set.</italic></p>
</list-item><list-item>
<p><italic>Second, we present strong dichotomy results, and derive mathematically the most precise and strongest form of the boundary condition.</italic></p>
</list-item></list>
<p>Let us now give a numerical example to demonstrate that exponential time is intractable. Bacterial life on earth has been around for at least 3.5 billion years, which correspond to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e090" xlink:type="simple"/></inline-formula> hours. Assuming fast bacterial cell division of 20–30 minutes on average we have at most <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e091" xlink:type="simple"/></inline-formula> generations. The expected discovery time for a sequence of length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e092" xlink:type="simple"/></inline-formula> with a very large broad peak of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e093" xlink:type="simple"/></inline-formula> is approximately <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e094" xlink:type="simple"/></inline-formula> generations; see <xref ref-type="table" rid="pcbi-1003818-t001">Table 1</xref>.</p>
<table-wrap id="pcbi-1003818-t001" position="float"><object-id pub-id-type="doi">10.1371/journal.pcbi.1003818.t001</object-id><label>Table 1</label><caption>
<title>Numerical data for discovery time in flat fitness landscapes.</title>
</caption><alternatives><graphic id="pcbi-1003818-t001-1" position="float" mimetype="image" xlink:href="info:doi/10.1371/journal.pcbi.1003818.t001" xlink:type="simple"/>
<table><colgroup span="1"><col align="left" span="1"/><col align="center" span="1"/><col align="center" span="1"/><col align="center" span="1"/></colgroup>
<thead>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e095" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e096" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e097" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e098" xlink:type="simple"/></inline-formula></td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e099" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e100" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e101" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e102" xlink:type="simple"/></inline-formula></td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e103" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e104" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e105" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e106" xlink:type="simple"/></inline-formula></td>
</tr>
</tbody>
</table>
</alternatives><table-wrap-foot><fn id="nt101"><label/><p>Numerical data for the discovery time of broad peaks with width <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e107" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e108" xlink:type="simple"/></inline-formula> embedded in flat fitness landscapes. First the discovery time is computed for small values of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e109" xlink:type="simple"/></inline-formula> as shown in <xref ref-type="fig" rid="pcbi-1003818-g002">Figure 2(C)</xref>. Then the exponential growth is extrapolated to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e110" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e111" xlink:type="simple"/></inline-formula>, respectively. We show the discovery times for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e112" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e113" xlink:type="simple"/></inline-formula>. For <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e114" xlink:type="simple"/></inline-formula> the values are polynomial in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e115" xlink:type="simple"/></inline-formula>.</p></fn></table-wrap-foot></table-wrap>
<p>If individual evolutionary processes cannot find targets in polynomial time, then perhaps the success of evolution is based on the fact that many populations are searching independently and in parallel for a particular adaptation. We prove that multiple, independent parallel searches are not the solution of the problem, if the starting sequence is far away from the target center. Formally we show the following result.</p>
</sec><sec id="s2a3">
<title>Theorem 2</title>
<p><italic>In all cases where the lower bound on the expected discovery time is exponential, for all polynomials </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e116" xlink:type="simple"/></inline-formula><italic>, </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e117" xlink:type="simple"/></inline-formula><italic> and </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e118" xlink:type="simple"/></inline-formula><italic>, for any starting sequence with Hamming distance at least </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e119" xlink:type="simple"/></inline-formula><italic> from the target center, the probability for any one out of </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e120" xlink:type="simple"/></inline-formula><italic> independent multiple searches to reach the target set within </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e121" xlink:type="simple"/></inline-formula><italic> steps is at most </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e122" xlink:type="simple"/></inline-formula><italic>.</italic></p>
<p>If an evolutionary process takes exponential time, then polynomially many independent searches do not find the target in polynomial time with reasonable probability (for details see Theorem S5 in the <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>). We also show an informal and approximate calculation of the success probability for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e123" xlink:type="simple"/></inline-formula> independent searches, as follows: if the expected discovery time is exponential (say, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e124" xlink:type="simple"/></inline-formula>), then the probability that all <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e125" xlink:type="simple"/></inline-formula> independent searches fail upto <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e126" xlink:type="simple"/></inline-formula> steps is at least <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e127" xlink:type="simple"/></inline-formula> (i.e., the success probability within <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e128" xlink:type="simple"/></inline-formula> steps of any of the searches is at most <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e129" xlink:type="simple"/></inline-formula>), when the starting sequence is far away from the target center. In such a case, one could quickly exhaust the physical resources of an entire planet. The estimated number of bacterial cells <xref ref-type="bibr" rid="pcbi.1003818-Whitman1">[36]</xref> on earth is about <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e130" xlink:type="simple"/></inline-formula>. To give a specific example let us assume that there are <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e131" xlink:type="simple"/></inline-formula> independent searches, each with population size <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e132" xlink:type="simple"/></inline-formula>. The probability that at least one of those independent searches succeeds within <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e133" xlink:type="simple"/></inline-formula> generations for sequence length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e134" xlink:type="simple"/></inline-formula> and broad peak of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e135" xlink:type="simple"/></inline-formula> is less than <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e136" xlink:type="simple"/></inline-formula>.</p>
<p>In our basic model, individual mutants are evaluated one at a time. The situation of many mutant lineages evolving in parallel is similar to the multiple searches described above. As we show that whenever a single search takes exponential time, multiple independent searches do not lead to polynomial time solutions, our results imply intractability for this case as well.</p>
<p>We now explore the case of multiple broad peaks that are uniformly and randomly distributed. Consider that there are <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e137" xlink:type="simple"/></inline-formula> target centers. Around each target center there is a selection gradient extending up to a distance <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e138" xlink:type="simple"/></inline-formula>. Formally we can consider any fitness function <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e139" xlink:type="simple"/></inline-formula> that assigns zero fitness to a sequence whose Hamming distance exceeds <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e140" xlink:type="simple"/></inline-formula> from all the target centers, which in particular is subsumed by considering the multiple broad peaks where around each center we consider a broad peak of target set with peak width <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e141" xlink:type="simple"/></inline-formula>. We establish the following result:</p>
</sec><sec id="s2a4">
<title>Theorem 3</title>
<p><italic>Consider a single search under the multiple broad peak fitness landscape of </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e142" xlink:type="simple"/></inline-formula><italic> target centers chosen uniformly at random, with peak width at most </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e143" xlink:type="simple"/></inline-formula><italic> for each center and </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e144" xlink:type="simple"/></inline-formula><italic>. Then with high probability, the expected discovery time of the target set is at least </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e145" xlink:type="simple"/></inline-formula><italic>.</italic></p>
<p>Whether or not the function <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e146" xlink:type="simple"/></inline-formula> is exponential in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e147" xlink:type="simple"/></inline-formula> depends on how <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e148" xlink:type="simple"/></inline-formula> changes with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e149" xlink:type="simple"/></inline-formula>. But even if we assume exponentially many broad peak centers, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e150" xlink:type="simple"/></inline-formula>, with peak width <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e151" xlink:type="simple"/></inline-formula> where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e152" xlink:type="simple"/></inline-formula>, we need not obtain polynomial time (<xref ref-type="fig" rid="pcbi-1003818-g003">Figure 3</xref> and Theorem S6 in <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>).</p>
<fig id="pcbi-1003818-g003" position="float"><object-id pub-id-type="doi">10.1371/journal.pcbi.1003818.g003</object-id><label>Figure 3</label><caption>
<title>The search for randomly, uniformly distributed targets in sequence space.</title>
<p>(A) The target set consists of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e153" xlink:type="simple"/></inline-formula> random sequences; each one of them is surrounded by a broad peak of width up to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e154" xlink:type="simple"/></inline-formula>. The figure shows a pictorial illustration where the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e155" xlink:type="simple"/></inline-formula>-dimensional sequence space is projected onto two dimensions. From a randomly chosen starting sequence outside the target set, the expected discovery time is at least <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e156" xlink:type="simple"/></inline-formula>, which can be exponential in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e157" xlink:type="simple"/></inline-formula>. (B) Computer simulations showing the average discovery time of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e158" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e159" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e160" xlink:type="simple"/></inline-formula> targets, with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e161" xlink:type="simple"/></inline-formula>. We observe exponential dependency on <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e162" xlink:type="simple"/></inline-formula>. The discovery time is averaged over 200 runs. (C) Success probability estimated as the fraction of the 200 searches that succeed in finding one of the target sequences within <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e163" xlink:type="simple"/></inline-formula> generations. The success probability drops exponentially with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e164" xlink:type="simple"/></inline-formula>. (D) Success probability as a function of time for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e165" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e166" xlink:type="simple"/></inline-formula>. (E) Discovery time for a large number of randomly generated target sequences. Either <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e167" xlink:type="simple"/></inline-formula> or <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e168" xlink:type="simple"/></inline-formula> sequences were generated. For <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e169" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e170" xlink:type="simple"/></inline-formula> the target set consists of balls of Hamming distance <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e171" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e172" xlink:type="simple"/></inline-formula> (respectively) around each sequence. The figure shows the average discovery time of 100 runs. As expected we observe that the discovery time grows exponentially with sequence length, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e173" xlink:type="simple"/></inline-formula>.</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pcbi.1003818.g003" position="float" xlink:type="simple"/></fig>
<p>It is known that recombination may accelerate evolution on certain fitness landscapes <xref ref-type="bibr" rid="pcbi.1003818-Crow1">[28]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Smith1">[37]</xref>–<xref ref-type="bibr" rid="pcbi.1003818-Park2">[39]</xref>, and recombination may also slow down evolution on other fitness landscapes <xref ref-type="bibr" rid="pcbi.1003818-deVisser1">[40]</xref>. Recombination, however, reduces the discovery time only by at most a linear factor in sequence length <xref ref-type="bibr" rid="pcbi.1003818-Crow1">[28]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Smith1">[37]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Crow2">[38]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Neher1">[41]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Weissman1">[42]</xref>. A linear or even polynomial factor improvement over an exponential function does not convert the exponential function into a polynomial one. Hence, recombination can make a significant difference only if the underlying evolutionary process without recombination already operates in polynomial time.</p>
<p>What are then adaptive problems that can be solved by evolution in polynomial time? We propose a “regeneration process”. The basic idea is that evolution can solve a new problem efficiently, if it is has solved a similar problem already. Suppose gene duplication or genome rearrangement can give rise to starting sequences that are at most <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e174" xlink:type="simple"/></inline-formula> point mutations away from the target set, where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e175" xlink:type="simple"/></inline-formula> is a number that is independent of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e176" xlink:type="simple"/></inline-formula>. It is important that starting sequences can be regenerated again and again. We prove that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e177" xlink:type="simple"/></inline-formula> many searches are sufficient in order to find the target in polynomial time with high probability (see <xref ref-type="fig" rid="pcbi-1003818-g004">Figure 4</xref> and Section 10 in <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>). The upper bound, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e178" xlink:type="simple"/></inline-formula>, holds even for neutral drift (without selection). Note that in this case, the expected discovery time for any single search is still exponential. Therefore, most of the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e179" xlink:type="simple"/></inline-formula> searches do not succeed in polynomial time; however, with high probability one of the searches succeeds in polynomial time. There are two key aspects to the “regeneration process”: (a) the starting sequence is only a small number of steps away from the target; and (b) the starting sequence can be generated repeatedly. This process enables evolution to overcome the exponential barrier. The upper bound, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e180" xlink:type="simple"/></inline-formula>, may possibly be further reduced, if selection and/or recombination are included.</p>
<fig id="pcbi-1003818-g004" position="float"><object-id pub-id-type="doi">10.1371/journal.pcbi.1003818.g004</object-id><label>Figure 4</label><caption>
<title>Regeneration process.</title>
<p>Gene duplication (or possibly some other process) generates a steady stream of starting sequences that are a constant number <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e181" xlink:type="simple"/></inline-formula> of mutations away from the target. Many searches drift away from the target, but some will succeed in polynomially many steps. We prove that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e182" xlink:type="simple"/></inline-formula> searches ensure that with high probability some search succeed in polynomially many steps.</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pcbi.1003818.g004" position="float" xlink:type="simple"/></fig></sec></sec></sec><sec id="s3">
<title>Discussion</title>
<p>The regeneration process formalizes the role of several existing ideas. First, it ties in with the proposal that gene duplications and genome rearrangements are major events leading to the emergence of new genes <xref ref-type="bibr" rid="pcbi.1003818-Ohno1">[43]</xref>. Second, evolution can be seen as a tinkerer playing around with small modifications of existing sequences rather than creating entirely new ones <xref ref-type="bibr" rid="pcbi.1003818-Jacob1">[44]</xref>. Third, the process is related to Gillespie's suggestion <xref ref-type="bibr" rid="pcbi.1003818-Gillespie1">[29]</xref> that the starting sequence for an evolutionary search must have high fitness. In our theory, proximity in fitness value is replaced by proximity in sequence space. However, our results show that proximity alone is insufficient to break the exponential barrier, and only when combined with the process of regeneration it yields polynomial discovery time with high probability. Our process can also explain the emergence of orphan genes arising from non-coding regions <xref ref-type="bibr" rid="pcbi.1003818-Tautz1">[45]</xref>. Section 12 of the <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref> discusses the connection of our approach to existing results.</p>
<p>There is one other scenario that must be mentioned. It is possible that certain biological functions are hyper-abundant in sequence space <xref ref-type="bibr" rid="pcbi.1003818-Fontana2">[21]</xref> and that a process generating a large number of random sequences will find the function with high probability. For example, Bartel &amp; Szostak <xref ref-type="bibr" rid="pcbi.1003818-Bartel1">[46]</xref> isolated a new ribozyme from a pool of about <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e183" xlink:type="simple"/></inline-formula> random sequences of length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e184" xlink:type="simple"/></inline-formula>. While such a process is conceivable for small effective sequence length, it cannot represent a general solution for large <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e185" xlink:type="simple"/></inline-formula>.</p>
<p>Our theory has clear empirical implications. The regeneration process can be tested in systems of in vitro evolution <xref ref-type="bibr" rid="pcbi.1003818-Leconte1">[47]</xref>. A starting sequence can be generated by introducing <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e186" xlink:type="simple"/></inline-formula> point mutations in a known protein encoding sequence of length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e187" xlink:type="simple"/></inline-formula>. If these point mutations destroy the function of the protein, then the expected discovery time of any one attempt to find the original sequence should be exponential in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e188" xlink:type="simple"/></inline-formula>. But only polynomially many searches in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e189" xlink:type="simple"/></inline-formula> are required to find the target with high probability in polynomially many steps. The same setup can be used to explore whether the biological function can be found elsewhere in sequence space: the evolutionary trajectory beginning with the starting sequence could discover new solutions. Our theory also highlights how important it is to explore the distribution of biological functions in sequence space both for RNA <xref ref-type="bibr" rid="pcbi.1003818-Fontana1">[20]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Fontana2">[21]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Jimenez1">[35]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Bartel1">[46]</xref> and in the protein universe <xref ref-type="bibr" rid="pcbi.1003818-Povolotskaya1">[48]</xref>.</p>
<p>In summary, we have developed a theory that allows us to estimate time scales of evolutionary trajectories. We have shown that various natural processes of evolution take exponential time as function of the sequence length, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e190" xlink:type="simple"/></inline-formula>. In some cases we have established strong dichotomy results for precise boundary conditions. We have proposed a mechanism that allows evolution in polynomial time scales. Some interesting directions of future work are as follows: (1) Consider various forms of rugged fitness landscapes and study more refined approximations as compared to the ones we consider; and then estimate the expected discovery time for the refined approximations. (2) While in this paper we characterize the difference between exponential and polynomial for the expected discovery time, more refined analysis (such as efficiency for polynomial time, like cubic vs quadratic time) for specific fitness landscapes using mechanisms like recombination is another interesting problem.</p>
</sec><sec id="s4" sec-type="materials|methods">
<title>Materials and Methods</title>
<p>Our results are based on a mathematical analysis of the underlying stochastic processes. For Markov chains on the one-dimensional grid, we describe recurrence relations for the expected hitting time and present lower and upper bounds on the expected hitting time using combinatorial analysis (see <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref> for details). We now present the basic intuitive arguments of the main results.</p>
</sec><sec id="s5">
<title/>
<sec id="s5a">
<title>Markov chain on the one-dimensional grid</title>
<p>For a single broad peak, due to symmetry we can interpret the evolutionary random walk as a Markov chain on the one-dimensional grid. A sequence of type <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e191" xlink:type="simple"/></inline-formula> is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e192" xlink:type="simple"/></inline-formula> steps away from the target, where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e193" xlink:type="simple"/></inline-formula> is the Hamming distance between this sequence and the target. The probability that a type <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e194" xlink:type="simple"/></inline-formula> sequence mutates to a type <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e195" xlink:type="simple"/></inline-formula> sequence is given by <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e196" xlink:type="simple"/></inline-formula>. The stochastic process of the evolutionary random walk is a Markov chain on the one-dimensional grid <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e197" xlink:type="simple"/></inline-formula>.</p>
</sec><sec id="s5b">
<title>The basic recurrence relation</title>
<p>Consider a Markov chain on the one-dimensional grid, and let <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e198" xlink:type="simple"/></inline-formula> denote the expected hitting time from <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e199" xlink:type="simple"/></inline-formula> to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e200" xlink:type="simple"/></inline-formula>. The general recurrence relation for the expected hitting time is as follows:<disp-formula id="pcbi.1003818.e201"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pcbi.1003818.e201" xlink:type="simple"/><label>(1)</label></disp-formula>for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e202" xlink:type="simple"/></inline-formula>, with boundary condition <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e203" xlink:type="simple"/></inline-formula>. The interpretation is as follows. Given the current state <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e204" xlink:type="simple"/></inline-formula>, if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e205" xlink:type="simple"/></inline-formula>, at least one transition will be made to a neighboring state <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e206" xlink:type="simple"/></inline-formula>, with probability <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e207" xlink:type="simple"/></inline-formula>, from which the hitting time is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e208" xlink:type="simple"/></inline-formula>.</p>
</sec><sec id="s5c">
<title>Intuition behind Theorem 1</title>
<p>Theorem 1 is derived by obtaining precise bounds for the recurrence relation of the hitting time (<xref ref-type="disp-formula" rid="pcbi.1003818.e201">Equation 1</xref>). Consider that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e209" xlink:type="simple"/></inline-formula> for all <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e210" xlink:type="simple"/></inline-formula> (i.e., progress towards state <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e211" xlink:type="simple"/></inline-formula> is always possible), as otherwise <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e212" xlink:type="simple"/></inline-formula> is never reached from <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e213" xlink:type="simple"/></inline-formula>. We show (see Lemma 2 in the <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>) that we can write <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e214" xlink:type="simple"/></inline-formula> as a sum, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e215" xlink:type="simple"/></inline-formula>, where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e216" xlink:type="simple"/></inline-formula> is the sequence defined as:<disp-formula id="pcbi.1003818.e217"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pcbi.1003818.e217" xlink:type="simple"/><label>(2)</label></disp-formula></p>
<p>The basic intuition obtained from <xref ref-type="disp-formula" rid="pcbi.1003818.e217">Equation 2</xref> is as follows: (i) If <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e218" xlink:type="simple"/></inline-formula>, for some constant <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e219" xlink:type="simple"/></inline-formula>, then the sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e220" xlink:type="simple"/></inline-formula> grows at least as fast as a geometric series with factor <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e221" xlink:type="simple"/></inline-formula>. (ii) On the other hand, if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e222" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e223" xlink:type="simple"/></inline-formula> for some constant <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e224" xlink:type="simple"/></inline-formula>, then the sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e225" xlink:type="simple"/></inline-formula> grows at most as fast as an arithmetic series with difference <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e226" xlink:type="simple"/></inline-formula>. From the above case analysis the result for Theorem 1 is obtained as follows: If <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e227" xlink:type="simple"/></inline-formula>, then for all <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e228" xlink:type="simple"/></inline-formula>, we have <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e229" xlink:type="simple"/></inline-formula> for some <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e230" xlink:type="simple"/></inline-formula>, and hence the sequence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e231" xlink:type="simple"/></inline-formula> grows geometrically for a linear length in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e232" xlink:type="simple"/></inline-formula>. Then, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e233" xlink:type="simple"/></inline-formula> for all states <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e234" xlink:type="simple"/></inline-formula> (i.e., for all sequences outside of the target set). This corresponds to case 1 of Theorem 1. On the other hand, if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e235" xlink:type="simple"/></inline-formula>, then it is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e236" xlink:type="simple"/></inline-formula>, and case 2 of Theorem 1 is derived (for details see Corollary 2 in <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>).</p>
</sec><sec id="s5d">
<title>Intuition behind Theorem 2</title>
<p>The basic intuition for the result is as follows: consider a single search for which the expected hitting time is exponential. Then for the single search the probability to succeed in polynomially many steps is negligible (as otherwise the expectation would not have been exponential). In case of independent searches, the independence ensures that the probability that all searches fail is the product of the probabilities that every single search fails. Using the above arguments we establish Theorem 2 (for details see Section 8 in <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>).</p>
</sec><sec id="s5e">
<title>Intuition behind Theorem 3</title>
<p>For this result, it is first convenient to view the evolutionary walk taking place in the sequence space of all sequences of length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e237" xlink:type="simple"/></inline-formula>, under no selection. Each sequence has <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e238" xlink:type="simple"/></inline-formula> neighbors, and considering that a point mutation happens, the transition probability to each of them is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e239" xlink:type="simple"/></inline-formula>. The underlying Markov chain due to symmetry has fast mixing time, i.e., the number of steps to converge to the stationary distribution (the mixing time) is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e240" xlink:type="simple"/></inline-formula>. Again by symmetry the stationary distribution is the uniform distribution. If <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e241" xlink:type="simple"/></inline-formula>, then from Theorem 1 we obtain that the expected time to reach a single broad peak is exponential. By union bound, if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e242" xlink:type="simple"/></inline-formula>, the probability to reach any of the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e243" xlink:type="simple"/></inline-formula> broad peaks within <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e244" xlink:type="simple"/></inline-formula> steps is negligible. Since after the first <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e245" xlink:type="simple"/></inline-formula> steps the Markov chain converges to the stationary distribution, then each step of the process can be interpreted as selection of sequences uniformly at random among all sequences. Using Hoeffding's inequality, we show that with high probability, in expectation <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pcbi.1003818.e246" xlink:type="simple"/></inline-formula> such steps are required before a sequence is found that belongs to the target set. Thus we obtain the result of Theorem 3 (for details see Section 9 in <xref ref-type="supplementary-material" rid="pcbi.1003818.s001">Text S1</xref>).</p>
</sec><sec id="s5f">
<title>Remark about techniques</title>
<p>An important aspect of our work is that we establish our results using elementary techniques for analysis of Markov chains. The use of more advanced mathematical machinery, such as martingales <xref ref-type="bibr" rid="pcbi.1003818-Williams1">[49]</xref> or drift analysis <xref ref-type="bibr" rid="pcbi.1003818-Hajek1">[50]</xref>, <xref ref-type="bibr" rid="pcbi.1003818-Lehre1">[51]</xref>, can possibly be used to derive more refined results. While in this work our goal is to distinguish between exponential and polynomial time, whether the techniques from <xref ref-type="bibr" rid="pcbi.1003818-Williams1">[49]</xref>–<xref ref-type="bibr" rid="pcbi.1003818-Lehre1">[51]</xref> can lead to a more refined characterization within polynomial time is an interesting direction for future work.</p>
</sec></sec><sec id="s6">
<title>Supporting Information</title>
<supplementary-material id="pcbi.1003818.s001" mimetype="application/pdf" xlink:href="info:doi/10.1371/journal.pcbi.1003818.s001" position="float" xlink:type="simple"><label>Text S1</label><caption>
<p>Detailed proofs for “The Time Scale of Evolutionary Innovation.”</p>
<p>(PDF)</p>
</caption></supplementary-material></sec></body>
<back>
<ack>
<p>We thank Nick Barton and Daniel Weissman for helpful discussions and pointing us to relevant literature.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="pcbi.1003818-Allwood1"><label>1</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Allwood</surname><given-names>AC</given-names></name>, <name name-style="western"><surname>Grotzinger</surname><given-names>JP</given-names></name>, <name name-style="western"><surname>Knoll</surname><given-names>AH</given-names></name>, <name name-style="western"><surname>Burch</surname><given-names>IW</given-names></name>, <name name-style="western"><surname>Anderson</surname><given-names>MS</given-names></name>, <etal>et al</etal>. (<year>2009</year>) <article-title>Controls on development and diversity of early archean stromatolites</article-title>. <source>Proc Natl Acad Sci USA</source> <volume>106</volume>: <fpage>9548</fpage>–<lpage>9555</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Schopf1"><label>2</label>
<mixed-citation publication-type="journal" xlink:type="simple"><article-title>Schopf JW (August 2006) The first billion years: When did life emerge?</article-title> <source>Elements</source> <volume>2</volume>: <fpage>229</fpage>–<lpage>233</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Kimura1"><label>3</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Kimura</surname><given-names>M</given-names></name> (<year>1968</year>) <article-title>Evolutionary rate at the molecular level</article-title>. <source>Nature</source> <volume>217</volume>: <fpage>624</fpage>–<lpage>626</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Ewens1"><label>4</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Ewens</surname><given-names>WJ</given-names></name> (<year>1967</year>) <article-title>The probability of survival of a new mutant in a uctuating environment</article-title>. <source>Heredity</source> <volume>22</volume>: <fpage>438</fpage>–<lpage>443</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Barton1"><label>5</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Barton</surname><given-names>NH</given-names></name> (<year>1995</year>) <article-title>Linkage and the limits to natural selection</article-title>. <source>Genetics</source> <volume>140</volume>: <fpage>821</fpage>–<lpage>41</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Campos1"><label>6</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Campos</surname><given-names>PR</given-names></name> (<year>2004</year>) <article-title>Fixation of beneficial mutations in the presence of epistatic interactions</article-title>. <source>Bull Math Biol</source> <volume>66</volume>: <fpage>473</fpage>–<lpage>486</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Antal1"><label>7</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Antal</surname><given-names>T</given-names></name>, <name name-style="western"><surname>Scheuring</surname><given-names>I</given-names></name> (<year>2006</year>) <article-title>Fixation of strategies for an evolutionary game in finite populations</article-title>. <source>Bull Math Biol</source> <volume>68</volume>: <fpage>1923</fpage>–<lpage>1944</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Whitlock1"><label>8</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Whitlock</surname><given-names>MC</given-names></name> (<year>2003</year>) <article-title>Fixation probability and time in subdivided populations</article-title>. <source>Genetics</source> <volume>164</volume>: <fpage>767</fpage>–<lpage>779</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Altrock1"><label>9</label>
<mixed-citation publication-type="other" xlink:type="simple">Altrock PM, Traulsen A (2009) Fixation times in evolutionary games under weak selection. New J Phys <volume>11</volume> . doi:10.1088/1367-2630/11/1/013012</mixed-citation>
</ref>
<ref id="pcbi.1003818-Kimura2"><label>10</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Kimura</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Ohta</surname><given-names>T</given-names></name> (<year>1969</year>) <article-title>Average number of generations until fixation of a mutant gene in a finite population</article-title>. <source>Genetics</source> <volume>61</volume>: <fpage>763</fpage>–<lpage>771</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Johnson1"><label>11</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Johnson</surname><given-names>T</given-names></name>, <name name-style="western"><surname>Gerrish</surname><given-names>P</given-names></name> (<year>2002</year>) <article-title>The fixation probability of a beneficial allele in a population dividing by binary fission</article-title>. <source>Genetica</source> <volume>115</volume>: <fpage>283</fpage>–<lpage>287</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Orr1"><label>12</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Orr</surname><given-names>HA</given-names></name> (<year>2000</year>) <article-title>The rate of adaptation in asexuals</article-title>. <source>Genetics</source> <volume>155</volume>: <fpage>961</fpage>–<lpage>968</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Wilke1"><label>13</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Wilke</surname><given-names>CO</given-names></name> (<year>2004</year>) <article-title>The speed of adaptation in large asexual populations</article-title>. <source>Genetics</source> <volume>167</volume>: <fpage>2045</fpage>–<lpage>2053</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Desai1"><label>14</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Desai</surname><given-names>MM</given-names></name>, <name name-style="western"><surname>Fisher</surname><given-names>DS</given-names></name>, <name name-style="western"><surname>Murray</surname><given-names>AW</given-names></name> (<year>2007</year>) <article-title>The speed of evolution and maintenance of variation in asexual populations</article-title>. <source>Curr Biol</source> <volume>17</volume>: <fpage>385</fpage>–<lpage>394</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Ohta1"><label>15</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Ohta</surname><given-names>T</given-names></name> (<year>1972</year>) <article-title>Population size and rate of evolution</article-title>. <source>J Mol Evol</source> <volume>1</volume>: <fpage>305</fpage>–<lpage>314</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Papadimitriou1"><label>16</label>
<mixed-citation publication-type="other" xlink:type="simple">Papadimitriou C (1994) Computational complexity. Addison-Wesley.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Cormen1"><label>17</label>
<mixed-citation publication-type="other" xlink:type="simple">Cormen T, Leiserson C, Rivest R, Stein C (2009) Introduction to Algorithms. MIT Press.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Valiant1"><label>18</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Valiant</surname><given-names>LG</given-names></name> (<year>2009</year>) <article-title>Evolvability</article-title>. <source>J ACM</source> 56: 3:1–3:21.</mixed-citation>
</ref>
<ref id="pcbi.1003818-MaynardSmith1"><label>19</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Maynard Smith</surname><given-names>J</given-names></name> (<year>1970</year>) <article-title>Natural selection and the concept of a protein space</article-title>. <source>Nature</source> <volume>225</volume>: <fpage>563</fpage>–<lpage>564</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Fontana1"><label>20</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Fontana</surname><given-names>W</given-names></name>, <name name-style="western"><surname>Schuster</surname><given-names>P</given-names></name> (<year>1987</year>) <article-title>A computer model of evolutionary optimization</article-title>. <source>Biophys Chem</source> <volume>26</volume>: <fpage>123</fpage>–<lpage>147</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Fontana2"><label>21</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Fontana</surname><given-names>W</given-names></name>, <name name-style="western"><surname>Schuster</surname><given-names>P</given-names></name> (<year>1998</year>) <article-title>Continuity in evolution: On the nature of transitions</article-title>. <source>Science</source> <volume>280</volume>: <fpage>1451</fpage>–<lpage>1455</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Eigen1"><label>22</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Eigen</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Mccaskill</surname><given-names>J</given-names></name>, <name name-style="western"><surname>Schuster</surname><given-names>P</given-names></name> (<year>1988</year>) <article-title>Molecular quasi-species</article-title>. <source>J Phys Chem</source> <volume>92</volume>: <fpage>6881</fpage>–<lpage>6891</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Eigen2"><label>23</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Eigen</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Schuster</surname><given-names>P</given-names></name> (<year>1978</year>) <article-title>The hypercycle</article-title>. <source>Naturwissenschaften</source> <volume>65</volume>: <fpage>7</fpage>–<lpage>41</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Park1"><label>24</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Park</surname><given-names>SC</given-names></name>, <name name-style="western"><surname>Simon</surname><given-names>D</given-names></name>, <name name-style="western"><surname>Krug</surname><given-names>J</given-names></name> (<year>2010</year>) <article-title>The speed of evolution in large asexual populations</article-title>. <source>J Stat Phys</source> <volume>138</volume>: <fpage>381</fpage>–<lpage>410</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Derrida1"><label>25</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Derrida</surname><given-names>B</given-names></name>, <name name-style="western"><surname>Peliti</surname><given-names>L</given-names></name> (<year>1991</year>) <article-title>Evolution in a at fitness landscape</article-title>. <source>Bull Math Biol</source> <volume>53</volume>: <fpage>355</fpage>–<lpage>382</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Stadler1"><label>26</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Stadler</surname><given-names>PF</given-names></name> (<year>2002</year>) <article-title>Fitness landscapes</article-title>. <source>Appl Math &amp; Comput</source> <volume>117</volume>: <fpage>187</fpage>–<lpage>207</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Worden1"><label>27</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Worden</surname><given-names>RP</given-names></name> (<year>1995</year>) <article-title>A speed limit for evolution</article-title>. <source>J Theor Biol</source> <volume>176</volume>: <fpage>137</fpage>–<lpage>152</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Crow1"><label>28</label>
<mixed-citation publication-type="other" xlink:type="simple">Crow JF, Kimura M (1965) Evolution in sexual and asexual populations. Am Nat <volume>99</volume> : pp. 439–450.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Gillespie1"><label>29</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Gillespie</surname><given-names>JH</given-names></name> (<year>1984</year>) <article-title>Molecular evolution over the mutational landscape</article-title>. <source>Evolution</source> <volume>38</volume>: <fpage>1116</fpage>–<lpage>1129</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Kauffman1"><label>30</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Kauffman</surname><given-names>S</given-names></name>, <name name-style="western"><surname>Levin</surname><given-names>S</given-names></name> (<year>1987</year>) <article-title>Towards a general theory of adaptive walks on rugged landscapes</article-title>. <source>Journal of Theoretical Biology</source> <volume>128</volume>: <fpage>11</fpage>–<lpage>45</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Orr2"><label>31</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Orr</surname><given-names>HA</given-names></name> (<year>2000</year>) <article-title>A minimum on the mean number of steps taken in adaptive walks</article-title>. <source>Journal of Theoretical Biology</source> <volume>220</volume>: <fpage>241</fpage>–<lpage>247</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Weinreich1"><label>32</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Weinreich</surname><given-names>DM</given-names></name>, <name name-style="western"><surname>Watson</surname><given-names>RA</given-names></name>, <name name-style="western"><surname>Chao</surname><given-names>L</given-names></name> (<year>2005</year>) <article-title>Perspective:sign epistasis and genetic constraint on evolutionary trajectories</article-title>. <source>Evolution</source> <volume>59</volume>: <fpage>1165</fpage>–<lpage>1174</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Poelwijk1"><label>33</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Poelwijk</surname><given-names>FJ</given-names></name>, <name name-style="western"><surname>Kiviet</surname><given-names>DJ</given-names></name>, <name name-style="western"><surname>Weinreich</surname><given-names>DM</given-names></name>, <name name-style="western"><surname>Tans</surname><given-names>SJ</given-names></name> (<year>2007</year>) <article-title>Empirical fitness landscapes reveal accessible evolutionary paths</article-title>. <source>Nature</source> <volume>445</volume>: <fpage>383</fpage>–<lpage>386</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Woods1"><label>34</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Woods</surname><given-names>RJ</given-names></name>, <name name-style="western"><surname>Barrick</surname><given-names>JE</given-names></name>, <name name-style="western"><surname>Cooper</surname><given-names>TF</given-names></name>, <name name-style="western"><surname>Shrestha</surname><given-names>U</given-names></name>, <name name-style="western"><surname>Kauth</surname><given-names>MR</given-names></name>, <etal>et al</etal>. (<year>2011</year>) <article-title>Second-order selection for evolvability in a large escherichia coli population</article-title>. <source>Science</source> <volume>331</volume>: <fpage>1433</fpage>–<lpage>1436</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Jimenez1"><label>35</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Jimenez</surname><given-names>JI</given-names></name>, <name name-style="western"><surname>Xulvi-Brunet</surname><given-names>R</given-names></name>, <name name-style="western"><surname>Campbell</surname><given-names>GW</given-names></name>, <name name-style="western"><surname>Turk-MacLeod</surname><given-names>R</given-names></name>, <name name-style="western"><surname>Chen</surname><given-names>IA</given-names></name> (<year>2013</year>) <article-title>Comprehensive experimental fitness landscape and evolutionary network for small rna</article-title>. <source>Proc Natl Acad Sci USA</source>. <volume>110(37)</volume>: <fpage>14984</fpage>–<lpage>9</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Whitman1"><label>36</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Whitman</surname><given-names>WB</given-names></name>, <name name-style="western"><surname>Coleman</surname><given-names>DC</given-names></name>, <name name-style="western"><surname>Wiebe</surname><given-names>WJ</given-names></name> (<year>1998</year>) <article-title>Prokaryotes: The unseen majority</article-title>. <source>Proc Natl Acad Sci USA</source> <volume>95</volume>: <fpage>6578</fpage>–<lpage>6583</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Smith1"><label>37</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Smith</surname><given-names>JM</given-names></name> (<year>1974</year>) <article-title>Recombination and the rate of evolution</article-title>. <source>Genetics</source> <volume>78</volume>: <fpage>299</fpage>–<lpage>304</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Crow2"><label>38</label>
<mixed-citation publication-type="other" xlink:type="simple">Crow JF, Kimura M (1970) An introduction to population genetics theory. Burgess Publishing Company.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Park2"><label>39</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Park</surname><given-names>SC</given-names></name>, <name name-style="western"><surname>Krug</surname><given-names>J</given-names></name> (<year>2013</year>) <article-title>Rate of adaptation in sexuals and asexuals: A solvable model of the fishermuller effect</article-title>. <source>Genetics</source> <volume>195</volume>: <fpage>941</fpage>–<lpage>955</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-deVisser1"><label>40</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>de Visser</surname><given-names>JAGM</given-names></name>, <name name-style="western"><surname>Park</surname><given-names>S</given-names></name>, <name name-style="western"><surname>Krug</surname><given-names>J</given-names></name> (<year>2009</year>) <article-title>Exploring the effect of sex on empirical fitness landscapes</article-title>. <source>The American Naturalist</source> <volume>174</volume>: <fpage>S15</fpage>–<lpage>S30</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Neher1"><label>41</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Neher</surname><given-names>RA</given-names></name>, <name name-style="western"><surname>Shraiman</surname><given-names>BI</given-names></name>, <name name-style="western"><surname>Fisher</surname><given-names>DS</given-names></name> (<year>2010</year>) <article-title>Rate of adaptation in large sexual populations</article-title>. <source>Genetics</source> <volume>184</volume>: <fpage>467</fpage>–<lpage>481</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Weissman1"><label>42</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Weissman</surname><given-names>DB</given-names></name>, <name name-style="western"><surname>Hallatschek</surname><given-names>O</given-names></name> (<year>2014</year>) <article-title>The rate of adaptation in large sexual populations with linear chromosomes</article-title>. <source>Genetics</source> <volume>196</volume>: <fpage>1167</fpage>–<lpage>1183</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Ohno1"><label>43</label>
<mixed-citation publication-type="other" xlink:type="simple">Ohno S (1970) Evolution by gene duplication. Springer-Verlag.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Jacob1"><label>44</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Jacob</surname><given-names>F</given-names></name> (<year>1977</year>) <article-title>Evolution and tinkering</article-title>. <source>Science</source> <volume>196</volume>: <fpage>1161</fpage>–<lpage>1166</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Tautz1"><label>45</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Tautz</surname><given-names>D</given-names></name>, <name name-style="western"><surname>Domazet-Lošo</surname><given-names>T</given-names></name> (<year>2011</year>) <article-title>The evolutionary origin of orphan genes</article-title>. <source>Nat Rev Genet</source> <volume>12</volume>: <fpage>692</fpage>–<lpage>702</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Bartel1"><label>46</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Bartel</surname><given-names>D</given-names></name>, <name name-style="western"><surname>Szostak</surname><given-names>J</given-names></name> (<year>1993</year>) <article-title>Isolation of new ribozymes from a large pool of random sequences</article-title>. <source>Science</source> <volume>261</volume>: <fpage>1411</fpage>–<lpage>1418</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Leconte1"><label>47</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Leconte</surname><given-names>AM</given-names></name>, <name name-style="western"><surname>Dickinson</surname><given-names>BC</given-names></name>, <name name-style="western"><surname>Yang</surname><given-names>DD</given-names></name>, <name name-style="western"><surname>Chen</surname><given-names>IA</given-names></name>, <name name-style="western"><surname>Allen</surname><given-names>B</given-names></name>, <etal>et al</etal>. (<year>2013</year>) <article-title>A population-based experimental model for protein evolution: Effects of mutation rate and selection stringency on evolutionary outcomes</article-title>. <source>Biochemistry</source> <volume>52</volume>: <fpage>1490</fpage>–<lpage>1499</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Povolotskaya1"><label>48</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Povolotskaya</surname><given-names>IS</given-names></name>, <name name-style="western"><surname>Kondrashov</surname><given-names>FA</given-names></name> (<year>2010</year>) <article-title>Sequence space and the ongoing expansion of the protein universe</article-title>. <source>Nature</source> <volume>465</volume>: <fpage>922</fpage>–<lpage>926</lpage>.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Williams1"><label>49</label>
<mixed-citation publication-type="other" xlink:type="simple">Williams D (1991) Probability with Martingales. Cambridge mathematical textbooks. Cambridge University Press.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Hajek1"><label>50</label>
<mixed-citation publication-type="other" xlink:type="simple">Hajek B (1982) Hitting-time and occupation-time bounds implied by drift analysis with applications. Advances in Applied Probability <volume>14</volume> : pp. 502–525.</mixed-citation>
</ref>
<ref id="pcbi.1003818-Lehre1"><label>51</label>
<mixed-citation publication-type="other" xlink:type="simple">Lehre PK, Witt C (2013) General drift analysis with tail bounds. CoRR abs/1307.2559.</mixed-citation>
</ref>
</ref-list></back>
</article>