<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="EN">
  <front>
    <journal-meta><journal-id journal-id-type="nlm-ta">PLoS ONE</journal-id><journal-id journal-id-type="publisher-id">plos</journal-id><journal-id journal-id-type="pmc">plosone</journal-id><!--===== Grouping journal title elements =====--><journal-title-group><journal-title>PLoS ONE</journal-title></journal-title-group><issn pub-type="epub">1932-6203</issn><publisher>
        <publisher-name>Public Library of Science</publisher-name>
        <publisher-loc>San Francisco, USA</publisher-loc>
      </publisher></journal-meta>
    <article-meta><article-id pub-id-type="publisher-id">09-PONE-RA-10011</article-id><article-id pub-id-type="doi">10.1371/journal.pone.0020660</article-id><article-categories>
        <subj-group subj-group-type="heading">
          <subject>Research Article</subject>
        </subj-group>
        <subj-group subj-group-type="Discipline-v2">
          <subject>Biology</subject>
          <subj-group>
            <subject>Computational biology</subject>
            <subj-group>
              <subject>Genomics</subject>
            </subj-group>
          </subj-group>
          <subj-group>
            <subject>Genetics</subject>
            <subj-group>
              <subject>Genetic mutation</subject>
            </subj-group>
          </subj-group>
          <subj-group>
            <subject>Genomics</subject>
            <subj-group>
              <subject>Comparative genomics</subject>
              <subject>Genome evolution</subject>
            </subj-group>
          </subj-group>
        </subj-group>
        <subj-group subj-group-type="Discipline">
          <subject>Genetics and Genomics</subject>
          <subject>Computational Biology</subject>
        </subj-group>
      </article-categories><title-group><article-title>SNPs Occur in Regions with Less Genomic Sequence Conservation</article-title><alt-title alt-title-type="running-head">SNPs and Conservation</alt-title></title-group><contrib-group>
        <contrib contrib-type="author" xlink:type="simple">
          <name name-style="western">
            <surname>Castle</surname>
            <given-names>John C.</given-names>
          </name>
          <xref ref-type="aff" rid="aff1"/>
          <xref ref-type="corresp" rid="cor1">
            <sup>*</sup>
          </xref>
          <xref ref-type="fn" rid="fn1">
            <sup>¤</sup>
          </xref>
        </contrib>
      </contrib-group><aff id="aff1">          <addr-line>Rosetta Inpharmatics LLC, a wholly owned subsidiary of Merck &amp; Co., Inc., Seattle, Washington, United States of America</addr-line>       </aff><contrib-group>
        <contrib contrib-type="editor" xlink:type="simple">
          <name name-style="western">
            <surname>Ruvinsky</surname>
            <given-names>Ilya</given-names>
          </name>
          <role>Editor</role>
          <xref ref-type="aff" rid="edit1"/>
        </contrib>
      </contrib-group><aff id="edit1">The University of Chicago, United States of America</aff><author-notes>
        <corresp id="cor1">* E-mail: <email xlink:type="simple">john.castle@tron-mainz.de</email></corresp>
        <fn fn-type="con">
          <p>Conceived and designed the experiments: JCC. Performed the experiments: JCC. Analyzed the data: JCC. Wrote the paper: JCC.</p>
        </fn>
        <fn fn-type="current-aff" id="fn1">
          <label>¤</label>
          <p>Current address: Computational Medicine &amp; Medical Genomics, TRON - Translational Oncology at the Johannes Gutenberg University of Mainz Medicine, Mainz, Rheinland Palatine, Germany</p>
        </fn>
      <fn fn-type="conflict">
        <p>The author has declared that no competing interests exist.</p>
      </fn></author-notes><pub-date pub-type="collection">
        <year>2011</year>
      </pub-date><pub-date pub-type="epub">
        <day>6</day>
        <month>6</month>
        <year>2011</year>
      </pub-date><volume>6</volume><issue>6</issue><elocation-id>e20660</elocation-id><history>
        <date date-type="received">
          <day>6</day>
          <month>1</month>
          <year>2011</year>
        </date>
        <date date-type="accepted">
          <day>6</day>
          <month>5</month>
          <year>2011</year>
        </date>
      </history><!--===== Grouping copyright info into permissions =====--><permissions><copyright-year>2011</copyright-year><copyright-holder>John C. Castle</copyright-holder><license><license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</license-p></license></permissions><abstract>
        <p>Rates of SNPs (single nucleotide polymorphisms) and cross-species genomic sequence conservation reflect intra- and inter-species variation, respectively. Here, I report SNP rates and genomic sequence conservation adjacent to mRNA processing regions and show that, as expected, more SNPs occur in less conserved regions and that functional regions have fewer SNPs. <xref ref-type="sec" rid="s2">Results</xref> are confirmed using both mouse and human data. Regions include protein start codons, 3′ splice sites, 5′ splice sites, protein stop codons, predicted miRNA binding sites, and polyadenylation sites. Throughout, SNP rates are lower and conservation is higher at regulatory sites. Within coding regions, SNP rates are highest and conservation is lowest at codon position three and the fewest SNPs are found at codon position two, reflecting codon degeneracy for amino acid encoding. Exon splice sites show high conservation and very low SNP rates, reflecting both splicing signals and protein coding. Relaxed constraint on the codon third position is dramatically seen when separating exonic SNP rates based on intron phase. At polyadenylation sites, a peak of conservation and low SNP rate occurs from 30 to 17 nt preceding the site. This region is highly enriched for the sequence AAUAAA, reflecting the location of the conserved polyA signal. miRNA 3′ UTR target sites are predicted incorporating interspecies genomic sequence conservation; SNP rates are low in these sites, again showing fewer SNPs in conserved regions. Together, these results confirm that SNPs, reflecting recent genetic variation, occur more frequently in regions with less evolutionarily conservation.</p>
      </abstract><funding-group><funding-statement>During the data analysis and writing of this manuscript, the author was a paid employee of Rosetta Inpharmatics and TRON gGmbH. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</funding-statement></funding-group><counts>
        <page-count count="12"/>
      </counts></article-meta>
  </front>
  <body>
    <sec id="s1">
      <title>Introduction</title>
      <p>Single nucleotide polymorphisms (SNPs) are intra-species sequence variation. Conversely, cross-species genomic conservation reflects longer term inter-species evolution. SNPs mark genomic locations where intra-species variability is permissible; fewer SNPs should occur in regions encoding sequence-dependent functions. Similarly, insofar that many functional molecular processes are sequence dependent and behave similarly across species, higher genome sequence conservation should occur in functional regions. Therefore, SNP occurrence rates and genomic conservation should be anti-correlated but similarly delineate functional regions. Indeed, there is an emerging field to actively integrate intra- and inter-genomic variation, with applications ranging from disease to phenotype to functional molecular biology to evolution to crop breeding to conservation <xref ref-type="bibr" rid="pone.0020660-Durbin1">[1]</xref>, <xref ref-type="bibr" rid="pone.0020660-Morin1">[2]</xref>, <xref ref-type="bibr" rid="pone.0020660-Miller1">[3]</xref>, <xref ref-type="bibr" rid="pone.0020660-Allendorf1">[4]</xref>, <xref ref-type="bibr" rid="pone.0020660-Stapley1">[5]</xref>, <xref ref-type="bibr" rid="pone.0020660-Chasman1">[6]</xref>.</p>
      <p>Each step in mRNA-processing relies on specific sequences. After transcription, the spliceosome binds RNA to splice exons into a mature transcript; sequence at the transcript is recognized and polyadenylated. The ribosome scans the mature transcript, starting protein translation at a start codon and ending translation at a stop codon. One pathway for mRNA degradation involves miRNA recognition of sequences in the mRNA 3′ untranslated region (3′ UTR).</p>
      <p>The human and mouse genomes were assembled in 2001 <xref ref-type="bibr" rid="pone.0020660-Lander1">[7]</xref> <xref ref-type="bibr" rid="pone.0020660-Venter1">[8]</xref> and 2002 <xref ref-type="bibr" rid="pone.0020660-Waterston1">[9]</xref>, respectively. The University of California Santa Cruz genome databases <xref ref-type="bibr" rid="pone.0020660-Kuhn1">[10]</xref> provide access to many genome-wide resources, including the genomic locations of mouse and human SNPs and a nucleotide-by-nucleotide cross-species conservation score. The databases also lists the locations of predicted miRNA target locations and human transcription factor binding sites and polyadenylation sites, along with genomic alignments of protein-coding RefSeq (NM) transcripts, including start codons, splice sites, and stop codons.</p>
      <p>Here, I integrate these genomic coordinates to determine SNP rates and conservation scores surrounding mRNA processing elements. I expected that, for example, SNPs would occur more frequently in codon position three, the degenerate position, and that conservation would be higher at this position. The results indeed confirm expectations and clearly show that SNP rates and conservation are anti-correlated, that both mark functional elements, and demonstrate the single-nucleotide, sequence-level genetic constraints imposed by protein-coding, by splicing, and by transcript processing.</p>
    </sec>
    <sec id="s2">
      <title>Results</title>
      <p>Using mouse and human data available in the UCSC genome databases, I calculate SNP rates and conservation scores at single nucleotide resolution across six mRNA processing regions, for both species (<xref ref-type="sec" rid="s4">Methods</xref>). The SNP rate is the percentage of nucleotides at a given relative position that overlap a SNP. The conservation score at a given relative position is the average inter-species genome conservation score, where the score is taken from the Vertebrate Multiz Alignment &amp; PhastCons Conservation (28 Species) table <xref ref-type="bibr" rid="pone.0020660-Siepel1">[11]</xref>. Summary diagrams for mouse start, stop, 3′, and 5′ splice sites are shown in <xref ref-type="fig" rid="pone-0020660-g001">Figure 1</xref>.</p>
      <fig id="pone-0020660-g001" position="float">
        <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g001</object-id>
        <label>Figure 1</label>
        <caption>
          <title>The regions examined, including, protein start sites, 3′ splice sites (3′SS), 5′ splice sites (5′SS), protein stop sites, predicted miRNA binding sites, and polyadenylation sites.</title>
          <p><xref ref-type="sec" rid="s2">Results</xref> for mouse start and stop (middle) and splice sites (lower) are shown.</p>
        </caption>
        <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g001" xlink:type="simple"/>
      </fig>
      <sec id="s2a">
        <title>Start codon (<xref ref-type="fig" rid="pone-0020660-g002">Figure 2</xref>)</title>
        <fig id="pone-0020660-g002" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g002</object-id>
          <label>Figure 2</label>
          <caption>
            <title>The SNP rate (top) and cross-species conservation (bottom) across mouse protein start sites.</title>
            <p>Symbol color and shape indicate codon position.</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g002" xlink:type="simple"/>
        </fig>
        <p>Within the region adjacent mouse protein coding start sites, the SNP rate is lowest and the conservation highest at the start codon. A phase three periodicity in SNP rate and conservation levels occurs after the start codon, with more frequent SNPs and less conservation at codon position three. Within each codon, the majority of positions two, and not positions one, show the lowest SNP rate. After peaking at the start codon to near 0.7, conservation stays high in the coding region at 0.6. The SNP rate is on average lower in positions one and two after the start codon, while the SNP rate in position three is similar before and after the start codon.</p>
      </sec>
      <sec id="s2b">
        <title>3′ splice site, phase 0 introns (<xref ref-type="fig" rid="pone-0020660-g003">Figure 3</xref>)</title>
        <fig id="pone-0020660-g003" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g003</object-id>
          <label>Figure 3</label>
          <caption>
            <title>The SNP rate (top) and cross-species conservation (bottom) across mouse 3′ splice sites.</title>
            <p>Only internal exons longer than 60 nt with phase 0 introns longer than 300 nt were included. Symbol color and shape indicate codon position.</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g003" xlink:type="simple"/>
        </fig>
        <p><xref ref-type="fig" rid="pone-0020660-g003">Figure 3</xref> shows the SNP rate and conservation at mouse 3′ splice sites. Introns can be classified according to where they occur relative to the intertwined codons. Introns of phase zero, one, and two represent whether the intron occurs between two codons, after the first nucleotide of a codon, or after the second nucleotide of a codon, respectively. <xref ref-type="fig" rid="pone-0020660-g003">Figure 3</xref> shows results for 3′ splice sites preceded by phase 0 introns. Additionally, for consistency, only exons that are longer than 60 nt and entirely protein coding are included. The SNP rate is lowest and conservation peaks at the splice site AG nucleotides. After the splice site, both the SNP rate and conservation show a clear phase three periodicity. The SNP rate is highest at codon position three. For every codon, the SNP rate is lowest at position 2, lower than the rate for position 1. Conservation stays constant in the exon at over 0.8.</p>
      </sec>
      <sec id="s2c">
        <title>5′ splice site, phase 0 introns (<xref ref-type="fig" rid="pone-0020660-g004">Figure 4</xref>)</title>
        <fig id="pone-0020660-g004" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g004</object-id>
          <label>Figure 4</label>
          <caption>
            <title>The SNP rate (top) and cross-species conservation (bottom) across mouse 5′ splice sites.</title>
            <p>Only internal exons longer than 60 nt with phase 0 introns longer than 300 nt were included. Symbol color and shape indicate codon position.</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g004" xlink:type="simple"/>
        </fig>
        <p><xref ref-type="fig" rid="pone-0020660-g004">Figure 4</xref> shows results for mouse 5′ splice sites preceding phase 0 introns, using protein coding exons greater than 60 nt long. The SNP rate is lowest and the conservation highest at the splice site GT nucleotides. In the exonic region before the splice site, there is a clear phase three periodicity in both SNP rate and conservation. The SNP rate is highest at codon position three and lowest at position two. The exonic SNP rate at codon positions two and three is low throughout the exon while the position three SNP rate is similar to the intronic rate. Conservation in the exon is high at over 0.8, peaks at the splice site, rapidly decreases, and levels after 15 nt into the intron.</p>
      </sec>
      <sec id="s2d">
        <title>Splice sites, all intron phases (<xref ref-type="fig" rid="pone-0020660-g005">Figure 5</xref>)</title>
        <fig id="pone-0020660-g005" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g005</object-id>
          <label>Figure 5</label>
          <caption>
            <title>The SNP rate across mouse 3′ (top) and 5′ (bottom) splice sites.</title>
            <p>Only internal exons longer than 60 nt with introns longer than 300 nt were included. Symbol color and shape indicate codon position. Intron phase is marked by line color. Web-logos were generated using sequences from the phase 0 introns.</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g005" xlink:type="simple"/>
        </fig>
        <p>The set of mouse RefSeq transcripts contain 50,779, 40,890, and 22,078 phase zero, one, and two introns, respectively, adjacent protein coding exons greater than 60 nt long. <xref ref-type="fig" rid="pone-0020660-g005">Figure 5</xref> shows the SNP rate at 3′ and 5′ splice sites, color coded based on the phase of preceding/following intron. The phase three periodicity in the protein coding region is obvious for all cases, and shifts according to the intron phase. The most SNPs occur at position three and the least at position two. Within intronic regions, all profiles are very similar and no phase three periodicity exists, reflecting the lack of protein coding constraint. Extracting the nucleotide sequences to examine intra-genome variation, I used WebLogo <xref ref-type="bibr" rid="pone.0020660-Crooks1">[12]</xref> to examine the sequence content and found the expected splice site motifs. Interestingly, at position −4 preceding the 3′ splice site (middle plot), the logo shows no nucleotide preference and, correspondingly, the SNP rate at this specific position is higher than adjacent positions. At position 5 following the 5′ splice site (lower plot), G is the preferred nucleotide; the SNP rate at position 5 is lower than at the neighboring positions.</p>
      </sec>
      <sec id="s2e">
        <title>Human Splice sites, phase 0 introns (<xref ref-type="fig" rid="pone-0020660-g006">Figure 6</xref>)</title>
        <fig id="pone-0020660-g006" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g006</object-id>
          <label>Figure 6</label>
          <caption>
            <title>The SNP rate across human 3′ (left) and 5′ (right) splice sites.</title>
            <p>Only internal exons longer than 60 nt with phase 0 introns longer than 300 nt were included. All SNPs (top), validated SNPs (middle), and 1000 Genomes identified SNPs (bottom) were considered.</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g006" xlink:type="simple"/>
        </fig>
        <p>The corresponding human genomic regions show results similar to the mouse results (<xref ref-type="fig" rid="pone-0020660-g006">Figure 6</xref>). As for mouse, only protein coding exons longer than 60 nt neighboring introns greater than 300 nt long were used. Additionally, three sets of SNPs were examined: all SNPs from dbSNP build 130 <xref ref-type="bibr" rid="pone.0020660-Sherry1">[13]</xref>; the subset of validated SNPs; and the subset found in the 1000 Genomes Project <xref ref-type="bibr" rid="pone.0020660-Pennisi1">[14]</xref>. The results are similar to the mouse results, including the lowest SNP rates at the splice sites and the period three ringing in protein coding regions. Intriguingly, the main difference between results from the three sets of SNPs is the average difference between exonic and intronic SNP rates, where the SNP rate at exonic codon positions one and two from the validated and 1000 Genomes SNPs is much lower than the intronic rate.</p>
      </sec>
      <sec id="s2f">
        <title>Stop codon (<xref ref-type="fig" rid="pone-0020660-g007">Figure 7</xref>)</title>
        <fig id="pone-0020660-g007" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g007</object-id>
          <label>Figure 7</label>
          <caption>
            <title>The SNP rate (top) and cross-species conservation (bottom) across mouse protein stop sites.</title>
            <p>Symbol color and shape indicate codon position.</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g007" xlink:type="simple"/>
        </fig>
        <p>Near mouse stop codons, the SNP rate is lower and conservation higher before the stop codon. The SNP rate and conservation score display a phase three periodicity before the stop codon. Position three shows a greater SNP frequency and lower conservation; position two shows the lowest SNP rate. The region with both the highest SNP rate and least conservation occurs from 5 to 20 nt after the stop codon, which is found in both human and mouse, suggesting that this region is under less evolutionary control.</p>
      </sec>
      <sec id="s2g">
        <title>Predicted miRNA binding sites (<xref ref-type="fig" rid="pone-0020660-g008">Figure 8</xref>)</title>
        <fig id="pone-0020660-g008" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g008</object-id>
          <label>Figure 8</label>
          <caption>
            <title>The SNP rate (top) and cross-species conservation (bottom) across predicted mouse miRNA binding sites.</title>
            <p>Symbol color and shape indicate codon position. The vertical axis extends to 0.2.</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g008" xlink:type="simple"/>
        </fig>
        <p>TargetScan <xref ref-type="bibr" rid="pone.0020660-Lewis1">[15]</xref> predicts mouse miRNA 3′ UTR binding sites using genomic conservation and thus genomic conservation should be high at predicted binding locations (<xref ref-type="fig" rid="pone-0020660-g008">Figure 8</xref>). I find that the SNP rate is lowest within the predicted miRNA binding locations. The lowest SNP rate occurs across a 5 nt window residing within a larger 8 nt window. The colored symbols in the plot, which marked codon position in previous plots, signify the distance from the 3′ edge, modulus 3, of the miRNA binding site (there are no codons here). As expected, no phase-three periodicity exists in either SNP rate or conservation.</p>
      </sec>
      <sec id="s2h">
        <title>Transcription termination (<xref ref-type="fig" rid="pone-0020660-g009">Figure 9</xref>) and human polyadenylation sites (<xref ref-type="fig" rid="pone-0020660-g010">Figure 10</xref>)</title>
        <fig id="pone-0020660-g009" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g009</object-id>
          <label>Figure 9</label>
          <caption>
            <title>The SNP rate (top) and cross-species conservation (bottom) across mouse transcript 3′ ends.</title>
            <p>Symbol color and shape indicate codon position.</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g009" xlink:type="simple"/>
        </fig>
        <fig id="pone-0020660-g010" position="float">
          <object-id pub-id-type="doi">10.1371/journal.pone.0020660.g010</object-id>
          <label>Figure 10</label>
          <caption>
            <title>The SNP rate (top) and cross-species conservation (bottom) across human polyadenylation sites.</title>
            <p>The percentage of sites with the A[A/U]UAAA motif is inset, including the distribution (solid bars) and cumulative level (line).</p>
          </caption>
          <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pone.0020660.g010" xlink:type="simple"/>
        </fig>
        <p>The UCSC genome databases list transcript 3′ ends for mouse transcripts and predicted polyadenylation sites for human transcripts <xref ref-type="bibr" rid="pone.0020660-Zhang1">[16]</xref> (<xref ref-type="fig" rid="pone-0020660-g009">Figures 9</xref> and <xref ref-type="fig" rid="pone-0020660-g010">10</xref>). Again, positions 1, 2, and 3 in the plot signify the distance, modulus 3, from the sites (there are no codons here) and, as expected, no phase-three periodicity exists in either SNP rate or conservation. The SNP rate shows a local minimum and the conservation a peak at 20 nt before the polyA site/transcript ends. Before this point, the human 3′ UTR shows higher conservation relative to the region after the polyA site. Beyond the conservation peak, conservation falls but with a 20 nt plateau after the polyA site, followed by a further decrease into the intergenic region (<xref ref-type="sec" rid="s3">Discussion</xref>).</p>
        <p>Using the predicted human polyA site coordinates, I extracted the sequence from the UCSC genome databases and searched for the polyA signal motifs AAUAAA and AUUAAA. The motif occurs between 30 and 17 nt before the polyA site (<xref ref-type="fig" rid="pone-0020660-g010">Figure 10</xref> inset) in the majority of the sequences. Thus, the conservation peak at 20 nt before polyA sites likely marks the location of the polyA signal.</p>
      </sec>
    </sec>
    <sec id="s3">
      <title>Discussion</title>
      <p>Here, I used mouse and human genomic resources to examine SNP rates and interspecies genomic sequence conservation in RNA-relevant regions. The results confirm expectations that the SNP rate and conservation score show anti-correlation: fewer SNPs occur in conserved regions. This is in agreement with our understanding that both processes mark regions under evolutionary conservation: SNPs reflect allowed intra-species genome variation and cross-species conservation demarcates long-term evolutionary stability.</p>
      <p><xref ref-type="sec" rid="s2">Results</xref> from human and mouse datasets are very similar. One difference is that when all human SNPs are considered, the SNP rate in introns is surprisingly lower than in exons (<xref ref-type="fig" rid="pone-0020660-g006">Figure 6</xref>, top), However, examination of only high quality SNPs shows SNP rates lower in exons, as expected (<xref ref-type="fig" rid="pone-0020660-g006">Figure 6</xref>, middle and bottom). SNPs discovery experiments are often biased based on genomic location: many previous and upcoming studies concentrate on variations (and mutations) in known coding exonic regions <xref ref-type="bibr" rid="pone.0020660-Stratton1">[17]</xref>, further biasing our knowledge of SNPs to those occurring in mRNAs. Thus the appearance of higher SNP rates in exons relative to introns may be a simple sampling bias.</p>
      <p>Additional biases may result from the genome of the laboratory mouse genome. Inbred laboratory mouse strains have a genome that is mosaic of two sub-species <xref ref-type="bibr" rid="pone.0020660-Wade1">[18]</xref>. The polymorphisms contained in dbSNP are a broad and diverse collection of variations submitted from many contributors. Thus a given SNP may represent either an inter-species variation (e.g., within <italic>Mus musculus domesticus</italic>) or a variation between sub-species (e.g. between <italic>M. m. domesticus</italic> and <italic>M. m. musculus</italic>). The latter represent variation somewhat between a true intra-species polymorphism and cross-species variation.</p>
      <p>Furthermore, changes in mutation rate variation could introduce biases in these observations (e.g., <xref ref-type="bibr" rid="pone.0020660-Lynch1">[19]</xref>). The data could be skewed if, for example, the genome of one species experienced an extreme mutation rate, biasing the genomic conservation; by the erroneous submission to dbSNP of somatic mutations from cancerous cell lines, impacting SNP rates; or by variation in evolution across a genome, such as from variable GC content <xref ref-type="bibr" rid="pone.0020660-Baer1">[20]</xref>, <xref ref-type="bibr" rid="pone.0020660-Chamary1">[21]</xref>. The biases likely exist; however, the aggregated results presented here are likely robust to these biases.</p>
      <p>Predicted miRNA binding sites <xref ref-type="bibr" rid="pone.0020660-Lewis1">[15]</xref> incorporate cross-species conservation, along with location relative to transcripts (e.g., in 3′ UTRs) and the underlying nucleic acid sequence, and thus genomic conservation is automatically high. However, as the prediction algorithms do not take SNPs into account, it is satisfying to observe that the SNP rate falls within the predicted mouse binding locations, as has been previously observed for human miRNA binding sites <xref ref-type="bibr" rid="pone.0020660-Chen1">[22]</xref>, <xref ref-type="bibr" rid="pone.0020660-Saunders1">[23]</xref>. Indeed, the low SNP rate in these conserved locations further demonstrates that SNPs and interspecies genomic sequence conservation are not independent processes; rather, SNPs occur in less conserved regions. Similarly, SNP rates and conservation around predicted human transcription factor binding sites (TFBSs) shows similar results: SNP rates are low in predicted TFBSs.</p>
      <p>The profiles of conservation scores and composite SNP rates delineate established functional elements. Conservation can be used to identify individual elements whereas SNP rates can identify elements when consolidated across the genome. The sharp peak in conservation and the SNP rate trough occurring 20 nt before the polyadenylation site identifies a region highly enriched for the polyA signal AA[U/A]AAA. Motif enrichment, conservation, and the low SNP rate all occur over a narrow range from 30 to 17 nt preceding the polyadenylation site. Variations near 5′ and 3′ splice sites may impact pre-mRNA splicing and can cause disease <xref ref-type="bibr" rid="pone.0020660-Cartegni1">[24]</xref>; correspondingly, high conservation and low SNP rates occur here. A lower SNP rate and higher conservation occurs from +/− 6 nt adjacent the splice sites, similar to previous findings that suggest that this is due to the occurrence of splicing signals <xref ref-type="bibr" rid="pone.0020660-Fairbrother1">[25]</xref>.</p>
      <p>Low genomic variability exists in protein coding regions: SNP rates are lower and conservation higher in all protein coding regions examined. Near the protein start site, conservation is sharply higher precisely at the start site and maintains a high level into the protein coding region. Similar profiles, but in reverse, occur at the protein stop site. However, while the stop codon shows a low SNP rate, conservation at the stop site codon itself is not higher than at previous positions, unlike the spike at the start codon. This may reflect either the stop codon degeneracy or that a mutated stop codon will likely be followed by a second stop codon. Furthermore, compared to the conservation in the internal protein coding exons (e.g., <xref ref-type="fig" rid="pone-0020660-g003">Figure 3</xref>), the conservation is lower in the coding regions adjacent the start and the stop codons. Speculatively, this may be due either to a relative decrease in functional importance at the protein N- and C-terminus (unlikely) or simply that our knowledge of the start and stop sites is more uncertain than the coordinates of internal exons, which are clearly defined by bounding AG-GT splice sites.</p>
      <p>Additionally, a 15 nt trough in conservation and corresponding high in the SNP rate occur from 5 to 20 nt after the stop site in both mouse and human results. That it occurs in both SNPs and conservation suggests that it is real and, taken at face value, shows that more changes have occurred in the region between the stop codon and 20 nt into the 3′ UTR. Previous work has shown that miRNA binding sites at positioned within the 3′UTR occur at least 15 nt from the stop codon <xref ref-type="bibr" rid="pone.0020660-Grimson1">[26]</xref>, suggesting that the increase in sequence variability in this zone is a result of fewer miRNA binding sites.</p>
      <p>Inter-species (conservation), intra-species (SNPs), and intra-genome (webLogos) data all demonstrate the low variation at splice sites. The SNP rate increases almost 3-fold from within the coding exons to the rate at 100-nt into introns. However, the low variability occurs not only in protein-coding regions but also into the adjacent intron, accentuating the need to study variation beyond non-synonymous changes (e.g., <xref ref-type="bibr" rid="pone.0020660-Chamary1">[21]</xref>). Indeed, the highest conservation and lowest SNP rate occur outside of the protein-coding exons, at the splice sites. This is in line with expectations that while polymorphisms in protein-coding regions could impact amino-acid selection or codon usage, polymorphism and mutations in splicing motifs could deregulate or cause the skipping of entire exons.</p>
      <p>Finally, all protein coding regions display a phase three periodicity with increased variability at position three. The periodicity is not observed outside of the protein coding regions, such as adjacent polyadenylation sites. That the SNP rates are highest at codon position three is undoubtedly a result of the degeneracy for amino acid coding at the third position (19 of 22 amino acids are degenerate at codon position 3). Moreover, codon degeneracy is slightly higher at codon position one relative to position two: I correspondingly find a lower SNP rate at the second position. Thus, more SNPs occur in degenerate positions for amino acid encoding, and are thus less likely to have a functional impact <xref ref-type="bibr" rid="pone.0020660-Chasman1">[6]</xref>.</p>
      <p>Together, these findings reinforce our understanding of evolution and functional elements in the genome. These include that SNPs occur more frequently in less conserved regions; that both genomic conservation and aggregated SNPs rates can be used to identify functional elements; and that codon position three in protein coding regions is more degenerate.</p>
    </sec>
    <sec id="s4" sec-type="methods">
      <title>Methods</title>
      <p>Genomic coordinates were downloaded from UCSC databases <xref ref-type="bibr" rid="pone.0020660-Kuhn1">[10]</xref>, including human dbSNP build 130 and mouse dbSNP build 128 <xref ref-type="bibr" rid="pone.0020660-Sherry1">[13]</xref>, human predicted transcription factor binding sites (TFBS) from the HMR Conserved Transcription Factor Binding Sites table <xref ref-type="bibr" rid="pone.0020660-Matys1">[27]</xref>, mouse and human predicted miRNA target sites from the TargetScan table <xref ref-type="bibr" rid="pone.0020660-Lewis1">[15]</xref>, reported human polyadenylation sites <xref ref-type="bibr" rid="pone.0020660-Zhang1">[16]</xref>, alternative transcriptional element alignments (derived in the UCSC databases from the unpublished txgAnalyse and txGraph programs written by Jim Kent at UCSC), and RefSeq NM transcript and EST sequence alignments. I accessed the nucleotide-specific cross-species conservation scores from the Vertebrate Multiz Alignment &amp; PhastCons Conservation (28 Species) table <xref ref-type="bibr" rid="pone.0020660-Siepel1">[11]</xref>. As mentioned, the predictions of both TFBS and miRNA binding sites explicitly use cross-species genome conservation as criteria. Polyadenylation site predictions incorporate the coordinates of 3′ ESTs, including the presence of polyA tails. Validated human SNPs are those containing the word “by” in the description (e.g., “by-hapmap”) and “1000 Genome” SNPs are those listing “by-1000genomes”. The SNPs identified in the 1000 Genomes Project are not error-free; however, I expect no systematic error that would impact the presented aggregated SNP rates. Transcript alignments were performed by UCSC using BLAT.</p>
      <p>For the human predicted polyadenylation sites, I selected only those polyadenylation sites falling within 100 nucleotides of the 3′ terminus of a RefSeq transcript. This excluded alternative polyadenylation sites occurring in the interior of 3′ UTRs and also allowed a strand assignment to each polyadenylation site, based on the RefSeq transcript orientation.</p>
      <p>I used RefSeq “NM” transcript alignments which include exon boundaries, transcription start and stop coordinates, coding start and stop coordinates, and intron phase. I required at least 60 nt of coding sequence after and before start and stop codons, respectively. For the 5′ and 3′ splice sites, I included only internal exons longer than 60 nt surrounded by introns greater than 300 nt long.</p>
      <p>I examined several mRNA processing regions (<xref ref-type="fig" rid="pone-0020660-g001">Figure 1</xref>): protein start codons, internal 3′ splice sites, internal 5′ splice sites, protein stop codons, predicted miRNA binding sites, predicted polyadenylation sites (human), transcript 3′ termini (mouse), and predicted transcription factor binding sites (human), For each region, the algorithm looped over each individual site, counting SNPs and the level of conservation at each position between −500 and +500 nt around the site. Reported values reflect the percentage (SNP rate) and average (genomic conservation) at each position relative to the specific site location. The UCSC genome database also includes the underlying nucleotide sequence, which was used for the WebLogos.</p>
    </sec>
  </body>
  <back>
    <ack>
      <p>This work was enabled by the UCSC genomic databases and the willingness of investigators worldwide to contribute data to these databases. I thank those parties. Feedback from Michael Koslowski, Chris Armour, Chris Raymond, Carol Rohl, Jason Johnson, and Will Fairbrother was appreciated.</p>
    </ack>
    <ref-list>
      <title>References</title>
      <ref id="pone.0020660-Durbin1">
        <label>1</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Durbin</surname><given-names>RM</given-names></name><name name-style="western"><surname>Abecasis</surname><given-names>GR</given-names></name><name name-style="western"><surname>Altshuler</surname><given-names>DL</given-names></name><name name-style="western"><surname>Auton</surname><given-names>A</given-names></name><name name-style="western"><surname>Brooks</surname><given-names>LD</given-names></name><etal/></person-group>             <year>2010</year>             <article-title>A map of human genome variation from population-scale sequencing.</article-title>             <source>Nature</source>             <volume>467</volume>             <fpage>1061</fpage>             <lpage>1073</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Morin1">
        <label>2</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Morin</surname><given-names>PA</given-names></name><name name-style="western"><surname>Luikart</surname><given-names>G</given-names></name><name name-style="western"><surname>Wayne</surname><given-names>RK</given-names></name></person-group>             <collab xlink:type="simple">group Sw</collab>             <year>2004</year>             <article-title>SNPs in ecology, evolution and conservation.</article-title>             <source>TRENDS in Ecology and Evolution</source>             <volume>19</volume>             <fpage>208</fpage>             <lpage>216</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Miller1">
        <label>3</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Miller</surname><given-names>MP</given-names></name><name name-style="western"><surname>Kumar</surname><given-names>S</given-names></name></person-group>             <year>2001</year>             <article-title>Understanding human disease mutations through the use of interspecific genetic variation.</article-title>             <source>Hum Mol Genet</source>             <volume>10</volume>             <fpage>2319</fpage>             <lpage>2328</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Allendorf1">
        <label>4</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Allendorf</surname><given-names>FW</given-names></name><name name-style="western"><surname>Hohenlohe</surname><given-names>PA</given-names></name><name name-style="western"><surname>Luikart</surname><given-names>G</given-names></name></person-group>             <year>2010</year>             <article-title>Genomics and the future of conservation genetics.</article-title>             <source>Nat Rev Genet</source>             <volume>11</volume>             <fpage>697</fpage>             <lpage>709</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Stapley1">
        <label>5</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Stapley</surname><given-names>J</given-names></name><name name-style="western"><surname>Reger</surname><given-names>J</given-names></name><name name-style="western"><surname>Feulner</surname><given-names>PG</given-names></name><name name-style="western"><surname>Smadja</surname><given-names>C</given-names></name><name name-style="western"><surname>Galindo</surname><given-names>J</given-names></name><etal/></person-group>             <year>2010</year>             <article-title>Adaptation genomics: the next generation.</article-title>             <source>Trends Ecol Evol</source>             <volume>25</volume>             <fpage>705</fpage>             <lpage>712</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Chasman1">
        <label>6</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Chasman</surname><given-names>D</given-names></name><name name-style="western"><surname>Adams</surname><given-names>RM</given-names></name></person-group>             <year>2001</year>             <article-title>Predicting the functional consequences of non-synonymous single nucleotide polymorphisms: structure-based assessment of amino acid variation.</article-title>             <source>J Mol Biol</source>             <volume>307</volume>             <fpage>683</fpage>             <lpage>706</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Lander1">
        <label>7</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Lander</surname><given-names>ES</given-names></name><name name-style="western"><surname>Linton</surname><given-names>LM</given-names></name><name name-style="western"><surname>Birren</surname><given-names>B</given-names></name><name name-style="western"><surname>Nusbaum</surname><given-names>C</given-names></name><name name-style="western"><surname>Zody</surname><given-names>MC</given-names></name><etal/></person-group>             <year>2001</year>             <article-title>Initial sequencing and analysis of the human genome.</article-title>             <source>Nature</source>             <volume>409</volume>             <fpage>860</fpage>             <lpage>921</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Venter1">
        <label>8</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Venter</surname><given-names>JC</given-names></name><name name-style="western"><surname>Adams</surname><given-names>MD</given-names></name><name name-style="western"><surname>Myers</surname><given-names>EW</given-names></name><name name-style="western"><surname>Li</surname><given-names>PW</given-names></name><name name-style="western"><surname>Mural</surname><given-names>RJ</given-names></name><etal/></person-group>             <year>2001</year>             <article-title>The sequence of the human genome.</article-title>             <source>Science</source>             <volume>291</volume>             <fpage>1304</fpage>             <lpage>1351</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Waterston1">
        <label>9</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Waterston</surname><given-names>RH</given-names></name><name name-style="western"><surname>Lindblad-Toh</surname><given-names>K</given-names></name><name name-style="western"><surname>Birney</surname><given-names>E</given-names></name><name name-style="western"><surname>Rogers</surname><given-names>J</given-names></name><name name-style="western"><surname>Abril</surname><given-names>JF</given-names></name><etal/></person-group>             <year>2002</year>             <article-title>Initial sequencing and comparative analysis of the mouse genome.</article-title>             <source>Nature</source>             <volume>420</volume>             <fpage>520</fpage>             <lpage>562</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Kuhn1">
        <label>10</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Kuhn</surname><given-names>RM</given-names></name><name name-style="western"><surname>Karolchik</surname><given-names>D</given-names></name><name name-style="western"><surname>Zweig</surname><given-names>AS</given-names></name><name name-style="western"><surname>Wang</surname><given-names>T</given-names></name><name name-style="western"><surname>Smith</surname><given-names>KE</given-names></name><etal/></person-group>             <year>2009</year>             <article-title>The UCSC Genome Browser Database: update 2009.</article-title>             <source>Nucleic Acids Res</source>             <volume>37</volume>             <fpage>D755</fpage>             <lpage>761</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Siepel1">
        <label>11</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Siepel</surname><given-names>A</given-names></name><name name-style="western"><surname>Bejerano</surname><given-names>G</given-names></name><name name-style="western"><surname>Pedersen</surname><given-names>JS</given-names></name><name name-style="western"><surname>Hinrichs</surname><given-names>AS</given-names></name><name name-style="western"><surname>Hou</surname><given-names>M</given-names></name><etal/></person-group>             <year>2005</year>             <article-title>Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.</article-title>             <source>Genome Res</source>             <volume>15</volume>             <fpage>1034</fpage>             <lpage>1050</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Crooks1">
        <label>12</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Crooks</surname><given-names>GE</given-names></name><name name-style="western"><surname>Hon</surname><given-names>G</given-names></name><name name-style="western"><surname>Chandonia</surname><given-names>JM</given-names></name><name name-style="western"><surname>Brenner</surname><given-names>SE</given-names></name></person-group>             <year>2004</year>             <article-title>WebLogo: a sequence logo generator.</article-title>             <source>Genome Res</source>             <volume>14</volume>             <fpage>1188</fpage>             <lpage>1190</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Sherry1">
        <label>13</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Sherry</surname><given-names>ST</given-names></name><name name-style="western"><surname>Ward</surname><given-names>MH</given-names></name><name name-style="western"><surname>Kholodov</surname><given-names>M</given-names></name><name name-style="western"><surname>Baker</surname><given-names>J</given-names></name><name name-style="western"><surname>Phan</surname><given-names>L</given-names></name><etal/></person-group>             <year>2001</year>             <article-title>dbSNP: the NCBI database of genetic variation.</article-title>             <source>Nucleic Acids Res</source>             <volume>29</volume>             <fpage>308</fpage>             <lpage>311</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Pennisi1">
        <label>14</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Pennisi</surname><given-names>E</given-names></name></person-group>             <year>2010</year>             <article-title>Genomics. 1000 Genomes Project gives new map of genetic diversity.</article-title>             <source>Science</source>             <volume>330</volume>             <fpage>574</fpage>             <lpage>575</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Lewis1">
        <label>15</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Lewis</surname><given-names>BP</given-names></name><name name-style="western"><surname>Shih</surname><given-names>IH</given-names></name><name name-style="western"><surname>Jones-Rhoades</surname><given-names>MW</given-names></name><name name-style="western"><surname>Bartel</surname><given-names>DP</given-names></name><name name-style="western"><surname>Burge</surname><given-names>CB</given-names></name></person-group>             <year>2003</year>             <article-title>Prediction of mammalian microRNA targets.</article-title>             <source>Cell</source>             <volume>115</volume>             <fpage>787</fpage>             <lpage>798</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Zhang1">
        <label>16</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Zhang</surname><given-names>H</given-names></name><name name-style="western"><surname>Hu</surname><given-names>J</given-names></name><name name-style="western"><surname>Recce</surname><given-names>M</given-names></name><name name-style="western"><surname>Tian</surname><given-names>B</given-names></name></person-group>             <year>2005</year>             <article-title>PolyA_DB: a database for mammalian mRNA polyadenylation.</article-title>             <source>Nucleic Acids Res</source>             <volume>33</volume>             <fpage>D116</fpage>             <lpage>120</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Stratton1">
        <label>17</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Stratton</surname><given-names>M</given-names></name></person-group>             <year>2008</year>             <article-title>Genome resequencing and genetic variation.</article-title>             <source>Nat Biotechnol</source>             <volume>26</volume>             <fpage>65</fpage>             <lpage>66</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Wade1">
        <label>18</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Wade</surname><given-names>CM</given-names></name><name name-style="western"><surname>Kulbokas</surname><given-names>EJ</given-names><suffix>3rd</suffix></name><name name-style="western"><surname>Kirby</surname><given-names>AW</given-names></name><name name-style="western"><surname>Zody</surname><given-names>MC</given-names></name><name name-style="western"><surname>Mullikin</surname><given-names>JC</given-names></name><etal/></person-group>             <year>2002</year>             <article-title>The mosaic structure of variation in the laboratory mouse genome.</article-title>             <source>Nature</source>             <volume>420</volume>             <fpage>574</fpage>             <lpage>578</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Lynch1">
        <label>19</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Lynch</surname><given-names>M</given-names></name></person-group>             <year>2010</year>             <article-title>Evolution of the mutation rate.</article-title>             <source>Trends in genetics: TIG</source>             <volume>26</volume>             <fpage>345</fpage>             <lpage>352</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Baer1">
        <label>20</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Baer</surname><given-names>CF</given-names></name><name name-style="western"><surname>Miyamoto</surname><given-names>MM</given-names></name><name name-style="western"><surname>Denver</surname><given-names>DR</given-names></name></person-group>             <year>2007</year>             <article-title>Mutation rate variation in multicellular eukaryotes: causes and consequences.</article-title>             <source>Nature reviews Genetics</source>             <volume>8</volume>             <fpage>619</fpage>             <lpage>631</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Chamary1">
        <label>21</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Chamary</surname><given-names>JV</given-names></name><name name-style="western"><surname>Parmley</surname><given-names>JL</given-names></name><name name-style="western"><surname>Hurst</surname><given-names>LD</given-names></name></person-group>             <year>2006</year>             <article-title>Hearing silence: non-neutral evolution at synonymous sites in mammals.</article-title>             <source>Nature reviews Genetics</source>             <volume>7</volume>             <fpage>98</fpage>             <lpage>108</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Chen1">
        <label>22</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Chen</surname><given-names>K</given-names></name><name name-style="western"><surname>Rajewsky</surname><given-names>N</given-names></name></person-group>             <year>2006</year>             <article-title>Natural selection on human microRNA binding sites inferred from SNP data.</article-title>             <source>Nat Genet</source>             <volume>38</volume>             <fpage>1452</fpage>             <lpage>1456</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Saunders1">
        <label>23</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Saunders</surname><given-names>MA</given-names></name><name name-style="western"><surname>Liang</surname><given-names>H</given-names></name><name name-style="western"><surname>Li</surname><given-names>WH</given-names></name></person-group>             <year>2007</year>             <article-title>Human polymorphism at microRNAs and microRNA target sites.</article-title>             <source>Proc Natl Acad Sci U S A</source>             <volume>104</volume>             <fpage>3300</fpage>             <lpage>3305</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Cartegni1">
        <label>24</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Cartegni</surname><given-names>L</given-names></name><name name-style="western"><surname>Chew</surname><given-names>SL</given-names></name><name name-style="western"><surname>Krainer</surname><given-names>AR</given-names></name></person-group>             <year>2002</year>             <article-title>Listening to silence and understanding nonsense: exonic mutations that affect splicing.</article-title>             <source>Nat Rev Genet</source>             <volume>3</volume>             <fpage>285</fpage>             <lpage>298</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Fairbrother1">
        <label>25</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Fairbrother</surname><given-names>WG</given-names></name><name name-style="western"><surname>Holste</surname><given-names>D</given-names></name><name name-style="western"><surname>Burge</surname><given-names>CB</given-names></name><name name-style="western"><surname>Sharp</surname><given-names>PA</given-names></name></person-group>             <year>2004</year>             <article-title>Single nucleotide polymorphism-based validation of exonic splicing enhancers.</article-title>             <source>PLoS Biol</source>             <volume>2</volume>             <fpage>E268</fpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Grimson1">
        <label>26</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Grimson</surname><given-names>A</given-names></name><name name-style="western"><surname>Farh</surname><given-names>KK</given-names></name><name name-style="western"><surname>Johnston</surname><given-names>WK</given-names></name><name name-style="western"><surname>Garrett-Engele</surname><given-names>P</given-names></name><name name-style="western"><surname>Lim</surname><given-names>LP</given-names></name><etal/></person-group>             <year>2007</year>             <article-title>MicroRNA targeting specificity in mammals: determinants beyond seed pairing.</article-title>             <source>Mol Cell</source>             <volume>27</volume>             <fpage>91</fpage>             <lpage>105</lpage>          </element-citation>
      </ref>
      <ref id="pone.0020660-Matys1">
        <label>27</label>
        <element-citation publication-type="journal" xlink:type="simple">             <person-group person-group-type="author"><name name-style="western"><surname>Matys</surname><given-names>V</given-names></name><name name-style="western"><surname>Kel-Margoulis</surname><given-names>OV</given-names></name><name name-style="western"><surname>Fricke</surname><given-names>E</given-names></name><name name-style="western"><surname>Liebich</surname><given-names>I</given-names></name><name name-style="western"><surname>Land</surname><given-names>S</given-names></name><etal/></person-group>             <year>2006</year>             <article-title>TRANSFAC and its module TRANSCompel: transcriptional gene regulation in eukaryotes.</article-title>             <source>Nucleic Acids Res</source>             <volume>34</volume>             <fpage>D108</fpage>             <lpage>110</lpage>          </element-citation>
      </ref>
    </ref-list>
    
  </back>
</article>