<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="nlm-ta">PLoS ONE</journal-id>
      <journal-id journal-id-type="publisher-id">plos</journal-id>
      <journal-id journal-id-type="pmc">plosone</journal-id>
      <journal-title-group>
        <journal-title>PLoS ONE</journal-title>
      </journal-title-group>
      <issn pub-type="epub">1932-6203</issn>
      <publisher>
        <publisher-name>Public Library of Science</publisher-name>
        <publisher-loc>San Francisco, USA</publisher-loc>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="publisher-id">PONE-D-12-13817</article-id>
      <article-id pub-id-type="doi">10.1371/journal.pone.0048638</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Research Article</subject>
        </subj-group>
        <subj-group subj-group-type="Discipline-v2">
          <subject>Biology</subject>
          <subj-group>
            <subject>Computational biology</subject>
            <subj-group>
              <subject>Evolutionary modeling</subject>
              <subject>Population genetics</subject>
            </subj-group>
          </subj-group>
          <subj-group>
            <subject>Evolutionary biology</subject>
            <subj-group>
              <subject>Evolutionary processes</subject>
              <subject>Population genetics</subject>
            </subj-group>
          </subj-group>
          <subj-group>
            <subject>Genetics</subject>
            <subj-group>
              <subject>Human genetics</subject>
              <subject>Population genetics</subject>
            </subj-group>
          </subj-group>
          <subj-group>
            <subject>Population biology</subject>
            <subj-group>
              <subject>Population genetics</subject>
            </subj-group>
          </subj-group>
        </subj-group>
        <subj-group subj-group-type="Discipline-v2">
          <subject>Mathematics</subject>
          <subj-group>
            <subject>Statistics</subject>
            <subj-group>
              <subject>Biostatistics</subject>
            </subj-group>
          </subj-group>
        </subj-group>
        <subj-group subj-group-type="Discipline">
          <subject>Genetics and Genomics</subject>
          <subject>Computational Biology</subject>
          <subject>Evolutionary Biology</subject>
          <subject>Mathematics</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Towards Improvements in the Estimation of the Coalescent: Implications for the Most Effective Use of Y Chromosome Short Tandem Repeat Mutation Rates</article-title>
        <alt-title alt-title-type="running-head">Improvements in the Estimation of the Coalescent</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" xlink:type="simple">
          <name name-style="western">
            <surname>Bird</surname>
            <given-names>Steven C.</given-names>
          </name>
          <xref ref-type="aff" rid="aff1"/>
          <xref ref-type="corresp" rid="cor1">
            <sup>*</sup>
          </xref>
        </contrib>
      </contrib-group>
      <aff id="aff1">
        <addr-line>Department of Biology, Texas State University, San Marcos, Texas, United States of America</addr-line>
      </aff>
      <contrib-group>
        <contrib contrib-type="editor" xlink:type="simple">
          <name name-style="western">
            <surname>Johnson</surname>
            <given-names>Norman</given-names>
          </name>
          <role>Editor</role>
          <xref ref-type="aff" rid="edit1"/>
        </contrib>
      </contrib-group>
      <aff id="edit1">
        <addr-line>University of Massachusetts, United States of America</addr-line>
      </aff>
      <author-notes>
        <corresp id="cor1">* E-mail: <email xlink:type="simple">sb1611@txstate.edu</email></corresp>
        <fn fn-type="conflict">
          <p>The author has declared that no competing interests exist.</p>
        </fn>
        <fn fn-type="con">
          <p>Conceived and designed the experiments: SCB. Performed the experiments: SCB. Analyzed the data: SCB. Contributed reagents/materials/analysis tools: SCB. Wrote the paper: SCB.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="collection">
        <year>2012</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>31</day>
        <month>10</month>
        <year>2012</year>
      </pub-date>
      <volume>7</volume>
      <issue>10</issue>
      <elocation-id>e48638</elocation-id>
      <history>
        <date date-type="received">
          <day>14</day>
          <month>5</month>
          <year>2012</year>
        </date>
        <date date-type="accepted">
          <day>1</day>
          <month>10</month>
          <year>2012</year>
        </date>
      </history>
      <permissions>
        <copyright-year>2012</copyright-year>
        <copyright-holder>Steven C</copyright-holder>
        <license xlink:type="simple">
          <license-p>Bird. This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</license-p>
        </license>
      </permissions>
      <abstract>
        <p>Over the past two decades, many short tandem repeat (STR) microsatellite loci on the human Y chromosome have been identified together with mutation rate estimates for the individual loci. These have been used to estimate the coalescent age, or the time to the most recent common ancestor (TMRCA) expressed in generations, in conjunction with the average square difference measure (ASD), an unbiased point estimator of TMRCA based upon the average within-locus allele variance between haplotypes. The ASD estimator, in turn, depends on accurate mutation rate estimates to be able to produce good approximations of the coalescent age of a sample. Here, a comparison is made between three published sets of per locus mutation rate estimates as they are applied to the calculation of the coalescent age for real and simulated population samples. A novel evaluation method is developed for estimating the degree of conformity of any Y chromosome STR locus of interest to the strict stepwise mutation model and specific recommendations are made regarding the suitability of thirty-two commonly used Y-STR loci for the purpose of estimating the coalescent. The use of the geometric mean for averaging ASD and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e001" xlink:type="simple"/></inline-formula> across loci is shown to improve the consistency of the resulting estimates, with decreased sensitivity to outliers and to the number of STR loci compared or the particular set of mutation rates selected.</p>
      </abstract>
      <funding-group>
        <funding-statement>The author has no support or funding to report.</funding-statement>
      </funding-group>
      <counts>
        <page-count count="11"/>
      </counts>
    </article-meta>
  </front>
  <body>
    <sec id="s1">
      <title>Introduction</title>
      <p>Two types of genetic markers found on the human Y chromosome are used extensively in the fields of forensics, genetic anthropology, population genetics and for identification of kinship among individuals; the first is the single nucleotide polymorphism (SNP) and the second is the short tandem repeat (STR), also called the microsatellite <xref ref-type="bibr" rid="pone.0048638-Jobling1">[1]</xref>. SNPs consist of random single point mutations within an individual genome that are passed down to the descendants of the first individual to acquire the specific SNP. SNPs are generally assumed to mutate approximately according to the Infinite Alleles Model (IAM), where each mutation event is unique and independent of all other SNPs <xref ref-type="bibr" rid="pone.0048638-Walsh1">[2]</xref>. The IAM was determined to be a reasonably accurate model of SNP mutation for either the Y chromosome or for mitochondrial DNA unless a very small effective (breeding) population size (N<sub>e</sub>&lt;200) was present <xref ref-type="bibr" rid="pone.0048638-Walsh1">[2]</xref>.</p>
      <p>The other type of genetic marker, the STR, consists of short, repeated segments of non-coding DNA (also called microsatellites) that increase or decrease in length by one or more sets of repeats (ex. GATA/GATA/GATA to GATA/GATA/GATA/GATA, etc.) which are typically characterized by the number of repeats that appear within a given chromosomal segment site <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>, <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>. These loci are classified according to the length of the relevant microsatellite repeat sequence, i.e., dinucleotides (AT/AT/AT etc.) trinucleotides (AAT/AAT/AAT etc.), tetranucleotides (GGAT/GGAT/GGAT etc.), pentanucleotides, and hexanucleotides. STRs are assumed, generally, to mutate according to a different model, the Stepwise Mutation Model (SMM) <xref ref-type="bibr" rid="pone.0048638-Ohta1">[5]</xref>, <xref ref-type="bibr" rid="pone.0048638-Kimura1">[6]</xref>. A full mathematical treatment of both mutation models, as they applied to the Y chromosome, was made by <xref ref-type="bibr" rid="pone.0048638-Walsh1">[2]</xref>.</p>
      <p>Human Y-SNPs have a low mutation rate, approximately 2.0E-8 <xref ref-type="bibr" rid="pone.0048638-Nachman1">[7]</xref>, but Y-STRs mutate much more quickly, on the order of 1.0E-4 to 1.0E-2, with each STR having its own mutation rate <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>, <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>. Because they are treated as unique evolutionary polymorphisms (UEPs), SNPs may be used to differentiate divisions (clades, also called haplogroups) of the overall human phylogenetic tree, particularly those SNPs associated with the two uniparentally inherited chromosomes, the Y chromosome (passed from the father only to his male offspring) <xref ref-type="bibr" rid="pone.0048638-Jobling1">[1]</xref> and the mitochondrial DNA chromosome (passed from the mother to all of her offspring, male or female) <xref ref-type="bibr" rid="pone.0048638-Behar1">[8]</xref>.</p>
      <p>Differences between human mitochondrial DNA chromosomes are defined solely by UEPs (SNPs) while the Y chromosome features both SNPs and STRs, allowing a more detailed resolution of the genetic history of a given individual or population of males <xref ref-type="bibr" rid="pone.0048638-Jobling1">[1]</xref>. Genetic profiles formed through the aggregation of various allele repeat values (e.g, 13-24-13-10-11-12-13-30 for eight STR loci) at multiple Y-STR locus sites across an individual’s Y chromosome are called haplotypes and are in 100% linkage disequilibrium with each other, due to non-recombination of the SRY region during meiosis; these allele values therefore are subject to change only when mutation occurs at a given individual STR locus <xref ref-type="bibr" rid="pone.0048638-Jobling1">[1]</xref>. The combination of SNPs and STRs provides a powerful tool for delineating male population substructure; male individuals who share the same Y-SNP also must share a common male lineal ancestor at the point of the SNP’s first appearance, while STRs can be used to estimate the number of generations that have passed since the haplogroup’s most recent common ancestor (MRCA) lived <xref ref-type="bibr" rid="pone.0048638-Jobling1">[1]</xref>, <xref ref-type="bibr" rid="pone.0048638-de1">[9]</xref>.</p>
      <p>For the Y chromosome, coalescence is the process by which two or more haploid male gene genealogies, drawn as a sample from a current population (generation 0,) converge after <italic>t</italic> generations at the MRCA <xref ref-type="bibr" rid="pone.0048638-Hudson1">[10]</xref>. Y-STR comparisons have been used for over a decade as a means of estimating the coalescent age parameter (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e002" xlink:type="simple"/></inline-formula>) in generations for a sample, e.g., <xref ref-type="bibr" rid="pone.0048638-Behar2">[11]</xref>, by measuring the average genetic distance between haplotypes. Two very similar measures of genetic distance for STR allele data, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e003" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e004" xlink:type="simple"/></inline-formula>were published almost simultaneously in 1995, with each measure referred to as “average squared difference” and “average squared distance” respectively <xref ref-type="bibr" rid="pone.0048638-Goldstein1">[12]</xref>, <xref ref-type="bibr" rid="pone.0048638-Slatkin1">[13]</xref>. The <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e005" xlink:type="simple"/></inline-formula>(henceforth ASD) method calculated the ratio of the observed within-locus allele variance, averaged across all sampled loci <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e006" xlink:type="simple"/></inline-formula> and divided by the mean of individual <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e007" xlink:type="simple"/></inline-formula> across all sampled loci <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e008" xlink:type="simple"/></inline-formula>, thus <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e009" xlink:type="simple"/></inline-formula>, with the parameter <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e010" xlink:type="simple"/></inline-formula> estimated in generations to the common ancestor (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e011" xlink:type="simple"/></inline-formula>). Unlike <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e012" xlink:type="simple"/></inline-formula>, ASD was determined to be independent of population size when populations were in mutation-drift equilibrium <xref ref-type="bibr" rid="pone.0048638-Goldstein1">[12]</xref>.</p>
      <p>Under the idealized Strict Stepwise Mutation Model (S-SMM) there is a linear relationship expected between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e013" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e014" xlink:type="simple"/></inline-formula>; any deviations from linearity are due only to stochasticity in the allele mutation process, genetic drift within the sampled population or sampling error <xref ref-type="bibr" rid="pone.0048638-Goldstein2">[14]</xref>. ASD’s linearity was lost only within a population undergoing size change and only until its variance assumed a new equilibrium value, but this loss of linearity was restricted only to the interval when the population was out of mutation-drift equilibrium <xref ref-type="bibr" rid="pone.0048638-Goldstein1">[12]</xref>. It was reported subsequently by the same authors that there was no evidence to support the hypothesis that Y microsatellites were out of mutation-drift equilibrium <xref ref-type="bibr" rid="pone.0048638-Goldstein3">[15]</xref>. ASD also has been shown empirically to be a reliable molecular clock and an unbiased estimator of TMRCA, with a linearity that extended to more than two million years before present in humans and a high correlation coefficient (R<sup>2</sup> = 0.97) with genetic sequence divergence measures based on nucleotide substitutions <xref ref-type="bibr" rid="pone.0048638-Sun1">[16]</xref>. Another attractive property of ASD, as a model-free means of calculating an unbiased point estimate for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e015" xlink:type="simple"/></inline-formula>, is its independence from the shape of a population’s genealogy, size or growth rate <xref ref-type="bibr" rid="pone.0048638-Stumpf1">[17]</xref>.</p>
      <p>Some problems remain with the accuracy of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e016" xlink:type="simple"/></inline-formula>, however. In particular, the selection of appropriate mutation rates for use when calculating <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e017" xlink:type="simple"/></inline-formula> is an ongoing topic of controversy <xref ref-type="bibr" rid="pone.0048638-Zhivotovsky1">[18]</xref>, <xref ref-type="bibr" rid="pone.0048638-Balaresque1">[19]</xref>, <xref ref-type="bibr" rid="pone.0048638-Busby1">[20]</xref>. An accurate estimate of the coalescent is strongly dependent upon the quality of the estimates of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e018" xlink:type="simple"/></inline-formula> <xref ref-type="bibr" rid="pone.0048638-Busby1">[20]</xref>. There has been no systematic comparison made yet, however, of different sets of published Y STR mutation rates with respect to their ability to produce an accurate estimate of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e019" xlink:type="simple"/></inline-formula>. Herein, I examined the effects of these differences in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e020" xlink:type="simple"/></inline-formula> on the calculation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e021" xlink:type="simple"/></inline-formula> by comparing estimates generated from these three published sets of Y-STR mutation rates while using the average square difference (ASD) approach to calculate the coalescent.</p>
      <p>Almost 200 STR loci and their individually estimated mutation rates (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e022" xlink:type="simple"/></inline-formula>) have been identified in the non-recombining (NRY) region of the human Y chromosome <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>, <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>. A recent study has estimated 186 Y-STR mutation rates using a Bayesian posterior distribution analysis, detecting 82 loci (44%) in the range of 1.0×10<sup>−4</sup>, 91 loci (48.9%) with mutation rates in the range of 1.0×10<sup>−3</sup> and 13 (6.9%) in the 1.0×10<sup>−2</sup> range, covering two orders or magnitude <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>. Another made estimates for 110 STR loci based on a logistic regression model and ranged from 3.6×10<sup>−4</sup> per generation to 9.6×10<sup>−3</sup> <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>. A third set of estimates maintained by the Y Chromosome Haplotype Reference Database (<ext-link ext-link-type="uri" xlink:href="http://www.YHRD.org" xlink:type="simple">www.YHRD.org</ext-link>), also is used widely. Two types of Y-STRs are used typically in genetic research and applications: (1) single-site STRs (ssSTRs) containing only one uninterrupted variable stretch of repeats and (2) multi-site STRs (msSTRs) where more than one site on the chromosome binds with the primer used to determine the number of allele repeats present at the STR of interest. Single-site STRs have a more linear relationship between mutation rates and accumulated variance than msSTRs <xref ref-type="bibr" rid="pone.0048638-Sun1">[16]</xref>, a finding with direct relevance to the calculation of the coalescent.</p>
      <p>The estimation and use of a correct mutation rate is problematic with regard to msSTRs, such as DYS 385a/b or DYS464a/b/c/d, because it is not possible under standard genotyping protocols (without direct sequencing) to distinguish which of the duplicated sites has mutated <xref ref-type="bibr" rid="pone.0048638-Goedbloed1">[21]</xref>, <xref ref-type="bibr" rid="pone.0048638-Vermeulen1">[22]</xref>. If independent mutational mechanisms occur with each part of a multi-part STR <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>, then without knowledge of the individual per-copy mutation rates or the ability (generally) to distinguish one copy of the msSTR from another, the ascertainment problem will be compounded by the potential for application of an incorrect mutation rate estimate to the wrong locus.</p>
      <p>A different type of problem arises with the multi-part locus DYS389. Rather than being a duplicated STR locus, DYS389 is a complex consisting of one polymorphic region, DYS389I, contained entirely within another, larger polymorphic region, DYS389II, both of which are amplified by the primer <xref ref-type="bibr" rid="pone.0048638-Goedbloed1">[21]</xref>. Because DYS389I is contained entirely within DYS389II, the allele value of DYS389II is partly dependent on the value of DYS389I. Rather than having a multipart locus with two identical or nearly identical parts that cannot be distinguished easily, the DYS389 locus has one independent and one dependent part. A mutation in DYS389I will change the allele value of DYS389II, but the reverse is not true; a mutation occurring outside of the DYS389I portion will not change the allele value of DYS389I. Thus, this locus will suffer from collinearity – two apparently independent variables that are in fact different expressions of the same variable – at least part of the time if both DYS389I and DYS389II are used in the calculation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e023" xlink:type="simple"/></inline-formula>.</p>
      <p>Another potential source of bias in the estimation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e024" xlink:type="simple"/></inline-formula>may be introduced by the use of one or more ssSTR loci that do not adhere closely to the S-SMM. For example, DYS392 and DYS438 were found to be bimodal in their distributions of allele repeat values, with certain pairs of allele values for both loci also having high linkage disequilibrium measures, D’ = 0.70 for {DYS392 = 11,DYS438 = 10} and D’ = 0.72 for {DYS392 = 13,DYS438 = 12}; both pairings were associated with each other strongly in real populations <xref ref-type="bibr" rid="pone.0048638-Gusmo1">[23]</xref>. The authors attributed these unusual allele distributions to ancient demographic events rather than departures from the stepwise mutation model, but the consequences for the accurate calculation of the coalescent were not explored in detail.</p>
      <p>The problem of bimodal allele value distributions also was described by <xref ref-type="bibr" rid="pone.0048638-de1">[9]</xref>, with the finding that the bimodal distribution of DYS392 differed substantially between two haplogroups as defined by “Unique Mutation Events” (UMEs, identical to UEPs or SNPs). The author of <xref ref-type="bibr" rid="pone.0048638-de1">[9]</xref> attributed this phenomenon to a “bottleneck” effect, with UMEs, defining newly emerged Y haplogroups as “bottlenecks.” (Note that the use of the term bottleneck in <xref ref-type="bibr" rid="pone.0048638-de1">[9]</xref> differs from the more conventional definition used within population biology, i.e., a sharp decline in breeding individuals within a population, leading to greatly reduced genetic diversity among the descendants of the surviving individuals.) <xref ref-type="bibr" rid="pone.0048638-de1">[9]</xref> noted that the net result of this effect was to cause an apparent breakdown of linkage disequilibrium, resulting in differences in genetic variation between Y haplogroups for DYS392. As part of this study, statistical analyses of STR data sets with and without these loci present (shown below) have identified significant biasing effects from DYS392 and DYS438, when used together, on the estimation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e025" xlink:type="simple"/></inline-formula>.</p>
      <p>The Strict Stepwise Mutation Model (S-SMM) first described the process of genetic mutation for STRs <xref ref-type="bibr" rid="pone.0048638-Ohta1">[5]</xref>, <xref ref-type="bibr" rid="pone.0048638-Kimura1">[6]</xref>. Alleles increased or decreased by one set of repeats when copied during meiosis according to an expected rate of mutation, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e026" xlink:type="simple"/></inline-formula>, equivalent to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e027" xlink:type="simple"/></inline-formula> (the per-locus mutation rate) in other studies <xref ref-type="bibr" rid="pone.0048638-Goldstein1">[12]</xref>, <xref ref-type="bibr" rid="pone.0048638-Goldstein3">[15]</xref>. The expected change in one generation was expressed as:<disp-formula id="pone.0048638.e028"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0048638.e028" xlink:type="simple"/></disp-formula>where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e029" xlink:type="simple"/></inline-formula> was for a starting allele repeat count in a population, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e030" xlink:type="simple"/></inline-formula> was for the per-locus mutation rate, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e031" xlink:type="simple"/></inline-formula> and <italic>x</italic><sub>+1</sub> denoted the repeat counts adjacent to <italic>x</italic>, either <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e032" xlink:type="simple"/></inline-formula> or <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e033" xlink:type="simple"/></inline-formula> repeats, respectively, and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e034" xlink:type="simple"/></inline-formula> was for change due to random sampling of gametes <xref ref-type="bibr" rid="pone.0048638-Kimura1">[6]</xref>.</p>
      <p>A modification of the S-SMM, called the Generalized Stepwise Mutation Model (G-SMM) was developed by <xref ref-type="bibr" rid="pone.0048638-Gusmo1">[23]</xref>. In the process of extending the G-SMM, several important characteristics of the S-SMM’s distribution were described by <xref ref-type="bibr" rid="pone.0048638-Calabrese1">[25]</xref>. If <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e035" xlink:type="simple"/></inline-formula> was large, then the Poisson probability distribution of the idealized S-SMM had a kurtosis (<italic>Κ</italic>) of 3.0. If <italic>K</italic> was large, however, then the distribution had a heavy tail and any estimate of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e036" xlink:type="simple"/></inline-formula>using STRs was difficult. For microsatellites, kurtosis became large when the mutation rate (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e037" xlink:type="simple"/></inline-formula>) multiplied by the number of generations <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e038" xlink:type="simple"/></inline-formula> to the coalescent, then divided by twice the allele repeat length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e039" xlink:type="simple"/></inline-formula> minus the allele minimum length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e040" xlink:type="simple"/></inline-formula> squared, was large relative to one <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e041" xlink:type="simple"/></inline-formula>. A large kurtosis also signalled that there was significant probability that the locus has undergone “microsatellite death,” or loss of its mutational activity <xref ref-type="bibr" rid="pone.0048638-Calabrese1">[25]</xref>. For diploid microsatellite data comparing African vs. non-African human populations first analyzed in <xref ref-type="bibr" rid="pone.0048638-Goldstein3">[15]</xref>, a kurtosis of 3.02 was calculated by <xref ref-type="bibr" rid="pone.0048638-Calabrese1">[25]</xref>, nearly identical to the expected kurtosis under the S-SMM. In comparison, the human-chimpanzee split exhibited a kurtosis of 3.93. It was reasonable to assume, therefore, that a kurtosis very near 3.0 was the correct expectation for an individual STR allele distribution adhering very closely to the S-SMM.</p>
      <p>Rather than attempting to calibrate the G-SMM to account for potential bias in some loci, a more practical approach to unbiased calculation of the coalescent may be to develop a method for identifying <italic>a priori</italic> those loci that conform most closely to expectations for the S-SMM. Using this approach, any potential bias is minimized by choosing, from available ssSTRs, those that exhibit the least deviation from the S-SMM in their observed allele distributions. We recall that the kurtosis of the idealized S-SMM is expected to be 3.0, identical to that of the normal distribution. A well-known property of the normal probability distribution is that the mean, median and mode values are the same <xref ref-type="bibr" rid="pone.0048638-Patel1">[26]</xref>. The arithmetic mean-geometric mean inequality (AM-GM) states that the geometric mean of a set of numbers is always less than the arithmetic mean, except when all of the numbers in the set are identical <xref ref-type="bibr" rid="pone.0048638-Uchida1">[27]</xref>. Jensen’s Inequality has been used both to prove the mean-median-mode equivalency for the normal distribution and to prove the AM-GM inequality <xref ref-type="bibr" rid="pone.0048638-Mallows1">[28]</xref>, <xref ref-type="bibr" rid="pone.0048638-Cartwright1">[29]</xref>. These properties suggest a hypothesis, that the arithmetic and geometric means will be identical for an infinitely large, quasi-normal Poisson distribution such as the S-SMM, but that the geometric mean will be lower for non-normally distributed data and the difference will be proportional to the degree of departure from the S-SMM by the sample data set. The hypothesis is empirically testable by using computer generated S-SMM and normal distributions to compare the arithmetic and geometric means for both.</p>
      <p>The duration of linearity for a microsatellite locus grows at the square of the allele’s range of values <xref ref-type="bibr" rid="pone.0048638-Goldstein3">[15]</xref> and the maximum range of any particular STR locus is considered to be the most important constraint on the linearity of the stepwise mutation model by <xref ref-type="bibr" rid="pone.0048638-Busby1">[20]</xref>. High kurtosis (&gt;&gt;3.0) signals an increased probability of microsatellite death and excessive skewness identifies those loci departing substantially from the expected allele distribution under the S-SMM. Incorporating all three of these factors into one estimator allows for a measure that can be used to identify the best available loci for improved accuracy in the calculation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e042" xlink:type="simple"/></inline-formula> from among those available.</p>
      <p>The present study considered the kurtosis and skewness of the probability distribution for alleles at each locus in relationship with the range of the allele values. The formula developed for this purpose was:<disp-formula id="pone.0048638.e043"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0048638.e043" xlink:type="simple"/><label>(1)</label></disp-formula>where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e044" xlink:type="simple"/></inline-formula> stood for kurtosis, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e045" xlink:type="simple"/></inline-formula> for skewness and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e046" xlink:type="simple"/></inline-formula> for range (highest minus lowest allele value), with the product <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e047" xlink:type="simple"/></inline-formula> (“quality”) equaling zero for a set of allele values within an idealized locus that conformed perfectly to the S-SMM. The new expression permitted the effects of kurtosis, skewness and range of any STR locus, for which these parameters can be estimated with reasonable accuracy, to be assessed using only the observable distribution of its allele values. Because any locus that conformed very closely to the S-SMM’s expectations for kurtosis and skewness had a numerator value near zero under <xref ref-type="disp-formula" rid="pone.0048638.e043">equation (1</xref>), while a larger range value <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e048" xlink:type="simple"/></inline-formula> caused the denominator to grow exponentially, the limit of ratios for which the numerator remained small compared to the denominator (and the allele distribution’s kurtosis remains near 3.0) was extended either by better adherence to the S-SMM or a larger range of possible allele values.</p>
      <p>The distribution of variances for Y-STR loci is not normal but rather log-normal <xref ref-type="bibr" rid="pone.0048638-Goldstein2">[14]</xref>. The geometric mean therefore may be a better choice for calculating the average between-locus variance than the arithmetic mean. The geometric mean is analogous to the median, as it measures the median tendency of random fluctuations in probabilities across populations for factors such as random selection of gametes, birth and death rates, mutation and sampling error <xref ref-type="bibr" rid="pone.0048638-Mills1">[30]</xref>, with the median itself minimizing the sum of the absolute deviations about any point <xref ref-type="bibr" rid="pone.0048638-Schwertman1">[31]</xref>. As there is substantial randomness (stochasticity) contributing to the variance of observed ASD (oASD) <xref ref-type="bibr" rid="pone.0048638-Goldstein2">[14]</xref>, the geometric mean may approximate the true value of the average between-locus variance better than the arithmetic mean. An approach to mitigating the problem of inaccurate estimates for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e049" xlink:type="simple"/></inline-formula> therefore is suggested. The use of the geometric mean to calculate oASD across loci and also the average mutation rate across loci <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e050" xlink:type="simple"/></inline-formula> may result in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e051" xlink:type="simple"/></inline-formula> being less sensitive to errors in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e052" xlink:type="simple"/></inline-formula>, to any deviations from the S-SMM by individual loci or any variation in oASD due to sampling error or stochasticity.</p>
      <p>In general, the current study sought to improve the accuracy of the coalescent age estimate by reducing variation caused by cryptic errors in mutation rate estimates and the use of loci that did not conform adequately to the S-SMM. The subsequent analysis identified specific procedural adjustments that were successful in reducing the amount of unexplained variation by a substantial amount, thus improving the accuracy of the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e053" xlink:type="simple"/></inline-formula> estimator.</p>
    </sec>
    <sec id="s2" sec-type="materials|methods">
      <title>Materials and Methods</title>
      <p>A comparison was made between the arithmetic and geometric means for a normal distribution by using a data set with 1.0E+8 data points, randomly generated by the “rnorm” function in the R statistical software package (<ext-link ext-link-type="uri" xlink:href="http://www.r-project.org" xlink:type="simple">http://www.r-project.org</ext-link>). Within the simulated normal distribution, true means of 8, 10, 20, 50, 100, 500 and 1000 were compared to calculated arithmetic and geometric means from the resulting normally distributed, randomly generated data sets. Next, the expectation that the arithmetic and geometric means were identical for an ideal S-SMM distribution was tested empirically by comparing the means from a large simulated Y-STR data set generated using the program SIMCOAL 2.0 (<ext-link ext-link-type="uri" xlink:href="http://cmpg.unibe.ch/software/simcoal2/" xlink:type="simple">http://cmpg.unibe.ch/software/simcoal2/</ext-link>). The simulated data were intended to reflect the mutational behavior of a fully linked set of 32 loci on the Y chromosome, under the idealized S-SMM, equivalent to approximately 500,000 years of descent (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e054" xlink:type="simple"/></inline-formula> = 18,738) from an initial breeding population of 100 males.</p>
      <p>Two sets of real population sample data, containing Y-STRs used frequently in population and forensic studies, were employed to identify loci that conformed most closely to the S-SMM. The first was a set of compiled STR allele repeat values extracted from the public online database SMGF (<ext-link ext-link-type="uri" xlink:href="http://www.smgf.org" xlink:type="simple">www.smgf.org</ext-link>), containing the number of times a given allele repeat value for a particular locus appeared in their Y chromosome database of approximately 35,600 haplotypes. Of the 48 STR loci available from the SMGF database, only the ssSTRs (32 loci) were used for analysis, following the recommendations of <xref ref-type="bibr" rid="pone.0048638-Vermeulen1">[22]</xref>. Fractional allele repeats (i.e., 10.1, 12.2), null alleles (0) and alleles with more than one repeat value detected (e.g., DYS19 with 12–15 or 13–14) were excluded from analysis, but constituted a negligible fraction of the STR values examined. Allele counts for all loci used are available online from SMGF. These data were evaluated by formula (1) to measure the degree of departure from the S-SMM, using a script written in R, available upon request. The results were sorted from lowest (“best”) to highest (“worst”) value according to the calculated <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e055" xlink:type="simple"/></inline-formula> statistic for each ssSTR locus.</p>
      <p>The second set of data was obtained from the “British Isles DNA Project,” a public database of haplotypes contributed by individuals who have been STR-tested through various commercial testing services containing 3,955 Y haplotypes with resolutions ranging from 11 to 102 STR loci per haplotype (<ext-link ext-link-type="uri" xlink:href="http://www.familytreedna.com/public/BritishIsles" xlink:type="simple">www.familytreedna.com/public/BritishIsles</ext-link>). From this group, all available haplotypes containing the identical set of ssSTRs reported in the SMGF database (henceforth, STR Data Set or SDS) were extracted (n = 245). To assess the effect of the number of individual STR loci compared within a data set on the value of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e056" xlink:type="simple"/></inline-formula>, an overall estimate for the STR Data Set was compared with a series of progressively smaller subsets of the SDS. First, the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e057" xlink:type="simple"/></inline-formula> values of the 32 individual loci were calculated using <xref ref-type="disp-formula" rid="pone.0048638.e043">equation (1</xref>), then sorted in ascending numerical order (“best to worst”). The value of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e058" xlink:type="simple"/></inline-formula> and 95% confidence intervals (CIs) for the entire set was calculated using the “Ytimeboot” program within the Matlab-based software package <italic>Ytime</italic> <xref ref-type="bibr" rid="pone.0048638-Behar2">[11]</xref>, with per-locus mutation rates taken from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>. A strict stepwise mutation model (S-SMM) was selected and a starlike (‘star’) genealogy was assumed. The ancestral haplotype for the most recent common ancestor was estimated from the median allele repeat value for each locus <xref ref-type="bibr" rid="pone.0048638-Sengupta1">[32]</xref>. The last locus (“worst”) in the STR data set then was removed (e.g., from 32 to 31 loci) and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e059" xlink:type="simple"/></inline-formula> was recalculated. The process was repeated (31 to 30 loci, etc.) until only one locus remained. The resulting <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e060" xlink:type="simple"/></inline-formula>statistics and 95% CIs from each trial were plotted for comparison.</p>
      <p>Next, the identical 32 loci from the SDS were processed in the same way, but with the per-locus <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e061" xlink:type="simple"/></inline-formula> taken from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>. Because the 32 locus data set entering the procedure each time was the same, any differences in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e062" xlink:type="simple"/></inline-formula>were attributable entirely to per locus differences in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e063" xlink:type="simple"/></inline-formula> between the two studies. The 15 STR loci ranked by <xref ref-type="bibr" rid="pone.0048638-Busby1">[20]</xref> then were subjected to the same procedure, according to the preferred order provided in <xref ref-type="table" rid="pone-0048638-t001">Table 1</xref> of that study, with the rates taken from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref> as in <xref ref-type="bibr" rid="pone.0048638-Busby1">[20]</xref>. Finally, the same 15 locus subset, containing the 15 ssSTR loci maintained by YHRD, was extracted and processed in the same way but using the YHRD-maintained mutation rate averages for the loci compared in place of those from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>.</p>
      <table-wrap id="pone-0048638-t001" position="float">
        <object-id pub-id-type="doi">10.1371/journal.pone.0048638.t001</object-id>
        <label>Table 1</label>
        <caption>
          <title>Results of normal distribution simulation.</title>
        </caption>
        <alternatives>
          <graphic id="pone-0048638-t001-1" position="float" mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0048638.t001" xlink:type="simple"/>
          <table>
            <colgroup span="1">
              <col align="left" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
            </colgroup>
            <thead>
              <tr>
                <td align="left" rowspan="1" colspan="1">True mean</td>
                <td align="left" rowspan="1" colspan="1">Arithmetic mean</td>
                <td align="left" rowspan="1" colspan="1">Geometric mean</td>
                <td align="left" rowspan="1" colspan="1">% difference</td>
              </tr>
            </thead>
            <tbody>
              <tr>
                <td align="left" rowspan="1" colspan="1">8</td>
                <td align="left" rowspan="1" colspan="1">7.999827</td>
                <td align="left" rowspan="1" colspan="1">7.936025</td>
                <td align="left" rowspan="1" colspan="1">0.099695%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">10</td>
                <td align="left" rowspan="1" colspan="1">10.000038</td>
                <td align="left" rowspan="1" colspan="1">9.949390</td>
                <td align="left" rowspan="1" colspan="1">0.050648%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">20</td>
                <td align="left" rowspan="1" colspan="1">20.000045</td>
                <td align="left" rowspan="1" colspan="1">19.974964</td>
                <td align="left" rowspan="1" colspan="1">0.006270%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">50</td>
                <td align="left" rowspan="1" colspan="1">50.000153</td>
                <td align="left" rowspan="1" colspan="1">49.990146</td>
                <td align="left" rowspan="1" colspan="1">0.000400%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">100</td>
                <td align="left" rowspan="1" colspan="1">100.000040</td>
                <td align="left" rowspan="1" colspan="1">99.995039</td>
                <td align="left" rowspan="1" colspan="1">0.000050%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">500</td>
                <td align="left" rowspan="1" colspan="1">499.999940</td>
                <td align="left" rowspan="1" colspan="1">499.998940</td>
                <td align="left" rowspan="1" colspan="1">&lt;0.000001%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">1000</td>
                <td align="left" rowspan="1" colspan="1">1000.000010</td>
                <td align="left" rowspan="1" colspan="1">999.999510</td>
                <td align="left" rowspan="1" colspan="1">&lt;0.000001%</td>
              </tr>
            </tbody>
          </table>
        </alternatives>
        <table-wrap-foot>
          <fn id="nt101">
            <p>Difference in arithmetic and geometric means for normal distributions (n = 100,000,000). Each new calculation was made with a <italic>de novo</italic> normal distribution generated.</p>
          </fn>
        </table-wrap-foot>
      </table-wrap>
      <p>The effects of differing mutation rate estimates <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e064" xlink:type="simple"/></inline-formula> for individual loci in the STR data set on <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e065" xlink:type="simple"/></inline-formula>were evaluated, generating per locus estimates for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e066" xlink:type="simple"/></inline-formula> ASD and the 95% CIs for each of the three mutation rate sets compared. A linear regression (LR) model was employed to predict ASD from <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e067" xlink:type="simple"/></inline-formula> for each STR. Observed ASD was compared to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e068" xlink:type="simple"/></inline-formula> to test for a significant relationship, with individual locus values of<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e069" xlink:type="simple"/></inline-formula> as the independent variable and ASD as the dependent variable. To test for any significant interaction between ASD and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e070" xlink:type="simple"/></inline-formula>, a multiple regression analysis also was conducted for each mutation rate set, using the model:<disp-formula id="pone.0048638.e071"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0048638.e071" xlink:type="simple"/><label>(2)</label></disp-formula></p>
      <p>R was used for the evaluation of all linear regression models.</p>
      <p>A hypothetical cause of differences in ssSTR mutation rates may be the minimum allele length<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e072" xlink:type="simple"/></inline-formula> for a given locus, with shorter minimum alleles possibly corresponding to lower mutation rates. To test for this possibility, LR analyses were performed for mutation rate estimates <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e073" xlink:type="simple"/></inline-formula> from both <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref> and <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e074" xlink:type="simple"/></inline-formula> as the independent variable and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e075" xlink:type="simple"/></inline-formula> as the dependent variable. A similar LR was performed to test for a possible correlation between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e076" xlink:type="simple"/></inline-formula> and the STR’s calculated <italic>q</italic> value, with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e077" xlink:type="simple"/></inline-formula> as the independent variable and <italic>q</italic> as the dependent variable. Both comparisons used the allele data from SMGF for the 32 locus set.</p>
      <p>The difference between the arithmetic and geometric means for a sample’s probability distribution was exploited to measure the percentage degree of departure by a sample set of haplotypes from the S-SMM. To assess the degree of departure for a given population sample from the S-STR, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e078" xlink:type="simple"/></inline-formula> was calculated for the set of simulated haplotypes using both the geometric and arithmetic means and the difference (<italic>D</italic>) between them was measured using the formula:<disp-formula id="pone.0048638.e079"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0048638.e079" xlink:type="simple"/><label>(3)</label></disp-formula>where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e080" xlink:type="simple"/></inline-formula> stood for the geometric average of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e081" xlink:type="simple"/></inline-formula>and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e082" xlink:type="simple"/></inline-formula> stood for the arithmetic average. <italic>D</italic> statistics were calculated for relevant combinations of actual STR data and mutation rates compared in the study.</p>
      <p>Although the central <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e083" xlink:type="simple"/></inline-formula>point estimate is independent of the shape of the genealogy of a population, the corresponding confidence intervals are strongly dependent on the (often unknown) growth rate of the sampled population. Except for the case of an extreme population bottleneck, though, the actual CIs will be bounded by the ones calculated for zero growth rate (a constant effective population size) and an exponentially growing population in which all lines survive to the present and are represented in the sample (a “starlike” genealogy) <xref ref-type="bibr" rid="pone.0048638-Stumpf1">[17]</xref>. Using <italic>Ytime</italic>, the effects of different population genealogies on both the central <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e084" xlink:type="simple"/></inline-formula>statistic and the resulting CIs were investigated. The <italic>Ytime</italic> parameter <italic>Rgrowth</italic> defined a scaled growth rate for the population sample being analyzed, with the parameter specified by <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e085" xlink:type="simple"/></inline-formula> where<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e086" xlink:type="simple"/></inline-formula>was the current effective population size and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e087" xlink:type="simple"/></inline-formula> was the instantaneous growth rate per generation. A no-growth, constant-sized population was specified in <italic>Ytime</italic> by using <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e088" xlink:type="simple"/></inline-formula>; any intermediate parameter value from zero to ‘star’ (quasi-infinite exponential) growth was available. The <italic>Rgrowth</italic> parameter was varied between zero and ‘star’ by eleven orders of magnitude (from 0.0E+0 to 1.0E+11) and the effects on <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e089" xlink:type="simple"/></inline-formula> and the 95% CIs were compared for each growth rate, using the SDS with 32 loci.</p>
    </sec>
    <sec id="s3">
      <title>Results</title>
      <p>Comparing the means of several sets of normal distributions (1.0E+8) data points randomly generated for each set, it was determined that the arithmetic and geometric means deviated by a maximum of 0.1% with a true mean value of 8 and with the difference diminishing to less than a millionth of a percent with a true mean value of 500 and 1000 (<xref ref-type="table" rid="pone-0048638-t001">Table 1</xref>). For the idealized S-SMM simulated data set, the ratio of the difference between the arithmetic mean and the geometric mean for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e090" xlink:type="simple"/></inline-formula> was calculated to be 0.06%, with a kurtosis of 3.027 and skewness of −0.002, conforming closely to expectations for the S-SMM as described by <xref ref-type="bibr" rid="pone.0048638-Kimmel1">[24]</xref>. These results supported the inference that an infinitely large data set conforming perfectly to the S-SMM would have identical geometric and arithmetic means, and that any difference in the means reflected the degree of divergence from the ideal S-SMM for the sample data set.</p>
      <p>Using <xref ref-type="disp-formula" rid="pone.0048638.e043">equation (1</xref>), the lowest calculated <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e091" xlink:type="simple"/></inline-formula> values were used to identify loci that departed the least from the S-SMM and therefore were preferred for the purpose of calculating <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e092" xlink:type="simple"/></inline-formula> (<xref ref-type="table" rid="pone-0048638-t002">Table 2</xref>). In graphing the relative values of the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e093" xlink:type="simple"/></inline-formula> statistic for the test loci, two noticeable changes in the slope of the curve occur; the first after DYS 462 and the second after DYS635. The resulting values for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e094" xlink:type="simple"/></inline-formula> thus fell into easily identifiable ranges of “best”, “average” and “worst” conformity to the S-SMM, for these 32 loci (<xref ref-type="fig" rid="pone-0048638-g001">Figure 1</xref>).</p>
      <fig id="pone-0048638-g001" position="float">
        <object-id pub-id-type="doi">10.1371/journal.pone.0048638.g001</object-id>
        <label>Figure 1</label>
        <caption>
          <title>Calculated <italic>q</italic> values for 32 loci, SMGF data set.</title>
          <p>A lower <italic>q</italic> value equals a better fit to the S-SMM for an individual locus.</p>
        </caption>
        <graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0048638.g001" position="float" xlink:type="simple"/>
      </fig>
      <table-wrap id="pone-0048638-t002" position="float">
        <object-id pub-id-type="doi">10.1371/journal.pone.0048638.t002</object-id>
        <label>Table 2</label>
        <caption>
          <title>32 commonly used, single-site STRs and their q values.</title>
        </caption>
        <alternatives>
          <graphic id="pone-0048638-t002-2" position="float" mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0048638.t002" xlink:type="simple"/>
          <table>
            <colgroup span="1">
              <col align="left" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
            </colgroup>
            <thead>
              <tr>
                <td align="left" rowspan="1" colspan="1">STR locus</td>
                <td align="left" rowspan="1" colspan="1"><italic>q</italic> value</td>
                <td align="left" rowspan="1" colspan="1">STR locus</td>
                <td align="left" rowspan="1" colspan="1"><italic>q</italic> value</td>
              </tr>
            </thead>
            <tbody>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS390</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.000412</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS426</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.041725</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS458</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.000603</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS389I</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.044951</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS439</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.000744</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS442</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.048485</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS438</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.003111</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS447</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.048934</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS456</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.005350</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS448</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.048934</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS392</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.008238</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS635</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.049764</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS391</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.009169</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS452</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.053431</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS449</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.010557</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS19</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.057823</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>GATA-H4</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.010700</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS461</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.062045</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>GATA-A10</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.014808</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS463</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.067503</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS460</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.015247</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS446</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.086166</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS437</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.017258</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS393</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.109470</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS462</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.019705</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS445</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.110377</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS444</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.032796</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS455</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.115019</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS389II</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.036689</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS454</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.154864</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS441</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.040277</td>
                <td align="left" rowspan="1" colspan="1">
                  <italic>DYS388</italic>
                </td>
                <td align="left" rowspan="1" colspan="1">0.155135</td>
              </tr>
            </tbody>
          </table>
        </alternatives>
        <table-wrap-foot>
          <fn id="nt102">
            <p>Loci are ranked in order of conformity to the strict stepwise mutation model (S-SMM) using <xref ref-type="disp-formula" rid="pone.0048638.e043">equation (1</xref>). The lowest <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e095" xlink:type="simple"/></inline-formula> value identifies the best fit to the S-SMM, using the SMGF summary statistics for 32 loci (n≈35,600).</p>
          </fn>
        </table-wrap-foot>
      </table-wrap>
      <p>Estimates of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e096" xlink:type="simple"/></inline-formula> based on diminishing numbers of loci extracted from the same data set, in which those loci with greater deviations from the S-SMM were removed first, varied from 220 generations to 380, a difference of 72% between the lowest and highest value when using the mutation rates from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> (<xref ref-type="fig" rid="pone-0048638-g002">Figure 2A</xref>). When the mutation rates from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref> were substituted, the range from the lowest to highest <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e097" xlink:type="simple"/></inline-formula> varied from 281 generations to 701 generations, a difference of 249% (<xref ref-type="fig" rid="pone-0048638-g002">Figure 2B</xref>). Similar results were obtained for the 15 locus data sets (<xref ref-type="fig" rid="pone-0048638-g002">Figure 2C, 2D</xref>.) The stability and linearity of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e098" xlink:type="simple"/></inline-formula> deteriorated with decreasing numbers of loci in all cases.</p>
      <fig id="pone-0048638-g002" position="float">
        <object-id pub-id-type="doi">10.1371/journal.pone.0048638.g002</object-id>
        <label>Figure 2</label>
        <caption>
          <title>Coalescent calculations (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e099" xlink:type="simple"/></inline-formula>) and confidence intervals.</title>
          <p>A) Burgarella mutation rates, 32 loci, best to worst order. B) Ballantyne mutation rates, 32 loci, best to worst order. C) Busby et al., 2012, 15 recommended loci, best to worst order. D) YHRD mutation rates, best to worst order, 15 loci. TMRCA is given in generations. Error bars represent 95% confidence intervals for each TMRCA estimate. Dashed lines indicate the median <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e100" xlink:type="simple"/></inline-formula> value for each comparison, shown in the box at the right end of the dashed line.</p>
        </caption>
        <graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0048638.g002" position="float" xlink:type="simple"/>
      </fig>
      <p>Mutation rates from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> showed a significant relationship between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e101" xlink:type="simple"/></inline-formula> and ASD (t = 4.645, df = 30, p&lt;0.0001, R<sup>2</sup> = 0.4184) but with a large portion of the variation (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e102" xlink:type="simple"/></inline-formula>) remaining unexplained (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3A</xref>). For the estimates from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>, ASD was not significantly related to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e103" xlink:type="simple"/></inline-formula> (t = 1.555, df = 30, p = 0.1304, R<sup>2</sup> = 0.0746), an unexpected result (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3B</xref>). The YHRD <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e104" xlink:type="simple"/></inline-formula> rates (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3C</xref>) also were not significantly related to ASD (t = 1.389, df = 30, p = 0.1880, R<sup>2</sup> = 0.1293). When data for the seven loci with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e105" xlink:type="simple"/></inline-formula>&lt;0.001 were removed from the LR data, however, the amount of explained variation in ASD due to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e106" xlink:type="simple"/></inline-formula> (rates from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>) increased sharply (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3D</xref>), with R<sup>2</sup> = 0.7535 (t = 8.20, df = 22, p&lt;0.0001). Subjecting rates from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref> to the removal of all loci with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e107" xlink:type="simple"/></inline-formula>&lt;0.001 did not improve the non-significant relationship between ASD and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e108" xlink:type="simple"/></inline-formula> (t = 1.147, df = 22, p = 0.2636, R<sup>2</sup> = 0.056). Removing two loci from the 15 YHRD loci with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e109" xlink:type="simple"/></inline-formula>&lt;0.001 (DYS438 and DYS392) and also DYS389I (collinear with DYS389II) improved the correlation coefficient between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e110" xlink:type="simple"/></inline-formula> and ASD for the remaining 12 loci to a statistically significant level (R<sup>2</sup> = 0.4921, t = 3.199, df = 10, p = 0.01101).</p>
      <fig id="pone-0048638-g003" position="float">
        <object-id pub-id-type="doi">10.1371/journal.pone.0048638.g003</object-id>
        <label>Figure 3</label>
        <caption>
          <title>Regression coefficients and confidence intervals.</title>
          <p>ASD for each locus compared with mutation rates for each locus, using three published studies. A) Burgarella rates, 32 loci. Loci identified in bold face (e.g., DYS448) fall within the 95% mean confidence interval for the locus set compared. B) Ballantyne rates, 32 loci. C) YHRD, 15 loci, YHRD rates. D) Burgarella rates, 24 loci, with loci that have <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e111" xlink:type="simple"/></inline-formula>0.001 removed from the calculation. Data set compared is from the “British Isles DNA Project,” n = 245, 32 locus haplotypes and subsets of the same data set (24 locus reduced set and 15 locus set corresponding to the YHRD ssSTR loci.).</p>
        </caption>
        <graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0048638.g003" position="float" xlink:type="simple"/>
      </fig>
      <p>With the mutation rate estimates from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>, there was a significant relationship detected between ASD and the response variable <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e112" xlink:type="simple"/></inline-formula> (t = −6.480, df = 28, p&lt;0.0001) and between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e113" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e114" xlink:type="simple"/></inline-formula> (t = 4.109, df = 28, p = 0.0003), but not for the interaction <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e115" xlink:type="simple"/></inline-formula> (t = 1.748, df = 28, p = 0.0914). With estimates from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>, there was a significant relationship between ASD and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e116" xlink:type="simple"/></inline-formula> (t = 3.682, df = 28, p = 0.0010) and between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e117" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e118" xlink:type="simple"/></inline-formula> (t = −2.423, df = 28, p = 0.0221) but the interaction <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e119" xlink:type="simple"/></inline-formula> again was not significant (t = −0.937, df = 28, p = 0.3570). The 15 locus YHRD mutation rate set showed a significant relationship between ASD and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e120" xlink:type="simple"/></inline-formula> (t = 3.598, df = 11, p = 0.0042), but this time the relationship between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e121" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e122" xlink:type="simple"/></inline-formula> was not significant (t = −0.302, df = 11, p = 0.7683) with 15 loci. However, when DYS438, DYS392 and DYS389I were removed from the allele set, the relationship between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e123" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e124" xlink:type="simple"/></inline-formula> became significant (t = −3.192, df = 8, p = 0.01276). The interaction <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e125" xlink:type="simple"/></inline-formula> remained non-significant (t = −0.583, df = 8, p = 0.57616) for the reduced data set.</p>
      <p>Linear regression analyses comparing the minimum allele length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e126" xlink:type="simple"/></inline-formula> for a given locus with its estimated mutation rate <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e127" xlink:type="simple"/></inline-formula> found no significant correlation between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e128" xlink:type="simple"/></inline-formula> and<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e129" xlink:type="simple"/></inline-formula> using rate estimate from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref> (t = 1.275, df = 30, p = 0.212) and <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> (t = 0.061, df = 30, p = 0.952). A comparison of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e130" xlink:type="simple"/></inline-formula> for a given locus with its calculated <italic>q</italic> value (which is not dependent on <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e131" xlink:type="simple"/></inline-formula>) also found that the correlation was not significant (t = −1.027, df = 30, p = 0.313).</p>
      <p>The calculation of <italic>D</italic> for several sets of loci and mutation rate combinations determined that the closest agreement between the two averages was calculated by using mutation rates from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> with the full 32 locus set (<xref ref-type="table" rid="pone-0048638-t003">Table 3</xref>). The overall geometric mean of the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e132" xlink:type="simple"/></inline-formula> rates for the 32 loci from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> was 0.0020 per generation, and for the rates from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref> 0.0021, a 5% difference, compared to a mean of 0.0026 for <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> and 0.0032 for <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>, a 23% difference for the arithmetic average.</p>
      <table-wrap id="pone-0048638-t003" position="float">
        <object-id pub-id-type="doi">10.1371/journal.pone.0048638.t003</object-id>
        <label>Table 3</label>
        <caption>
          <title>Comparison of arithmetic and geometric means for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e133" xlink:type="simple"/></inline-formula>.</title>
        </caption>
        <alternatives>
          <graphic id="pone-0048638-t003-3" position="float" mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0048638.t003" xlink:type="simple"/>
          <table>
            <colgroup span="1">
              <col align="left" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
              <col align="center" span="1"/>
            </colgroup>
            <thead>
              <tr>
                <td align="left" rowspan="1" colspan="1">Mutation rates</td>
                <td align="left" rowspan="1" colspan="1">No. of loci</td>
                <td colspan="2" align="left" rowspan="1">Geometric mean</td>
                <td colspan="2" align="left" rowspan="1">Arithmetic mean</td>
                <td colspan="2" align="left" rowspan="1"><italic>D</italic> (% difference)</td>
              </tr>
            </thead>
            <tbody>
              <tr>
                <td align="left" rowspan="1" colspan="1">Burgarella</td>
                <td colspan="2" align="left" rowspan="1">32</td>
                <td align="left" rowspan="1" colspan="1">337.8274</td>
                <td align="left" rowspan="1" colspan="1">348.7863</td>
                <td colspan="3" align="left" rowspan="1">3.1%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">Ballantyne</td>
                <td colspan="2" align="left" rowspan="1">15</td>
                <td align="left" rowspan="1" colspan="1">294.3013</td>
                <td align="left" rowspan="1" colspan="1">259.0993</td>
                <td colspan="3" align="left" rowspan="1">13.6%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">Ballantyne</td>
                <td colspan="2" align="left" rowspan="1">32</td>
                <td align="left" rowspan="1" colspan="1">321.066</td>
                <td align="left" rowspan="1" colspan="1">280.7067</td>
                <td colspan="3" align="left" rowspan="1">14.4%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">Burgarella</td>
                <td colspan="2" align="left" rowspan="1">15</td>
                <td align="left" rowspan="1" colspan="1">360.613</td>
                <td align="left" rowspan="1" colspan="1">314.5164</td>
                <td colspan="3" align="left" rowspan="1">14.7%</td>
              </tr>
              <tr>
                <td align="left" rowspan="1" colspan="1">YHRD</td>
                <td colspan="2" align="left" rowspan="1">15</td>
                <td align="left" rowspan="1" colspan="1">367.785</td>
                <td align="left" rowspan="1" colspan="1">317.3084</td>
                <td colspan="3" align="left" rowspan="1">15.9%</td>
              </tr>
            </tbody>
          </table>
        </alternatives>
        <table-wrap-foot>
          <fn id="nt103">
            <p>Difference (<italic>D</italic>) from <xref ref-type="disp-formula" rid="pone.0048638.e071">equation (2</xref>) is a measure of the overall conformity of the data to the distribution expected under S-SMM. The lowest percentage difference indicates the least amount of departure from the S-SMM.</p>
          </fn>
        </table-wrap-foot>
      </table-wrap>
      <p>To test for the effect of different growth model selections on <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e134" xlink:type="simple"/></inline-formula> <italic>Ytime</italic> was run under several growth rate scenarios, varying by orders of magnitude from zero growth to ‘star’ for the STR data set (n = 235). Substitution of the no-growth model (<italic>Rgrowth</italic> = 0) for the ‘star’ genealogy model produced no change in the point estimate <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e135" xlink:type="simple"/></inline-formula> but did increase the width of the 95% confidence intervals substantially, by approximately 54% for the no-growth model compared to the star genealogy. For any value of <italic>Rgrowth</italic> at or above 1000, the average increase in 95% confidence intervals over the star genealogy was 21% or less, approaching a limit of about 7% above 1.0E+10 (<xref ref-type="fig" rid="pone-0048638-g004">Figure 4</xref>).</p>
      <fig id="pone-0048638-g004" position="float">
        <object-id pub-id-type="doi">10.1371/journal.pone.0048638.g004</object-id>
        <label>Figure 4</label>
        <caption>
          <title>32 locus STR data set with changing growth rates and genealogy shapes.</title>
          <p>A) <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e136" xlink:type="simple"/></inline-formula>and 95% confidence intervals plotted for the STR data set (n = 245, 32 loci) as the <italic>Rgrowth</italic> (N<sub>e</sub>*r) parameter is varied between <italic>Rgrowth</italic> = 0 (no growth, constant sized population) to a “star” genealogy (approximating infinite exponential growth by all descendant lines) by orders of magnitude from 1.0E+1 to 1.0E+11. <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e137" xlink:type="simple"/></inline-formula> = 345.09 was the median value for all <italic>Rgrowth</italic> models and was shown empirically to be independent of the shape of the sampled genealogy, growth rate and population size. B) Percentage change of 95% CI overestimation compared to ‘star’ genealogy for the full range of comparisons in 4A.</p>
        </caption>
        <graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0048638.g004" position="float" xlink:type="simple"/>
      </fig>
    </sec>
    <sec id="s4">
      <title>Discussion</title>
      <p>A fundamental question for any Y-STR based research study or application is: how many STR loci are needed and which of the available loci will be the most useful for the purpose at hand? The purpose of this study was to provide an exploratory framework for the analysis and selection of the most suitable currently available Y-STRs for calculation of the coalescent and for the future evaluation of those STRs on the Y chromosome yet to be identified. In the process, two statistics for quantifying the degree of departure by any Y-STR from the S-SMM were developed, with <italic>q</italic> (<xref ref-type="disp-formula" rid="pone.0048638.e043">equation 1</xref>) dependent only upon the availability of empirical data concerning allele repeat values and frequencies and <italic>D</italic> (<xref ref-type="disp-formula" rid="pone.0048638.e079">equation 3</xref>), a normalized statistic independent of any parameter save the need to provide a reasonable estimate of the average mutation rate <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e138" xlink:type="simple"/></inline-formula> across all loci compared.</p>
      <p>In principle, the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e139" xlink:type="simple"/></inline-formula> statistic developed here can be used effectively with any STR to evaluate its suitability for the purpose of calculating <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e140" xlink:type="simple"/></inline-formula> and, given a large enough allele data sample, to allow an evaluation of its behavior with regard to the S-SMM. <xref ref-type="disp-formula" rid="pone.0048638.e043">Equation (1</xref>) may be incorporated easily into a spreadsheet with the caveat that the method of calculating kurtosis used by many spreadsheets returns an “excess” kurtosis value (kurtosis in excess of 3.0), therefore the parameter<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e141" xlink:type="simple"/></inline-formula> (kurtosis minus 3.0) from <xref ref-type="disp-formula" rid="pone.0048638.e043">Equation (1</xref>) becomes <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e142" xlink:type="simple"/></inline-formula> (kurtosis in excess of 3.0).</p>
      <p>To illustrate the procedure, 93 ssSTRs routinely available from commercial testing companies were evaluated using those haplotypes from the SDS that were tested for all 93 loci (n = 235) in <italic>Excel</italic>, with the results ranked from best to worst in Supplementary <xref ref-type="supplementary-material" rid="pone.0048638.s001">Table S1</xref> (ST1). Differences in rank order between <xref ref-type="table" rid="pone-0048638-t002">Table 2</xref> and ST1 were attributable to sample size and sampling error, with the SMGF sample far exceeding the SDS (n≈35,600 vs. n = 235), resulting in some variation in allele distribution at any given locus compared between samples. Because of these differences, the results in <xref ref-type="table" rid="pone-0048638-t002">Table 2</xref> were regarded as more robust. Some loci found in ST1 clearly do not provide an accurate estimate of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e143" xlink:type="simple"/></inline-formula> due to excessive kurtosis, skewness, a narrow range, or a combination of all three (e.g., DYS472). Inclusion of loci that conform poorly to the SMM in coalescent calculations will result in inaccurate estimates of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e144" xlink:type="simple"/></inline-formula>. Considering the results from the two data sets and <xref ref-type="fig" rid="pone-0048638-g001">Figure 1</xref> as a whole, a “cutoff” value of <italic>q</italic> = 0.07 is suggested, permitting only loci that have good to excellent conformity to the S-SMM to be used when estimating <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e145" xlink:type="simple"/></inline-formula>.</p>
      <p>ASD is not dependent on population size or the specific shape of the genealogy <xref ref-type="bibr" rid="pone.0048638-Sun1">[16]</xref>, <xref ref-type="bibr" rid="pone.0048638-Stumpf1">[17]</xref> and because <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e146" xlink:type="simple"/></inline-formula>is dependent only upon ASD and<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e147" xlink:type="simple"/></inline-formula> for its calculation, it too is independent of population size and genealogy shape. By examining different growth models, from a no-growth, constant size population to one with infinite growth, it was confirmed empirically (<xref ref-type="fig" rid="pone-0048638-g004">Figure 4</xref>) that the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e148" xlink:type="simple"/></inline-formula>estimate was unaffected by the specific genealogy of the sample but that confidence intervals were affected, with the no-growth, constant-size model having the most conservative (widest) CIs. Therefore, any comparison of the arithmetic and geometric means (<xref ref-type="disp-formula" rid="pone.0048638.e079">equation 3</xref>) will be independent of population size or genealogical history as only the point estimates <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e149" xlink:type="simple"/></inline-formula> are used in calculating <italic>D</italic>.</p>
      <p>A basic limitation in calculating the coalescent from ssSTRs is that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e150" xlink:type="simple"/></inline-formula>cannot be extended beyond the MRCA for all members of a given haplogroup. Each Y haplogroup must coalesce no earlier than the first individual male who acquired the defining Y-SNP <xref ref-type="bibr" rid="pone.0048638-de1">[9]</xref>. It is possible that the MRCA for a haplogroup currently represented in a living population was not the first person to acquire the SNP, but merely the most recent common ancestor of those particular Y lines that have survived to the present, with other lines from the same founding haplotype having gone extinct. Thus, although ASD is an unbiased estimator of the TMRCA for a specific haplogroup for any length of time up to at least 2 million years before present <xref ref-type="bibr" rid="pone.0048638-Sun1">[16]</xref>, it cannot always provide an estimate of the actual founding time of the haplogroup.</p>
      <p>At the opposite end of the time scale a very recent founding event, with the defining Y-SNP having emerged only in the past several generations, may result in a population that has not yet have acquired sufficient genetic diversity among descendants of the founder to calculate <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e151" xlink:type="simple"/></inline-formula> accurately, because STR mutation is a stochastic process and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e152" xlink:type="simple"/></inline-formula> is only the central estimate of the average mutation rate over time. This limitation is especially important for samples that have few ssSTR loci to compare. The largest differences between adjacent <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e153" xlink:type="simple"/></inline-formula> values, as the number of loci was reduced, occurred when there were fewer than six loci compared (<xref ref-type="fig" rid="pone-0048638-g002">Figure 2</xref>) and suggested that a minimum of six ssSTR loci with good conformity to the S-SMM were needed when calculating <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e154" xlink:type="simple"/></inline-formula> to avoid wholly unreliable estimates. More than six loci should be used whenever possible, regardless of time depth, to further improve the accuracy of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e155" xlink:type="simple"/></inline-formula>as long as the additional ssSTRs also conform closely to the S-SMM. The more loci compared, the narrower the confidence intervals will become, improving precision as well as accuracy. The use of many ssSTRs per haplotype reduces the problem of shallow TMRCAs because there is a higher cumulative probability that at least some loci have undergone mutation within the past few generations. Even after dropping DYS389I/II, DYS438 and DYS392 from the calculation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e156" xlink:type="simple"/></inline-formula>, there were still 45 ssSTRs remaining in Supplementary <xref ref-type="supplementary-material" rid="pone.0048638.s001">Table S1</xref> with <italic>q</italic> &lt;0.07, a sufficiently high number of loci for most purposes and greatly exceeding the resolution of the 16 “standard” YHRD loci used by many studies presently.</p>
      <p>The size of errors associated with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e157" xlink:type="simple"/></inline-formula> depended partly upon the range of the mutation rate, with higher rates having smaller errors <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>. No significant relationship was found, however, between the minimum allele length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e158" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e159" xlink:type="simple"/></inline-formula> for the 32 Y STR loci from SMGF for either <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref> or <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> or between the minimum allele length <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e160" xlink:type="simple"/></inline-formula> for a given locus and its calculated <italic>q</italic> value. Furthermore, the failure of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e161" xlink:type="simple"/></inline-formula> to correlate significantly with ASD for two of the three sets of mutation rate estimates (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3</xref>) suggested that some mutation rate estimates introduced substantial, uncorrected errors to the calculation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e162" xlink:type="simple"/></inline-formula>, as the allele values (and the resulting oASD statistics) entering each calculation were the same for the particular locus being compared and only <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e163" xlink:type="simple"/></inline-formula> was varied.</p>
      <p>The sharp increase in the proportion of explained variability between the 32 locus data set and the reduced 24 locus set (using <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e164" xlink:type="simple"/></inline-formula> from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> – <xref ref-type="fig" rid="pone-0048638-g003">Figure 3A,D</xref>) is notable and suggests that estimates for loci with mutation rates below 0.001 per generation may not be suitable presently for the calculation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e165" xlink:type="simple"/></inline-formula>, especially when there is little or no father-son direct observational evidence available for the estimate. The absence of a significant interaction between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e166" xlink:type="simple"/></inline-formula> and ASD for any of the comparisons made, though, suggests that the relationship between these two factors is not dependent upon any absolute value of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e167" xlink:type="simple"/></inline-formula>, but rather is independent of the rates’ ranges. Put differently, a low, medium or high <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e168" xlink:type="simple"/></inline-formula> rate is not likely to be the cause for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e169" xlink:type="simple"/></inline-formula>’s failure to predict ASD, but rather that the estimation errors for low rates are greater than for higher rates due to fewer available direct observations from which to estimate the “true” rate for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e170" xlink:type="simple"/></inline-formula>. If the true <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e171" xlink:type="simple"/></inline-formula> for all loci compared actually were available, any remaining errors in the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e172" xlink:type="simple"/></inline-formula>estimate would be reduced to those introduced by stochasticity, genetic drift and sampling error. By removing those loci for which <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e173" xlink:type="simple"/></inline-formula>&lt;0.001 from the calculation, the amount of unexplained variance due to factors such as error in estimates of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e174" xlink:type="simple"/></inline-formula>, sampling error, unequal numbers of offspring and stochasticity in the mutation process was reduced to 24.7% for the 24 locus subset (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3D</xref>, with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e175" xlink:type="simple"/></inline-formula> estimates from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>), a substantial improvement over the first comparison (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3A</xref>), where 58.2% of the total variation was left unexplained.</p>
      <p>Linear regression analyses of YHRD loci with and without DYS392, DYS438 and DYS389I highlighted significant bias in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e176" xlink:type="simple"/></inline-formula>when these loci first were included (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3C</xref>) and then removed (<xref ref-type="fig" rid="pone-0048638-g003">Figure 3D</xref>). Their removal from the YHRD loci strengthened correlation from a non-significant result to a significant one (R<sup>2</sup> = 0.4921, with mutation rates from <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>), suggesting that estimates based on the 12 selected YHRD ssSTR loci are reasonably reliable if no higher resolution data is available. Using mutation rate estimates from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> with the YHRD loci, the correlation coefficient for the 12 selected YHRD loci improved slightly more, with R<sup>2</sup> = 0.5027. Because many studies have been published with only the standard YHRD loci available in the data set, the improvements described here are directly applicable to the calculation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e177" xlink:type="simple"/></inline-formula> from these data. Further improvements seen here in the accuracy of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e178" xlink:type="simple"/></inline-formula> resulted directly from an increase in the number of loci from 12 to 24 (from 50.27% to 75.35% of variation in ASD explained by <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e179" xlink:type="simple"/></inline-formula>, mutation rate estimates from <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>), strongly supporting the need for increased numbers of ssSTR loci when planning research projects that include a calculation of the coalescent.</p>
      <p>Given the significant relationship detected between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e180" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e181" xlink:type="simple"/></inline-formula>, it was reasonable to infer that the specific choice of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e182" xlink:type="simple"/></inline-formula> for each locus had the strongest effect on the calculated value of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e183" xlink:type="simple"/></inline-formula>. Because <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e184" xlink:type="simple"/></inline-formula>, any upward bias in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e185" xlink:type="simple"/></inline-formula> introduced by use of the arithmetic mean when averaging mutation rate estimates across loci had the effect of biasing <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e186" xlink:type="simple"/></inline-formula> downward. Use of the geometric mean provided a more consistent measure of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e187" xlink:type="simple"/></inline-formula> when using different sets of loci because it returned the median value for the average of the observed per-locus variances (oASD) and for the individual <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e188" xlink:type="simple"/></inline-formula> values rather than a potentially skewed <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e189" xlink:type="simple"/></inline-formula> based on the arithmetic means for one or both variables.</p>
      <p>While using the geometric mean to determine the average within-locus observed variance across loci (oASD), precautions were taken to remove any ssSTRs from the sample data set that had zero variance (i.e., all haplotypes have the same allele value as the ancestral haplotype at the locus of interest) because normal computer-based methods of calculating the geometric mean would have resulted in an undefined result for log(0). The issue of zero allelic variance within a specific locus is more likely to occur within data sets that have only a few haplotypes or loci being compared. Within a larger sample, any locus exhibiting zero variance between the sample and the ancestral (median) allele value is fundamentally uninformative with regard to the calculation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e190" xlink:type="simple"/></inline-formula> and may be deviating substantially from the S-SMM. These uninformative loci, lacking in variation, should be dropped from calculations regardless (even while using the arithmetic mean) because a substantial downward bias of the average between-locus variance may be introduced by them, resulting in an underestimation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e191" xlink:type="simple"/></inline-formula>. For example, it was clear from <xref ref-type="fig" rid="pone-0048638-g003">Figure 3A and 3B</xref> that DYS454 did not appear to be accumulating any appreciable amount of variance, especially given its relatively fast mutation rate estimate (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e192" xlink:type="simple"/></inline-formula> 0.002182 by <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref>). It was suggested by the empirical data that this locus may be undergoing microsatellite death. Regardless of the cause, DYS454 was uninformative with regard to the coalescent in this context and its removal from STR haplotype data sets was necessary to avoid a downward bias in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e193" xlink:type="simple"/></inline-formula>.</p>
      <p>Though there is no reason to believe that mutational behavior for ssSTRs will differ between male Y haplogroups or geographic regions, it should be noted that both population samples used here (SMGF, SDS) contain Y haplotypes dominated by male lineages originating in Europe and may not be fully representative of ssSTR allele distributions in other regions of the world. Some Y-STR databases, such as SMGF, have many loci per haplotype available and contain many haplotypes, but are biased toward descendants of a particular ancestral population (European), due to their sample collection strategies. Others, such as YHRD, have excellent coverage of many geographic regions but relatively few loci available. Because of this, <italic>q</italic> values developed using the SMGF and SDS data regarding the suitability of a particular Y STR locus for estimation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e194" xlink:type="simple"/></inline-formula>using the S-SMM and reported in <xref ref-type="table" rid="pone-0048638-t002">Table 2</xref> and ST1 should be viewed as preliminary, though the very large sample size for the SMGF data set is likely to compensate for any unknown bias introduced by incomplete geographic coverage. Further research, including a worldwide survey of Y STR allele frequencies at each of the loci described in <xref ref-type="table" rid="pone-0048638-t002">Table 2</xref> and in ST1, will help to further refine the results presented here concerning the conformity of a given ssSTR locus to the S-SMM.</p>
      <p>Even within the geographic limitations of the available test data, though, the use of <xref ref-type="disp-formula" rid="pone.0048638.e043">equations (1</xref>) and (3) in choosing and evaluating STRs lead to substantially improved estimates of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e195" xlink:type="simple"/></inline-formula>. When combined with the use of the geometric mean in place of the arithmetic mean for averaging both ASD and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e196" xlink:type="simple"/></inline-formula>, the overall percentage of unexplained error for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e197" xlink:type="simple"/></inline-formula> becomes mainly dependent on sample size, variations in numbers of offspring, and the inherent stochasticity of the mutation and gamete selection processes, leading to more robust estimates of the coalescent <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e198" xlink:type="simple"/></inline-formula> with less unexplained error. Having taken into account factors of skew, kurtosis, allele value ranges and having identified those loci that are likely to contribute bias when calculating the coalescent, it also is evident from the statistical analysis presented here that the accuracy of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e199" xlink:type="simple"/></inline-formula> calculations will benefit most from future refinements in the estimation of ssSTR mutation rates, with the logistic regression model <xref ref-type="bibr" rid="pone.0048638-Burgarella1">[4]</xref> currently providing <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e200" xlink:type="simple"/></inline-formula> estimates that are more consistent with ASD than the Bayesian posterior distribution analysis approach <xref ref-type="bibr" rid="pone.0048638-Ballantyne1">[3]</xref>. The problem is mitigated, however, by the use of the geometric mean for averaging both ASD and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e201" xlink:type="simple"/></inline-formula> when calculating <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0048638.e202" xlink:type="simple"/></inline-formula> rather than the arithmetic mean.</p>
    </sec>
    <sec id="s5">
      <title>Supporting Information</title>
      <supplementary-material id="pone.0048638.s001" mimetype="application/x-excel" xlink:href="info:doi/10.1371/journal.pone.0048638.s001" position="float" xlink:type="simple">
        <label>Table S1</label>
        <caption>
          <p><bold>List of 93 commercially available Y chromosomes ssSTRs ranked by </bold><bold><italic>q</italic></bold><bold> value.</bold> Lowest <italic>q</italic> value equals best conformity to the S-SMM.</p>
          <p>(XLS)</p>
        </caption>
      </supplementary-material>
    </sec>
  </body>
  <back>
    <ack>
      <p>I would like to extend my greatest appreciation and thanks to Dr. Floyd W. Weckerly, Texas State University Department of Biology, for many helpful comments and conversations concerning the statistical analysis used in this study. Thank you to the academic reviewer and the two anonymous reviewers for their helpful and insightful comments, which greatly improved sections of this study. Lastly, thanks go to the “British Isles DNA Project,” Roy Keys and Andrew Lancaster, project administrators, for their assistance in providing Y-STR data collected by the project for this analysis.</p>
    </ack>
    <ref-list>
      <title>References</title>
      <ref id="pone.0048638-Jobling1">
        <label>1</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Jobling</surname><given-names>MA</given-names></name>, <name name-style="western"><surname>Tyler-Smith</surname><given-names>C</given-names></name> (<year>2003</year>) <article-title>The human Y chromosome: an evolutionary marker comes of age</article-title>. <source>Nat Rev Genet</source> <volume>4</volume>: <fpage>598</fpage>–<lpage>612</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Walsh1">
        <label>2</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Walsh</surname><given-names>B</given-names></name> (<year>2000</year>) <article-title>Estimating the time to the most recent common ancestor for the Y chromosome or mitochondrial DNA for a pair of individuals</article-title>. <source>Genetics</source> <volume>156</volume>: <fpage>897</fpage>–<lpage>912</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Ballantyne1">
        <label>3</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Ballantyne</surname><given-names>K</given-names></name>, <name name-style="western"><surname>Goedbloed</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Fang</surname><given-names>R</given-names></name>, <name name-style="western"><surname>Schaap</surname><given-names>O</given-names></name>, <name name-style="western"><surname>Lao</surname><given-names>O</given-names></name>, <etal>et al</etal>. (<year>2010</year>) <article-title>Mutability of Y- Chromosomal Microsatellites: Rates, Characteristics, Molecular Bases, and Forensic Implications</article-title>. <source>Am J Hum Genet</source> <volume>87</volume>: <fpage>341</fpage>–<lpage>353</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Burgarella1">
        <label>4</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Burgarella</surname><given-names>C</given-names></name>, <name name-style="western"><surname>Navascués</surname><given-names>M</given-names></name> (<year>2011</year>) <article-title>Mutation rate estimates for 110 Y-chromosome STRs combining population and father–son pair data</article-title>. <source>Eur J Hum Genet</source> <volume>19</volume>: <fpage>70</fpage>–<lpage>75</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Ohta1">
        <label>5</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Ohta</surname><given-names>T</given-names></name>, <name name-style="western"><surname>Kimura</surname><given-names>M</given-names></name> (<year>1973</year>) <article-title>A model of mutation appropriate to estimate the number of electrophoretically detectable alleles in a finite population</article-title>. <source>Genet Res</source> <volume>22</volume>: <fpage>201</fpage>–<lpage>204</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Kimura1">
        <label>6</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Kimura</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Ohta</surname><given-names>T</given-names></name> (<year>1978</year>) <article-title>Stepwise mutation model and distribution of allelic frequencies in a finite population</article-title>. <source>Proc Natl Acad Sci U S A</source> <volume>75</volume>: <fpage>2868</fpage>–<lpage>2872</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Nachman1">
        <label>7</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Nachman</surname><given-names>MW</given-names></name>, <name name-style="western"><surname>Crowell</surname><given-names>SI</given-names></name> (<year>2000</year>) <article-title>Estimate of the mutation rate per nucleotide in humans</article-title>. <source>Genetics</source> <volume>156</volume>: <fpage>297</fpage>–<lpage>304</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Behar1">
        <label>8</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Behar</surname><given-names>DM</given-names></name>, <name name-style="western"><surname>Metspalu</surname><given-names>E</given-names></name>, <name name-style="western"><surname>Kivisild</surname><given-names>T</given-names></name>, <name name-style="western"><surname>Achilli</surname><given-names>A</given-names></name>, <name name-style="western"><surname>Hadid</surname><given-names>Y</given-names></name>, <etal>et al</etal>. (<year>2006</year>) <article-title>“The matrilineal ancestry of Ashkenazi Jewry: portrait of a recent founder event”</article-title>. <source>Am J Hum Genet 78</source> <volume>(3)</volume>: <fpage>487</fpage>–<lpage>97</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-de1">
        <label>9</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>de</surname><given-names>Knijff</given-names></name> (<year>2000</year>) <collab xlink:type="simple">P</collab> (<year>2000</year>) <article-title>Messages through bottlenecks: on the combined use of slow and fast evolving polymorphic markers on the human Y chromosome</article-title>. <source>Am J Hum Genet</source> <volume>67</volume>: <fpage>1055</fpage>–<lpage>1061</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Hudson1">
        <label>10</label>
        <mixed-citation publication-type="other" xlink:type="simple">Hudson RR, (1991) Genetic genealogies and the coalescent process. Oxford Surveys in Evolutionary Biology 7: 1–44. Oxford: Oxford Univ Press.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Behar2">
        <label>11</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Behar</surname><given-names>DM</given-names></name>, <name name-style="western"><surname>Thomas</surname><given-names>MG</given-names></name>, <name name-style="western"><surname>Skorecki</surname><given-names>K</given-names></name>, <name name-style="western"><surname>Hammer</surname><given-names>MF</given-names></name>, <name name-style="western"><surname>Bulygina</surname><given-names>E</given-names></name>, <etal>et al</etal>. (<year>2003</year>) <article-title>Multiple origins of Ashkenazi Levites: Y chromosome evidence for both Near Eastern and European ancestries</article-title>. <source>Am J Hum Genet</source> <volume>73</volume>: <fpage>768</fpage>–<lpage>779</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Goldstein1">
        <label>12</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Goldstein</surname><given-names>DB</given-names></name>, <name name-style="western"><surname>Linares</surname><given-names>AR</given-names></name>, <name name-style="western"><surname>Cavalli-Sforza</surname><given-names>LL</given-names></name>, <name name-style="western"><surname>Feldman</surname><given-names>MW</given-names></name> (<year>1995</year>) <article-title>Genetic absolute dating based on microsatellites and the origin of modern humans</article-title>. <source>Proc Natl Acad Sci U S A</source> <volume>92</volume>: <fpage>6723</fpage>–<lpage>6727</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Slatkin1">
        <label>13</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><article-title>Slatkin, M. 1995. A measure of population subdivision based on microsatellite allele frequencies</article-title>. <source>Genetics</source> <volume>139</volume>: <fpage>457</fpage>–<lpage>462</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Goldstein2">
        <label>14</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Goldstein</surname><given-names>DF</given-names></name>, <name name-style="western"><surname>Zhivotovsky</surname><given-names>LA</given-names></name>, <name name-style="western"><surname>Nayar</surname><given-names>K</given-names></name>, <name name-style="western"><surname>Linares</surname><given-names>AR</given-names></name>, <name name-style="western"><surname>Cavalli-Sforza</surname><given-names>LL</given-names></name>, <etal>et al</etal>. (<year>1996</year>) <article-title>Statistical properties of the variation at linked microsatellite loci: implications for the history of human Y chromosomes</article-title>. <source>Mol Biol Evol</source> <volume>13</volume>: <fpage>1213</fpage>–<lpage>1218</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Goldstein3">
        <label>15</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Goldstein</surname><given-names>DF</given-names></name>, <name name-style="western"><surname>Linares</surname><given-names>AR</given-names></name>, <name name-style="western"><surname>Cavalli-Sforza</surname><given-names>LL</given-names></name>, <name name-style="western"><surname>Feldman</surname><given-names>MW</given-names></name> (<year>1995</year>) <article-title>An evaluation of genetic distances for use with microsatellite loci</article-title>. <source>Genetics</source> <volume>139</volume>: <fpage>463</fpage>–<lpage>471</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Sun1">
        <label>16</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sun</surname><given-names>JX</given-names></name>, <name name-style="western"><surname>Mullikin</surname><given-names>JC</given-names></name>, <name name-style="western"><surname>Patterson</surname><given-names>N</given-names></name> (<year>2009</year>) <name name-style="western"><surname>Reich</surname><given-names>DE</given-names></name> (<year>2009</year>) <article-title>Microsatellites are molecular clocks that support accurate inferences about history</article-title>. <source>Mol Biol Evol</source> <volume>26(5)</volume>: <fpage>1017</fpage>–<lpage>27</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Stumpf1">
        <label>17</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Stumpf</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Goldstein</surname><given-names>D</given-names></name> (<year>2001</year>) <article-title>Genealogical and Evolutionary Inference with the Human Y Chromosome</article-title>. <source>Science</source> <volume>291</volume>: <fpage>1738</fpage>–<lpage>1742</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Zhivotovsky1">
        <label>18</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Zhivotovsky</surname><given-names>LA</given-names></name>, <name name-style="western"><surname>Underhill</surname><given-names>PA</given-names></name>, <name name-style="western"><surname>Feldman</surname><given-names>MW</given-names></name> (<year>2006</year>) <article-title>Difference between evolutionarily effective and germ line mutation rate due to stochastically varying haplogroup size</article-title>. <source>Mol Biol Evol</source> <volume>23</volume>: <fpage>2268</fpage>–<lpage>2270</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Balaresque1">
        <label>19</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Balaresque</surname><given-names>P</given-names></name>, <name name-style="western"><surname>Bowden</surname><given-names>GR</given-names></name>, <name name-style="western"><surname>Adams</surname><given-names>SM</given-names></name>, <name name-style="western"><surname>Leung</surname><given-names>HY</given-names></name>, <name name-style="western"><surname>King</surname><given-names>TE</given-names></name>, <etal>et al</etal>. (<year>2010</year>) <article-title>A predominantly neolithic origin for European paternal lineages</article-title>. <source>PLoS Biol</source> <volume>8</volume>: <fpage>e1000285</fpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Busby1">
        <label>20</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Busby</surname><given-names>GB</given-names></name>, <name name-style="western"><surname>Brisighelli</surname><given-names>F</given-names></name>, <name name-style="western"><surname>Sánchez-Diz</surname><given-names>P</given-names></name>, <name name-style="western"><surname>Ramos-Luis</surname><given-names>E</given-names></name>, <name name-style="western"><surname>Martinez-Cadenas</surname><given-names>C</given-names></name>, <etal>et al</etal>. (<year>2012</year>) <article-title>The peopling of Europe and the cautionary tale of Y chromosome lineage R-M269</article-title>. <source>Proc Biol Sci</source> <volume>279</volume>: <fpage>884</fpage>–<lpage>92</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Goedbloed1">
        <label>21</label>
        <mixed-citation publication-type="other" xlink:type="simple">Goedbloed M, Vermeulen M, Fang, RN Lembring, M Wollstein, A, etal. (2009) Comprehensive mutation analysis of 17 Y-chromosomal short tandem repeat polymorphisms included in the AmpFlSTR Yfiler PCR amplification kit. Int J Legal Med 123, 471–482.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Vermeulen1">
        <label>22</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Vermeulen</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Wollstein</surname><given-names>A</given-names></name>, <name name-style="western"><surname>van der</surname><given-names>Gaag</given-names></name>, <name name-style="western"><surname>K</surname><given-names>Lao</given-names></name>, <name name-style="western"><surname>O</surname><given-names>Xue</given-names></name>, <name name-style="western"><surname>Y, et</surname><given-names>al</given-names></name> (<year>2009</year>) <article-title>Improving global and regional resolution of male lineage differentiation by simple single-copy Y-chromosomal short tandem repeat polymorphisms</article-title>. <source>Forensic Sci Int Genet</source> <volume>3</volume>: <fpage>205</fpage>–<lpage>213</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Gusmo1">
        <label>23</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Gusmão</surname><given-names>L</given-names></name>, <name name-style="western"><surname>Krawczak</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Sánchez-Diz</surname><given-names>P</given-names></name>, <name name-style="western"><surname>Alvas</surname><given-names>C</given-names></name>, <name name-style="western"><surname>Lopes</surname><given-names>A</given-names></name>, <etal>et al</etal>. (<year>2003</year>) <article-title>Bimodal allele frequency distribution at Y-STR loci DYS392 and DYS438: no evidence for a deviation from the stepwise mutation model</article-title>. <source>Int J Legal Med</source> <volume>117</volume>: <fpage>287</fpage>–<lpage>290</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Kimmel1">
        <label>24</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Kimmel</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Chakraborty</surname><given-names>R</given-names></name> (<year>1996</year>) <article-title>Measures of variation at DNA repeat loci under a general stepwise mutation model</article-title>. <source>Theor Popul Biol</source> <volume>50</volume>: <fpage>345</fpage>–<lpage>67</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Calabrese1">
        <label>25</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Calabrese</surname><given-names>PP</given-names></name>, <name name-style="western"><surname>Durrett</surname><given-names>RT</given-names></name>, <name name-style="western"><surname>Aquadro</surname><given-names>CF</given-names></name> (<year>2001</year>) <article-title>Dynamics of microsatellite divergence under stepwise mutation and proportional slippage/point mutation models</article-title>. <source>Genetics</source> <volume>159</volume>: <fpage>839</fpage>–<lpage>852</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Patel1">
        <label>26</label>
        <mixed-citation publication-type="other" xlink:type="simple">Patel JK, Read CB (1996) Handbook of the normal distribution, 2<sup>nd</sup> ed. Boca Raton: CRC Press. 431 p.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Uchida1">
        <label>27</label>
        <mixed-citation publication-type="other" xlink:type="simple">Uchida Y (2008) A simple proof of the geometric-arithmetic mean inequality. J Inequal Pure and Appl Math 9:Art. 56, 2 pp.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Mallows1">
        <label>28</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Mallows</surname><given-names>C</given-names></name> (<year>1991</year>) <article-title>Another comment on O’Cinneide</article-title>. <source>Am Stat</source> <volume>45</volume>: <fpage>257</fpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Cartwright1">
        <label>29</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Cartwright</surname><given-names>D</given-names></name>, <name name-style="western"><surname>Field</surname><given-names>M</given-names></name> (<year>1978</year>) <article-title>A Refinement of the Arithmetic Mean-Geometric Mean Inequality</article-title>. <source>Proc Am Math Soc</source> <volume>71</volume>: <fpage>36</fpage>–<lpage>38</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Mills1">
        <label>30</label>
        <mixed-citation publication-type="other" xlink:type="simple">Mills LS (2007) Conservation of Wildlife Populations: Demography, Genetics and Management. Maldon: Blackwell Publishing, 424 p.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Schwertman1">
        <label>31</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Schwertman</surname><given-names>NC</given-names></name>, <name name-style="western"><surname>Gilks</surname><given-names>AJ</given-names></name>, <name name-style="western"><surname>Cameron</surname><given-names>J</given-names></name> (<year>1990</year>) <article-title>A simple noncalculus proof that the median minimizes the sum of the absolute deviations</article-title>. <source>Am Statistician</source> <volume>44</volume>: <fpage>38</fpage>–<lpage>39</lpage>.</mixed-citation>
      </ref>
      <ref id="pone.0048638-Sengupta1">
        <label>32</label>
        <mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sengupta</surname><given-names>S</given-names></name>, <name name-style="western"><surname>Zhivotovsky</surname><given-names>LA</given-names></name>, <name name-style="western"><surname>King</surname><given-names>R</given-names></name>, <name name-style="western"><surname>Mehdi</surname><given-names>SQ</given-names></name>, <name name-style="western"><surname>Edmunds</surname><given-names>CA</given-names></name>, <etal>et al</etal>. (<year>2006</year>) <article-title>Polarity and temporality of high-resolution y-chromosome distributions in India identify both indigenous and exogenous expansions and reveal minor genetic influence of central Asian pastoralists</article-title>. <source>Am J Hum Genet</source> <volume>78</volume>: <fpage>202</fpage>–<lpage>221</lpage>.</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>