<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="discussion" dtd-version="3.0" xml:lang="EN">
    <front>
        <journal-meta><journal-id journal-id-type="publisher-id">plos</journal-id><journal-id journal-id-type="nlm-ta">PLoS Comput Biol</journal-id><journal-id journal-id-type="pmc">ploscomp</journal-id><!--===== Grouping journal title elements =====--><journal-title-group><journal-title>PLoS Computational Biology</journal-title></journal-title-group><issn pub-type="ppub">1553-734X</issn><issn pub-type="epub">1553-7358</issn><publisher>
                <publisher-name>Public Library of Science</publisher-name>
                <publisher-loc>San Francisco, USA</publisher-loc>
            </publisher></journal-meta>
        <article-meta><article-id pub-id-type="publisher-id">08-PLCB-MI-1074R3</article-id><article-id pub-id-type="doi">10.1371/journal.pcbi.1000366</article-id><article-categories>
                <subj-group subj-group-type="heading">
                    <subject>Message from ISCB</subject>
                </subj-group>
                <subj-group subj-group-type="Discipline">
                    <subject>Biochemistry/Bioinformatics</subject>
                    <subject>Computational Biology/Systems Biology</subject>
                </subj-group>
            </article-categories><title-group><article-title>Getting Started in Computational Mass Spectrometry–Based
                    Proteomics</article-title></title-group><contrib-group>
                <contrib contrib-type="author" xlink:type="simple">
                    <name name-style="western">
                        <surname>Vitek</surname>
                        <given-names>Olga</given-names>
                    </name>
                    <xref ref-type="aff" rid="aff1"/>
                    <xref ref-type="corresp" rid="cor1">
                        <sup>*</sup>
                    </xref>
                </contrib>
            </contrib-group><aff id="aff1">
                <addr-line>Departments of Statistics and Computer Science, Purdue University, West
                    Lafayette, Indiana, United States of America</addr-line>
            </aff><contrib-group>
                <contrib contrib-type="editor" xlink:type="simple">
                    <name name-style="western">
                        <surname>Troyanskaya</surname>
                        <given-names>Olga G.</given-names>
                    </name>
                    <role>Editor</role>
                    <xref ref-type="aff" rid="edit1"/>
                </contrib>
            </contrib-group><aff id="edit1">Princeton University, United States of America</aff><author-notes>
                <corresp id="cor1">* E-mail: <email xlink:type="simple">ovitek@stat.purdue.edu</email></corresp>
            <fn fn-type="conflict">
                <p>The author has declared that no competing interests exist.</p>
            </fn></author-notes><pub-date pub-type="collection">
                <month>5</month>
                <year>2009</year>
            </pub-date><pub-date pub-type="epub">
                <day>29</day>
                <month>5</month>
                <year>2009</year>
            </pub-date><volume>5</volume><issue>5</issue><elocation-id>e1000366</elocation-id><!--===== Grouping copyright info into permissions =====--><permissions><copyright-year>2009</copyright-year><copyright-holder>Vitek</copyright-holder><license><license-p>This is an open-access article distributed under the terms
                of the Creative Commons Attribution License, which permits unrestricted use,
                distribution, and reproduction in any medium, provided the original author and
                source are credited.</license-p></license></permissions><funding-group><funding-statement>The author received no specific funding for this article.</funding-statement></funding-group><counts>
                <page-count count="4"/>
            </counts></article-meta>
    </front>
    <body>
        <p><graphic mimetype="image" position="anchor" xlink:href="info:doi/10.1371/journal.pcbi.1000366.ISCB_logo" xlink:type="simple"/></p>
            
            
            
        <sec id="s1">
            <title>Introduction</title>
            <p><italic>Proteomics</italic> aims at a large-scale characterization of localization,
                abundance, post-translational modifications, and biomolecular interactions of the
                proteins in an organism, with the goal of understanding their function. An extensive
                insight can be obtained by identifying and quantifying the components of biological
                mixtures. For example, a) In studies of biomolecular networks, partners interacting
                with a protein can help determine its function. It is possible to experimentally
                isolate protein complexes, e.g., using tag affinity purification. Identification of
                the components of this mixture helps determine potential interactors <xref ref-type="bibr" rid="pcbi.1000366-Kumar1">[1]</xref>. b)
                Post-translational modifications such as phosphorylation play an important role in
                regulating biological processes, e.g., cellular growth and signaling. Identification
                and quantification of phosphorylated proteins and their substrates helps elucidate
                complex signaling pathway phosphorylation events <xref ref-type="bibr" rid="pcbi.1000366-Mann1">[2]</xref>. c) Molecular biomarkers,
                i.e., proteins for which changes in abundance are indicative of an early onset of a
                disease or a therapy response, are of interest in clinical research. Identifying and
                quantifying components of a biofluid such as serum helps detect proteins with such
                discriminative ability <xref ref-type="bibr" rid="pcbi.1000366-Rifai1">[3]</xref>. d) A goal of genome annotation is the discovery
                and validation of protein-coding regions. Identifying peptides and proteins in a
                cell helps confirm and improve the annotations at the translational level, e.g., by
                confirming the presence of intron boundaries or alternative splicings <xref ref-type="bibr" rid="pcbi.1000366-Ansong1">[4]</xref>.</p>
            <p><italic>Mass spectrometry</italic> is a method of choice for protein identification
                and quantification due to its sensitivity and to the versatility of the
                instrumentation <xref ref-type="bibr" rid="pcbi.1000366-Aebersold1">[5]</xref>,<xref ref-type="bibr" rid="pcbi.1000366-Steen1">[6]</xref>. A typical “bottom-up”
                workflow experimentally digests the proteins into a mixture of peptides with an
                enzyme such as trypsin. This is necessary, in part, because the sensitivity of the
                mass spectrometer is much higher for peptides than for proteins. The peptides are
                then injected onto a liquid chromatography (LC) column from which they elute
                sequentially. The eluted peptides are ionized and separated by the mass spectrometer
                according to their ratio of mass to charge (<italic>m/z</italic>) in a mass spectrum
                (MS).</p>
            <p>The collection of mass spectra obtained at different elution times forms an LC-MS run
                shown in <xref ref-type="fig" rid="pcbi-1000366-g001">Figure 1A</xref>. Peaks in the
                run correspond to peptide ions; however, the sequence of amino acids underlying each
                peak is unknown. For identification, the mass spectrometer isolates the biological
                material from a peak (called precursor ion in this context), and subjects it to a
                high-collision energy. The energy breaks the peptide at different amide bonds, and
                the resulting fragments are separated according to their <italic>m/z</italic> in a
                secondary spectrum (called MS2, MS/MS, or tandem MS), shown in <xref ref-type="fig" rid="pcbi-1000366-g001">Figure 1B</xref>. Distances between peaks in the MS/MS
                spectrum are used to infer the peptide sequence of the parent LC-MS peak.</p>
            <fig id="pcbi-1000366-g001" position="float">
                <object-id pub-id-type="doi">10.1371/journal.pcbi.1000366.g001</object-id>
                <label>Figure 1</label>
                <caption>
                    <title>Example of spectral data.</title>
                    <p>(A) LC-MS run. Features in the LC-MS space are peptide ions; their intensity
                        is related to peptide abundance. (B) MS/MS spectrum. The spectrum is
                        obtained by fragmenting the peptide ion isolated from an LC-MS peak. The
                        peaks are fragment ions; distances between peaks are used for peptide
                        sequence determination.</p>
                </caption>
                <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pcbi.1000366.g001" xlink:type="simple"/>
            </fig>
            <p>Peak intensity is related to the abundances of peptides, and can be used for relative
                quantification. With the label-free approach, a separate LC-MS run is obtained for
                each biological sample, and peaks are quantified and compared across runs. In stable
                isotopic labeling workflow, samples from different groups are labeled metabolically
                (e.g., in SILAC, where stable isotopes are included in the growth medium of an
                organism), or chemically (e.g., in ICAT or iTRAQ, where reacting chemical labels are
                applied after tryptic digestion). Several samples (e.g., one from each group) are
                then mixed, and their peaks are identified and quantified within the same run.
                Finally, a targeted workflow based, for example, on selected reaction monitoring
                (SRM) <xref ref-type="bibr" rid="pcbi.1000366-Pan1">[7]</xref>,
                increases sensitivity and specificity by monitoring signals from a list of
                predefined peptides.</p>
            <p>The design of proteomic experiments, and subsequent analysis of the spectra, involves
                extensive computation and requires expertise at the intersection of computer
                science, engineering, and statistics. It presents exciting opportunities for both
                methodological and applied computational research.</p>
        </sec>
        <sec id="s2">
            <title>Experimental Design</title>
            <p>Experimental design specifies how biological samples are selected and allocated in
                space and time during spectral acquisition. For example, a biomarker discovery
                project can produce biased conclusions if patients from different groups have
                different characteristics (such as prior medication), or their spectra are acquired
                under different conditions. Moreover, sample selection and allocation can be
                inefficient, and can undermine the ability to uncover the true differences between
                groups.</p>
            <p>Statistical experimental design avoids bias and optimizes efficiency by using
                replication, randomization, and blocking, and by choosing an appropriate type and
                number of replicates <xref ref-type="bibr" rid="pcbi.1000366-Oberg1">[8]</xref>. The need for a statistical design of proteomic
                experiments is increasingly emphasized <xref ref-type="bibr" rid="pcbi.1000366-Ransohoff1">[9]</xref>. Specific choices
                require a statistical model that describes the spectra, and development of such
                models is an important area of research.</p>
        </sec>
        <sec id="s3">
            <title>Open Data Formats</title>
            <p>After spectral acquisition, the first computational task is to extract and store peak
                information. Unfortunately, most mass spectrometer vendors have their own
                proprietary formats. An advance has been made by implementing open XML-based formats
                (such as mzXML), and the associated converters and validators, to store this
                information and to make the subsequent analysis vendor-neutral <xref ref-type="bibr" rid="pcbi.1000366-Deutsch1">[10]</xref>. These tools are
                available from <ext-link ext-link-type="uri" xlink:href="http://www.proteomecommons.org" xlink:type="simple">http://www.proteomecommons.org</ext-link>. Efforts are invested, for example, by
                the Proteomics Standards Initiative (<ext-link ext-link-type="uri" xlink:href="http://www.psidev.info" xlink:type="simple">http://www.psidev.info</ext-link>), in
                further developments of XML formats.</p>
        </sec>
        <sec id="s4">
            <title>Identification of Peptides and Proteins</title>
            <p>An MS/MS spectrum such as in <xref ref-type="fig" rid="pcbi-1000366-g001">Figure
                1A</xref> is generated by a series of peptide fragments. Thus, mass differences
                between neighboring MS/MS peaks are used to determine the underlying amino acid
                sequence. Typical approaches involve searches of an a priori–defined
                database, de novo identifications, and combinations of the two <xref ref-type="bibr" rid="pcbi.1000366-Nesvizhskii1">[11]</xref>. Here we focus on
                the database-based approach which compares each observed spectrum against entries in
                a database (<xref ref-type="fig" rid="pcbi-1000366-g002">Figure 2A</xref>). Several
                aspects of the procedure require consideration.</p>
            <fig id="pcbi-1000366-g002" position="float">
                <object-id pub-id-type="doi">10.1371/journal.pcbi.1000366.g002</object-id>
                <label>Figure 2</label>
                <caption>
                    <title>Example of a proteomic workflow using database-based identification and
                        label-free quantification.</title>
                    <p>(A) Identification of MS/MS spectra. Experimental spectra are compared to
                        peptides in a database, and the best-scoring PSMs are reported while
                        controlling the FDR. Protein sequences are identified from the peptides. (B)
                        Label-free quantification. Features in LC-MS runs (shown with circles) are
                        located, quantified, and aligned across runs. (C) LC-MS features are
                        annotated with peptide sequences when identifications are available (shown
                        with filled circles). The annotations are used to optimize the alignment of
                        features across runs. The list of quantified, identified, and aligned
                        features is then subjected to transformation, normalization, and
                        summarization. (D) The list of features is used as input to machine
                        learning, functional annotation, and data integration steps.</p>
                </caption>
                <graphic mimetype="image" position="float" xlink:href="info:doi/10.1371/journal.pcbi.1000366.g002" xlink:type="simple"/>
            </fig>
            <sec id="s4a">
                <title>Database of Candidate Peptides</title>
                <p>Protein sequence databases now exist for many organisms. One can digest the
                    sequences in silico into peptides, and construct a theoretical spectrum for each
                    peptide. Alternatively, one can use a library of peptides with associated
                    consensus experimental spectra derived from previous identifications <xref ref-type="bibr" rid="pcbi.1000366-Lam1">[12]</xref>. In
                    both cases, the number of candidate peptides increases exponentially when we
                    allow nonspecific enzymes and/or post-translational modifications (PTM) that
                    alter a theoretical mass.</p>
            </sec>
            <sec id="s4b">
                <title>Scoring Function</title>
                <p>Scoring functions quantify the similarity of a candidate peptide-spectrum match
                    (PSM). A typical two-stage procedure filters out PSMs with incompatible peptide
                    and precursor ion masses, and scores plausible PSMs using counts of shared MS/MS
                    peaks. Newer scores incorporate additional characteristics, e.g., peak intensity
                    (for spectral libraries) and empirical peptide detectability <xref ref-type="bibr" rid="pcbi.1000366-Tang1">[13]</xref>, and
                    learn the scores dynamically from the data <xref ref-type="bibr" rid="pcbi.1000366-Kll1">[14]</xref>,<xref ref-type="bibr" rid="pcbi.1000366-Ding1">[15]</xref>.</p>
            </sec>
            <sec id="s4c">
                <title>Search Algorithm</title>
                <p>For each observed spectrum, the algorithm scores its similarity to every
                    candidate peptide and returns the best-scoring PSM. Since typical experiments
                    produce hundreds of thousands of MS/MS spectra, development of efficient search
                    algorithms is an active area of research. Improvements include clustering the
                    observed spectra using a similarity metric, and only searching the resulting
                    consensus spectra <xref ref-type="bibr" rid="pcbi.1000366-Frank1">[16]</xref>. Another approach aligns the observed spectra
                    in a procedure similar to genomic sequence alignment, and creates meta-spectra
                    that cover longer protein segments <xref ref-type="bibr" rid="pcbi.1000366-Bandeira1">[17]</xref>. Finally, a de
                    novo identification of short sequence tags (e.g., three amino acids long)
                    combined with a subsequent database search also allows one to reduce the space
                        <xref ref-type="bibr" rid="pcbi.1000366-Kim1">[18]</xref>.</p>
            </sec>
            <sec id="s4d">
                <title>False Discovery Rate (FDR) of Spectral Identification</title>
                <p>Due to the stochastic variation in the spectra, deficiencies of the scoring
                    schemes, and possible incompleteness of the database, only a fraction of
                    best-scoring PSMs are typically correct. There is thus a need for a statistical
                    measure of “confidence” in a reported list of PSMs, and for
                    an inferential procedure that distinguishes “confident” PSMs
                    from noise.</p>
                <p>An accepted statistical measure is FDR, defined as the expected proportion of
                    incorrect identifications in a list of PSMs with scores above a cutoff. To
                    determine FDR-controlled lists of PSMs, the target–decoy strategy
                        <xref ref-type="bibr" rid="pcbi.1000366-Elias1">[19]</xref> appends a randomized version of the theoretical
                    database (decoy) to the actual database (target), and estimates the FDR as twice
                    the proportion of decoy matches among all the matches in the list.
                    Alternatively, Peptide Prophet <xref ref-type="bibr" rid="pcbi.1000366-Keller1">[20]</xref> fits an Empirical Bayes two-group mixture
                    model to scores of correct and incorrect identifications, and estimates the FDR
                    as the fitted probability of correct identifications for scores above a cutoff.
                    Numerous extensions are continually proposed (see, e.g., <ext-link ext-link-type="uri" xlink:href="http://pubs.acs.org/toc/jprobs/7/1" xlink:type="simple">http://pubs.acs.org/toc/jprobs/7/1</ext-link>), and in the future the focus
                    will likely broaden to the FDR of peptides, proteins, and protein sites.</p>
            </sec>
            <sec id="s4e">
                <title>Protein Inference</title>
                <p>Confidently identified peptides can be grouped to infer the protein components of
                    the mixture. This is nontrivial due to ambiguous mappings of peptides to
                    proteins, and to the insufficient discrimination of some proteins by the
                    identified peptides <xref ref-type="bibr" rid="pcbi.1000366-Nesvizhskii2">[21]</xref>. Current approaches use characteristics such
                    as the number of mapped peptides, protein length, and peptide detectability
                        <xref ref-type="bibr" rid="pcbi.1000366-Alves1">[22]</xref> to identify proteins. More research is needed to
                    control the FDR in the protein list.</p>
            </sec>
            <sec id="s4f">
                <title>Resources</title>
                <p>Extensive spectral databases are publicly available, e.g., the Peptide Atlas at
                        <ext-link ext-link-type="uri" xlink:href="http://www.peptideatlas.org" xlink:type="simple">http://www.peptideatlas.org</ext-link>, containing millions of spectra from
                    biological experiments, and <ext-link ext-link-type="uri" xlink:href="http://regis-web.systemsbiology.net/PublicDatasets" xlink:type="simple">http://regis-web.systemsbiology.net/PublicDatasets</ext-link>, containing
                    spectra from controlled protein mixtures.</p>
            </sec>
        </sec>
        <sec id="s5">
            <title>Quantification</title>
            <p>Quantitative proteomics monitors peptide and protein abundance across samples of
                multiple types. The goals are similar to other high-throughput experiments such as
                gene expression microarrays <xref ref-type="bibr" rid="pcbi.1000366-Simon1">[23]</xref>,<xref ref-type="bibr" rid="pcbi.1000366-Gillette1">[24]</xref>. A typical workflow
                    (<xref ref-type="fig" rid="pcbi-1000366-g002">Figure 2</xref>) involves multiple
                steps <xref ref-type="bibr" rid="pcbi.1000366-Listgarten1">[25]</xref>.</p>
            <sec id="s5a">
                <title>Signal Processing</title>
                <p>Quantitative workflows require signal processing beyond spectral identification.
                    Features in the spectra must be located and quantified, annotated when possible
                    with peptide sequences information, and aligned across runs. A variety of tools
                    have been implemented <xref ref-type="bibr" rid="pcbi.1000366-Mueller1">[26]</xref>; they are specific to label-free or labeling
                    workflows, but all output a list of detected features and their abundances
                    across samples.</p>
            </sec>
            <sec id="s5b">
                <title>Transformation, Normalization, and Summarization</title>
                <p>The biological effects are multiplicative in nature, and a logarithm transform of
                    intensities is frequently recommended. Feature intensities are further
                    normalized across runs, e.g., using quantile normalization <xref ref-type="bibr" rid="pcbi.1000366-Bolstad1">[27]</xref>. When multiple
                    features are observed within a sample for a same peptide or protein, they are
                    often summarized in one number.</p>
            </sec>
            <sec id="s5c">
                <title>Learning</title>
                <p>Statistical and machine learning tools are then applied for (1) <italic>class
                        comparison</italic>, e.g., determination of proteins that change in
                    abundance between healthy individuals and individuals with disease; (2)
                        <italic>class discovery</italic>, e.g., unsupervised detection of sample
                    subgroups with homogeneous quantitative protein profiles; and (3) <italic>class
                        prediction</italic>, e.g., a supervised prediction of a disease status of a
                    new sample based on its protein abundance. Here analysis issues are similar to,
                    e.g., gene expression microarrays, in that the features are interdependent, and
                    their number exceeds the number of samples. An example from this area of
                    research is <xref ref-type="bibr" rid="pcbi.1000366-Hand1">[28]</xref>.</p>
            </sec>
            <sec id="s5d">
                <title>Functional Annotation</title>
                <p>Database technologies connect the proteins to their annotations, e.g., from Gene
                    Ontology, or from databases of disease. The annotations can confirm the
                    plausibility of the identifications, and can enable tests for over-represented
                    functional categories in the protein list <xref ref-type="bibr" rid="pcbi.1000366-Nam1">[29]</xref>.</p>
            </sec>
            <sec id="s5e">
                <title>Data Integration</title>
                <p>Recent studies combine proteomic measurements with gene expression and
                    metabolomic profiles, and/or known biochemical networks, with the general goal
                    of protein function determination <xref ref-type="bibr" rid="pcbi.1000366-Sharan1">[30]</xref>. A number of tools
                    facilitate these tasks, which include proprietary databases GeneGo and
                    Ingenuity, and open-source Cytoscape at <ext-link ext-link-type="uri" xlink:href="http://www.cytoscape.org" xlink:type="simple">http://www.cytoscape.org</ext-link>
                    and Bioconductor at <ext-link ext-link-type="uri" xlink:href="http://www.bioconductor.org" xlink:type="simple">http://www.bioconductor.org</ext-link>.</p>
            </sec>
        </sec>
    </body>
    <back>
        <ref-list>
            <title>References</title>
            <ref id="pcbi.1000366-Kumar1">
                <label>1</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Kumar</surname>
                            <given-names>A</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Snyder</surname>
                            <given-names>M</given-names>
                        </name>
                    </person-group>
                    <year>2002</year>
                    <article-title>Protein complexes take the bait.</article-title>
                    <source>Nature</source>
                    <volume>415</volume>
                    <fpage>123</fpage>
                    <lpage>124</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Mann1">
                <label>2</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Mann</surname>
                            <given-names>M</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Jensen</surname>
                            <given-names>O</given-names>
                        </name>
                    </person-group>
                    <year>2003</year>
                    <article-title>Proteomic analysis of post-translational modifications.</article-title>
                    <source>Nat Biotechnol</source>
                    <volume>21</volume>
                    <fpage>255</fpage>
                    <lpage>261</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Rifai1">
                <label>3</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Rifai</surname>
                            <given-names>N</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Gillette</surname>
                            <given-names>MA</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Carr</surname>
                            <given-names>SA</given-names>
                        </name>
                    </person-group>
                    <year>2006</year>
                    <article-title>Protein biomarker discovery and validation: The long and
                        uncertain path to clinical utility.</article-title>
                    <source>Nat Biotechnol</source>
                    <volume>24</volume>
                    <fpage>971</fpage>
                    <lpage>983</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Ansong1">
                <label>4</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Ansong</surname>
                            <given-names>C</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Purvine</surname>
                            <given-names>SO</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Adkins</surname>
                            <given-names>JN</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Lipton</surname>
                            <given-names>MS</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Smith</surname>
                            <given-names>RD</given-names>
                        </name>
                    </person-group>
                    <year>2008</year>
                    <article-title>Proteogenomics: Needs and roles to be filled by proteomics in
                        genome annotation.</article-title>
                    <source>Brief Funct Genomics Proteomics</source>
                    <volume>7</volume>
                    <fpage>50</fpage>
                    <lpage>62</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Aebersold1">
                <label>5</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Aebersold</surname>
                            <given-names>R</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Mann</surname>
                            <given-names>M</given-names>
                        </name>
                    </person-group>
                    <year>2003</year>
                    <article-title>Mass spectrometry–based proteomics.</article-title>
                    <source>Nature</source>
                    <volume>422</volume>
                    <fpage>198</fpage>
                    <lpage>207</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Steen1">
                <label>6</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Steen</surname>
                            <given-names>H</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Mann</surname>
                            <given-names>M</given-names>
                        </name>
                    </person-group>
                    <year>2004</year>
                    <article-title>The ABCs (and XYZs) of peptide sequencing.</article-title>
                    <source>Nat Rev</source>
                    <volume>5</volume>
                    <fpage>699</fpage>
                    <lpage>711</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Pan1">
                <label>7</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Pan</surname>
                            <given-names>S</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Aebersold</surname>
                            <given-names>R</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Chen</surname>
                            <given-names>R</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Rush</surname>
                            <given-names>J</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Goodlett</surname>
                            <given-names>DR</given-names>
                        </name>
                        <etal/>
                    </person-group>
                    <year>2009</year>
                    <article-title>Mass spectrometry based targeted protein quantification: Methods
                        and applications.</article-title>
                    <source>J Proteome Res</source>
                    <volume>8</volume>
                    <fpage>787</fpage>
                    <lpage>797</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Oberg1">
                <label>8</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Oberg</surname>
                            <given-names>AL</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Vitek</surname>
                            <given-names>O</given-names>
                        </name>
                    </person-group>
                    <year>2009</year>
                    <article-title>Statistical design of quantitative mass
                        spectrometry–based proteomic experiments.</article-title>
                    <source>J Proteome Res</source>
                    <comment>In press</comment>
                    <comment>doi:10.1021/pr8010099</comment>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Ransohoff1">
                <label>9</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Ransohoff</surname>
                            <given-names>DF</given-names>
                        </name>
                    </person-group>
                    <year>2005</year>
                    <article-title>Lessons from controversy: Ovarian cancer screening and serum
                        proteomics.</article-title>
                    <source>J Natl Cancer Inst</source>
                    <volume>97</volume>
                    <fpage>315</fpage>
                    <lpage>319</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Deutsch1">
                <label>10</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Deutsch</surname>
                            <given-names>EW</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Lam</surname>
                            <given-names>H</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Aebersold</surname>
                            <given-names>R</given-names>
                        </name>
                    </person-group>
                    <year>2008</year>
                    <article-title>Data analysis and bioinformatics tools for tandem mass
                        spectrometry in proteomics.</article-title>
                    <source>Physiol Genomics</source>
                    <volume>33</volume>
                    <fpage>18</fpage>
                    <lpage>25</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Nesvizhskii1">
                <label>11</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Nesvizhskii</surname>
                            <given-names>A</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Vitek</surname>
                            <given-names>O</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Aebersold</surname>
                            <given-names>R</given-names>
                        </name>
                    </person-group>
                    <year>2007</year>
                    <article-title>Analysis and validation of proteomic data generated by tandem
                        mass spectrometry.</article-title>
                    <source>Nat Methods</source>
                    <volume>4</volume>
                    <fpage>787</fpage>
                    <lpage>797</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Lam1">
                <label>12</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Lam</surname>
                            <given-names>H</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Deutsch</surname>
                            <given-names>EW</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Eddes</surname>
                            <given-names>JS</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Eng</surname>
                            <given-names>JK</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Stein</surname>
                            <given-names>SE</given-names>
                        </name>
                        <etal/>
                    </person-group>
                    <year>2008</year>
                    <article-title>Building consensus spectral libraries for peptide identification
                        in proteomics.</article-title>
                    <source>Nat Methods</source>
                    <volume>5</volume>
                    <fpage>873</fpage>
                    <lpage>875</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Tang1">
                <label>13</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Tang</surname>
                            <given-names>H</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Arnold</surname>
                            <given-names>RJ</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Alves</surname>
                            <given-names>P</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Xun</surname>
                            <given-names>Z</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Clemmer</surname>
                            <given-names>DE</given-names>
                        </name>
                        <etal/>
                    </person-group>
                    <year>2006</year>
                    <article-title>A computational approach toward label-free protein quantification
                        using predicted peptide detectability.</article-title>
                    <source>Bioinformatics</source>
                    <volume>22</volume>
                    <fpage>e481</fpage>
                    <lpage>e488</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Kll1">
                <label>14</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Käll</surname>
                            <given-names>L</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Canterbury</surname>
                            <given-names>JD</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Weston</surname>
                            <given-names>J</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Noble</surname>
                            <given-names>WS</given-names>
                        </name>
                        <name name-style="western">
                            <surname>MacCoss</surname>
                            <given-names>MJ</given-names>
                        </name>
                    </person-group>
                    <year>2007</year>
                    <article-title>Semi-supervised learning for peptide identification from shotgun
                        proteomics datasets.</article-title>
                    <source>Nat Methods</source>
                    <volume>4</volume>
                    <fpage>923</fpage>
                    <lpage>925</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Ding1">
                <label>15</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Ding</surname>
                            <given-names>Y</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Choi</surname>
                            <given-names>H</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Nesvizhskii</surname>
                            <given-names>AI</given-names>
                        </name>
                    </person-group>
                    <year>2008</year>
                    <article-title>Adaptive discriminant function analysis and reranking of MS/MS
                        database search results for improved peptide identification in shotgun
                        proteomics.</article-title>
                    <source>J Proteome Res</source>
                    <volume>7</volume>
                    <fpage>4878</fpage>
                    <lpage>4889</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Frank1">
                <label>16</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Frank</surname>
                            <given-names>AM</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Bandeira</surname>
                            <given-names>N</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Shen</surname>
                            <given-names>Z</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Tanner</surname>
                            <given-names>S</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Briggs</surname>
                            <given-names>SP</given-names>
                        </name>
                        <etal/>
                    </person-group>
                    <year>2008</year>
                    <article-title>Clustering millions of tandem mass spectra.</article-title>
                    <source>J Proteome Res</source>
                    <volume>7</volume>
                    <fpage>113</fpage>
                    <lpage>122</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Bandeira1">
                <label>17</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Bandeira</surname>
                            <given-names>N</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Tsur</surname>
                            <given-names>D</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Frank</surname>
                            <given-names>A</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Pevzner</surname>
                            <given-names>PA</given-names>
                        </name>
                    </person-group>
                    <year>2007</year>
                    <article-title>Protein identification by spectral networks analysis.</article-title>
                    <source>Proc Natl Acad Sci U S A</source>
                    <volume>104</volume>
                    <fpage>6140</fpage>
                    <lpage>6145</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Kim1">
                <label>18</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Kim</surname>
                            <given-names>S</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Gupta</surname>
                            <given-names>N</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Bandeira</surname>
                            <given-names>N</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Pevzner</surname>
                            <given-names>PA</given-names>
                        </name>
                    </person-group>
                    <year>2009</year>
                    <article-title>Spectral dictionaries: Integrating de novo peptide sequencing
                        with database search of tandem mass spectra.</article-title>
                    <source>Mol Cell Proteomics</source>
                    <volume>8</volume>
                    <fpage>53</fpage>
                    <lpage>69</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Elias1">
                <label>19</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Elias</surname>
                            <given-names>JE</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Gygi</surname>
                            <given-names>SP</given-names>
                        </name>
                    </person-group>
                    <year>2007</year>
                    <article-title>Target-decoy search strategy for increased confidence in
                        large-scale protein identifications by mass spectrometry.</article-title>
                    <source>Nat Methods</source>
                    <volume>2</volume>
                    <fpage>207</fpage>
                    <lpage>214</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Keller1">
                <label>20</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Keller</surname>
                            <given-names>A</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Nesvizhskii</surname>
                            <given-names>AI</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Kolker</surname>
                            <given-names>E</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Aebersold</surname>
                            <given-names>R</given-names>
                        </name>
                    </person-group>
                    <year>2002</year>
                    <article-title>Empirical statistical model to estimate the accuracy of peptide
                        identifications made by MS/MS and database search.</article-title>
                    <source>Analytical Chemistry</source>
                    <volume>74</volume>
                    <fpage>5383</fpage>
                    <lpage>5392</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Nesvizhskii2">
                <label>21</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Nesvizhskii</surname>
                            <given-names>AI</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Aebersold</surname>
                            <given-names>R</given-names>
                        </name>
                    </person-group>
                    <year>2005</year>
                    <article-title>Interpretation of shotgun proteomic data: The protein inference
                        problem.</article-title>
                    <source>Mol Cell Proteomics</source>
                    <volume>4</volume>
                    <fpage>1419</fpage>
                    <lpage>1440</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Alves1">
                <label>22</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Alves</surname>
                            <given-names>P</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Arnold</surname>
                            <given-names>RJ</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Novotny</surname>
                            <given-names>MV</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Radivojac</surname>
                            <given-names>P</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Reilly</surname>
                            <given-names>JP</given-names>
                        </name>
                        <etal/>
                    </person-group>
                    <year>2007</year>
                    <article-title>Advancements in protein inference from shotgun proteomics using
                        peptide detectability.</article-title>
                    <source>Pac Symp Biocomput</source>
                    <volume>12</volume>
                    <fpage>409</fpage>
                    <lpage>420</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Simon1">
                <label>23</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Simon</surname>
                            <given-names>R</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Radmacher</surname>
                            <given-names>MD</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Dobbins</surname>
                            <given-names>K</given-names>
                        </name>
                    </person-group>
                    <year>2002</year>
                    <article-title>Design of studies using DNA microarrays.</article-title>
                    <source>Genet Epidemiol</source>
                    <volume>23</volume>
                    <fpage>21</fpage>
                    <lpage>36</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Gillette1">
                <label>24</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Gillette</surname>
                            <given-names>MA</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Mani</surname>
                            <given-names>DR</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Carr</surname>
                            <given-names>SA</given-names>
                        </name>
                    </person-group>
                    <year>2005</year>
                    <article-title>Place of pattern in proteomic biomarker discovery.</article-title>
                    <source>J Proteome Res</source>
                    <volume>4</volume>
                    <fpage>1143</fpage>
                    <lpage>1154</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Listgarten1">
                <label>25</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Listgarten</surname>
                            <given-names>J</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Emili</surname>
                            <given-names>A</given-names>
                        </name>
                    </person-group>
                    <year>2005</year>
                    <article-title>Statistical and computational methods for comparative proteomic
                        profiling using liquid chromatography–tandem mass spectrometry.</article-title>
                    <source>Mol Cell Proteomics</source>
                    <volume>4</volume>
                    <fpage>419</fpage>
                    <lpage>434</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Mueller1">
                <label>26</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Mueller</surname>
                            <given-names>LN</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Brusniak</surname>
                            <given-names>MY</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Mani</surname>
                            <given-names>DR</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Aebersold</surname>
                            <given-names>R</given-names>
                        </name>
                    </person-group>
                    <year>2008</year>
                    <article-title>An assessment of software solutions for the analysis of mass
                        spectrometry based quantitative proteomics data.</article-title>
                    <source>J Proteome Res</source>
                    <volume>7</volume>
                    <fpage>51</fpage>
                    <lpage>61</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Bolstad1">
                <label>27</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Bolstad</surname>
                            <given-names>BM</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Irizarry</surname>
                            <given-names>RA</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Astrand</surname>
                            <given-names>M</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Speed</surname>
                            <given-names>TP</given-names>
                        </name>
                    </person-group>
                    <year>2003</year>
                    <article-title>A comparison of normalization methods for high density
                        oligonucleotide array data based on variance and bias.</article-title>
                    <source>Bioinformatics</source>
                    <volume>19</volume>
                    <fpage>185</fpage>
                    <lpage>193</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Hand1">
                <label>28</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Hand</surname>
                            <given-names>DJ</given-names>
                        </name>
                    </person-group>
                    <year>2008</year>
                    <article-title>Breast cancer diagnosis from proteomic mass spectrometry data: A
                        comparative evaluation.</article-title>
                    <source>Stat Appl Genet Mol Biol</source>
                    <volume>7</volume>
                    <fpage>Article 15</fpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Nam1">
                <label>29</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Nam</surname>
                            <given-names>D</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Kim</surname>
                            <given-names>SY</given-names>
                        </name>
                    </person-group>
                    <year>2008</year>
                    <article-title>Gene-set approach for expression pattern analysis.</article-title>
                    <source>Brief Bioinformatics</source>
                    <volume>9</volume>
                    <fpage>189</fpage>
                    <lpage>197</lpage>
                </element-citation>
            </ref>
            <ref id="pcbi.1000366-Sharan1">
                <label>30</label>
                <element-citation publication-type="journal" xlink:type="simple">
                    <person-group person-group-type="author">
                        <name name-style="western">
                            <surname>Sharan</surname>
                            <given-names>R</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Ulitsky</surname>
                            <given-names>I</given-names>
                        </name>
                        <name name-style="western">
                            <surname>Shamir</surname>
                            <given-names>R</given-names>
                        </name>
                    </person-group>
                    <year>2007</year>
                    <article-title>Network-based prediction of protein function.</article-title>
                    <source>Mol Syst Biol</source>
                    <volume>3</volume>
                    <fpage>Article 88</fpage>
                    <comment>doi:10.1038/msb4100129</comment>
                </element-citation>
            </ref>
        </ref-list>
        
    </back>
</article>