<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">PLoS ONE</journal-id>
<journal-id journal-id-type="publisher-id">plos</journal-id>
<journal-id journal-id-type="pmc">plosone</journal-id><journal-title-group>
<journal-title>PLoS ONE</journal-title></journal-title-group>
<issn pub-type="epub">1932-6203</issn>
<publisher>
<publisher-name>Public Library of Science</publisher-name>
<publisher-loc>San Francisco, USA</publisher-loc></publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">PONE-D-13-29779</article-id>
<article-id pub-id-type="doi">10.1371/journal.pone.0084217</article-id>
<article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer science</subject><subj-group><subject>Algorithms</subject></subj-group><subj-group><subject>Computer modeling</subject></subj-group></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Engineering</subject><subj-group><subject>Signal processing</subject><subj-group><subject>Data mining</subject><subject>Statistical signal processing</subject></subj-group></subj-group></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Mathematics</subject><subj-group><subject>Applied mathematics</subject><subj-group><subject>Algorithms</subject><subject>Decision theory</subject></subj-group></subj-group><subj-group><subject>Statistics</subject><subj-group><subject>Contingency tables</subject><subject>Decision theory</subject><subject>Statistical methods</subject></subj-group></subj-group></subj-group></article-categories>
<title-group>
<article-title>100% Classification Accuracy Considered Harmful: The Normalized Information Transfer Factor Explains the Accuracy Paradox</article-title>
<alt-title alt-title-type="running-head">100% Classication Accuracy Considered Harmful</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Valverde-Albacete</surname><given-names>Francisco J.</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Peláez-Moreno</surname><given-names>Carmen</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib>
</contrib-group>
<aff id="aff1"><label>1</label><addr-line>Departamento de Lenguajes y Sistemas Informáticos, Universidad Nacional de Educación a Distancia, Madrid, Spain</addr-line></aff>
<aff id="aff2"><label>2</label><addr-line>Signal Theory and Communications Department, University Carlos III Madrid, Madrid, Spain</addr-line></aff>
<contrib-group>
<contrib contrib-type="editor" xlink:type="simple"><name name-style="western"><surname>Paris</surname><given-names>Matteo G. A.</given-names></name>
<role>Editor</role>
<xref ref-type="aff" rid="edit1"/></contrib>
</contrib-group>
<aff id="edit1"><addr-line>Università degli Studi di Milano (University of Milan), Italy</addr-line></aff>
<author-notes>
<corresp id="cor1">* E-mail: <email xlink:type="simple">fva@lsi.uned.es</email></corresp>
<fn fn-type="conflict"><p>The authors have declared that no competing interests exist.</p></fn>
<fn fn-type="con"><p>Conceived and designed the experiments: FJVA CPM. Performed the experiments: FJVA CPM. Analyzed the data: FJVA CPM. Contributed reagents/materials/analysis tools: FJVA CPM. Wrote the paper: FJVA CPM.</p></fn>
</author-notes>
<pub-date pub-type="collection"><year>2014</year></pub-date>
<pub-date pub-type="epub"><day>10</day><month>1</month><year>2014</year></pub-date>
<volume>9</volume>
<issue>1</issue>
<elocation-id>e84217</elocation-id>
<history>
<date date-type="received"><day>22</day><month>7</month><year>2013</year></date>
<date date-type="accepted"><day>13</day><month>11</month><year>2013</year></date>
</history>
<permissions>
<copyright-year>2014</copyright-year>
<copyright-holder>Valverde-Albacete, Peláez-Moreno</copyright-holder><license xlink:href="http://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/" xlink:type="simple">Creative Commons Attribution License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</license-p></license></permissions>
<abstract>
<p>The most widely spread measure of performance, accuracy, suffers from a paradox: predictive models with a given level of accuracy may have greater predictive power than models with higher accuracy. Despite optimizing classification error rate, high accuracy models may fail to capture crucial information transfer in the classification task. We present evidence of this behavior by means of a combinatorial analysis where every possible contingency matrix of 2, 3 and 4 classes classifiers are depicted on the entropy triangle, a more reliable information-theoretic tool for classification assessment.</p>
<p>Motivated by this, we develop from first principles a measure of classification performance that takes into consideration the information learned by classifiers. We are then able to obtain the entropy-modulated accuracy (EMA), a pessimistic estimate of the expected accuracy with the influence of the input distribution factored out, and the normalized information transfer factor (NIT), a measure of how efficient is the transmission of information from the input to the output set of classes.</p>
<p>The EMA is a more natural measure of classification performance than accuracy when the heuristic to maximize is the transfer of information through the classifier instead of classification error count. The NIT factor measures the effectiveness of the learning process in classifiers and also makes it harder for them to “cheat” using techniques like specialization, while also promoting the interpretability of results. Their use is demonstrated in a mind reading task competition that aims at decoding the identity of a video stimulus based on magnetoencephalography recordings. We show how the EMA and the NIT factor reject rankings based in accuracy, choosing more meaningful and interpretable classifiers.</p>
</abstract>
<funding-group><funding-statement>Francisco José Valverde-Albacete has been partially supported by EU FP7 project LiMoSINe (contract 288024): <ext-link ext-link-type="uri" xlink:href="http://www.limosine-project.eu" xlink:type="simple">www.limosine-project.eu</ext-link> Carmen Peláez Moreno has been partially supported by the Spanish Government-Comisión Interministerial de Ciencia y Tecnología project TEC2011–26807. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</funding-statement></funding-group><counts><page-count count="10"/></counts></article-meta>
</front>
<body><sec id="s1">
<title>Introduction</title>
<p>Classification is an ubiquitous task in Science, Technology and the Humanities <xref ref-type="bibr" rid="pone.0084217-Sokal1">[1]</xref>. Usage ranges from diagnosing diseases <xref ref-type="bibr" rid="pone.0084217-Huang1">[2]</xref> or the status of tumors using gene expression data <xref ref-type="bibr" rid="pone.0084217-West1">[3]</xref> to the actual classification of tumor classes <xref ref-type="bibr" rid="pone.0084217-Wei1">[4]</xref>; from analyzing human performance in perceptual tasks <xref ref-type="bibr" rid="pone.0084217-Miller1">[5]</xref> to analyzing that of automated remote sensors <xref ref-type="bibr" rid="pone.0084217-Congalton1">[6]</xref> or automatic speech recognition machines <xref ref-type="bibr" rid="pone.0084217-Jurafsky1">[7]</xref>. If follows that the assessment of the performance of classification processes is of paramount importance for Scientific, Technological and Societal reasons <xref ref-type="bibr" rid="pone.0084217-Sokal1">[1]</xref>, <xref ref-type="bibr" rid="pone.0084217-Swets1">[8]</xref>–<xref ref-type="bibr" rid="pone.0084217-Jurman1">[10]</xref>.</p>
<p>To set the theoretical backdrop for our discussion, consider a set of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e001" xlink:type="simple"/></inline-formula> <italic>prior, instance or true classes</italic> <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e002" xlink:type="simple"/></inline-formula> and a discrete random variable <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e003" xlink:type="simple"/></inline-formula> distributed according to a <italic>prior class distribution</italic> <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e004" xlink:type="simple"/></inline-formula>. Consider also a set of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e005" xlink:type="simple"/></inline-formula> <italic>instances or patterns</italic>, each belonging to only one of those classes, but we do not know precisely which. A <italic>classification</italic> is a process whereby each of those instances is assigned to one among a set of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e006" xlink:type="simple"/></inline-formula> <italic>decision or predicted classes</italic> <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e007" xlink:type="simple"/></inline-formula> generating a discrete random variable <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e008" xlink:type="simple"/></inline-formula> distributed according to a <italic>posterior class distribution, </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e009" xlink:type="simple"/></inline-formula>, so that the joint events of this classification process consist of “presenting one instance of an input class <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e010" xlink:type="simple"/></inline-formula> for classification and deciding the output class to be <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e011" xlink:type="simple"/></inline-formula>”.</p>
<p>To measure the performance of the classification process we use its <italic>confusion matrix</italic>, a special contingency table <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e012" xlink:type="simple"/></inline-formula> counting the occurrences of the joint events. Usually, the maximum likelihood estimate of the joint probability <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e013" xlink:type="simple"/></inline-formula> is used as summary data. <xref ref-type="fig" rid="pone-0084217-g001">Figure 1</xref> represents two such contingency matrices for a <italic>brain decoding</italic> or <italic>mind reading</italic> task consisting in automatically identifying the class of video stimulus shown to the subjects based on magnetoencephalography (MEG) data. Five different types of stimuli were presented: the first three ones (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e014" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e015" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e016" xlink:type="simple"/></inline-formula>) belonging to the category of <italic>short</italic> clips (6–26 s. long) and the last two (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e017" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e018" xlink:type="simple"/></inline-formula>) to the category of <italic>long</italic> clips (approximately 10 min. long).</p>
<fig id="pone-0084217-g001" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0084217.g001</object-id><label>Figure 1</label><caption>
<title>Heatmap of the best classifiers of the MEG mind reading competition <xref ref-type="bibr" rid="pone.0084217-Klami1">[23]</xref> according to accuracy (left) and the EMA and the NIT factor (right) criteria.</title>
<p>Rows correspond to stimulus <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e019" xlink:type="simple"/></inline-formula> and columns to the decision <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e020" xlink:type="simple"/></inline-formula> or response. Darker hues correlate with higher joint probability <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e021" xlink:type="simple"/></inline-formula>. The heat map on the left reveals that the best classifier according to accuracy does not capture the fact that stimuli <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e022" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e023" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e024" xlink:type="simple"/></inline-formula> belong to a particular category whilst <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e025" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e026" xlink:type="simple"/></inline-formula> belong to another. A<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e027" xlink:type="simple"/></inline-formula> B<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e028" xlink:type="simple"/></inline-formula> C<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e029" xlink:type="simple"/></inline-formula></p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0084217.g001" position="float" xlink:type="simple"/></fig>
<p>Performance evaluation takes the form of the exploratory analysis of this confusion matrix or joint distribution. For instance, the <italic>de facto</italic> standard for performance <italic>visualization</italic> for binary—that is, two-class—classification is the Receiver-Operating-Characteristic (ROC) <xref ref-type="bibr" rid="pone.0084217-Fawcett1">[11]</xref>, but its generalization to higher class numbers is not as effective. We have argued elsewhere that the De Finetti entropy triangle (ET) <xref ref-type="bibr" rid="pone.0084217-ValverdeAlbacete1">[12]</xref> is a better tool to analyze classifier performance, with a solid information-theoretical basis, and not plagued with the problems of the ROC—see <italic>sec:mms: sec:entropy-triangle</italic>. In any case, neither device provides a <italic>single</italic> figure-of-merit or performance measure to compare systems, a practice cherished by researchers.</p>
<p>As a single figure-of-merit, by far the most widespread performance criterion used is <italic>accuracy</italic>, defined as the fraction of correctly classified instances, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e030" xlink:type="simple"/></inline-formula>. This is probably due to its easy and intuitive nature, despite many reasons <italic>not</italic> to do so <xref ref-type="bibr" rid="pone.0084217-BenDavid1">[13]</xref>. In <xref ref-type="bibr" rid="pone.0084217-Sokolova1">[14]</xref>, this and many other performance measures were examined in the context of several machine learning tasks, but inconclusive results as to their fitness of purpose were reached. However, the comparison made evident that accuracy was one of the measures that possessed the least number of invariants with respect to changes in confusion matrix entries, a detrimental quality. An earlier paper <xref ref-type="bibr" rid="pone.0084217-Kononenko1">[15]</xref> had already argued for the factoring <italic>out</italic> of the influence of prior class distributions on similar measures.</p>
<p>It is now acknowledged that <italic>high accuracy is not necessarily an indicator of high classifier performance</italic> and therein lies the <italic>accuracy paradox</italic> <xref ref-type="bibr" rid="pone.0084217-Zhu1">[16]</xref>–<xref ref-type="bibr" rid="pone.0084217-Fernandes1">[18]</xref>. For instance, in a predictive classification setting, predictive models with a given (lower) level of accuracy may have greater predictive power than models with higher accuracy. This deleterious feature is explained in-depth in Section <italic>sec:crit-accur-using</italic>. In particular, if a single class contains most of the data, a <italic>majority classifier</italic> that assigns all input cases to this majority class (the one concentrating the probability mass of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e031" xlink:type="simple"/></inline-formula>) would produce an accurate result. Highly <italic>imbalanced</italic> or <italic>skewed</italic> training data is very commonly encountered in samples taken from natural phenomena. Moreover, the classes' distributions of the samples do not necessarily reflect the distributions in the whole population since most of the times the samples are gathered in very controlled conditions. This skewness in the data hinders the capability of statistical models to predict the behavior of the phenomena being modeled and data balancing strategies are then advisable <xref ref-type="bibr" rid="pone.0084217-GarciaMoral1">[19]</xref>.</p>
<p>In this paper, we claim that performance measures based in the statistical information transfer from <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e032" xlink:type="simple"/></inline-formula> to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e033" xlink:type="simple"/></inline-formula> may be better measures for classification if <italic>predictive classification error is not the paramount performance criterion</italic>. This is the case, for example, of classifiers not used to make final decisions but, instead designed to be components of more complex diagnostic systems (as in <xref ref-type="bibr" rid="pone.0084217-GarciaMoral1">[19]</xref>) or when the conditions in the experimentation stage during which the data is collected do not hold in the deployment stage, as mentioned before. For this purpose, in Section <italic>sec:perpl-its-prop</italic> we establish the basis of our analysis in the propagation of <italic>perplexity</italic>—the effective number of classes a classifier sees—a concept that is directly related to accuracy.</p>
<p>In Section <italic>sec:perf-meas-based</italic> we use the <italic>remaining input perplexity </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e034" xlink:type="simple"/></inline-formula> to claim that the <italic>entropy-modulated accuracy (EMA)</italic>, defined in (3), is a better measure of classifier performance than accuracy for several reasons: it is well-grounded in information-theoretical terms, it provides an intuitive interpretation of the statistical learning process as the transfer of the information from the phenomena that are being modeled over a virtual channel, it factors out the influence of the input and output class distributions, it is invariant to permutations in the columns of the confusion matrix enabling the identification of cross-labeling errors common in unsupervised learning methods, and it is a pessimistic estimate of accuracy. For the same reasons, the <italic>normalized information transfer factor ( NIT factor )</italic>, defined as in (5), adds to some of the previous advantages the fact that it is capable of assessing the effectiveness of the learning process in the classifier, it is co-variant with expected mutual information (MI) <xref ref-type="bibr" rid="pone.0084217-Fano1">[20]</xref>, and contra-variant with the variation of information <xref ref-type="bibr" rid="pone.0084217-Meila1">[21]</xref>.</p>
<p>In <italic>sec:example-use</italic>, we suggest how to apply these metrics to a classification task, instantiating the process for a mind-reading challenge using multi-classification on magnetoencephalography signals, that shows one clear instance where ranking by EMA and NIT factor provides a more interpretable classifier than accuracy-based ranking. We provide further evidence, examples and a comparison with other metrics in <italic><xref ref-type="supplementary-material" rid="pone.0084217.s006">File S1</xref></italic>. The paper is closed with a <italic>sec:discussion</italic> where we also compare EMA and the NIT factor with two previously proposed measures for classification assessment and show the superiority of our proposal.</p>
</sec><sec id="s2">
<title>Results</title>
<sec id="s2a">
<title>A critique of accuracy using information-theoretic principles</title>
<p>To assess the theoretical adequacy of accuracy, we generated some samples of the space of joint count distributions for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e035" xlink:type="simple"/></inline-formula> input and output classes and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e036" xlink:type="simple"/></inline-formula> instances of classification with a prescribed accuracy (see Section <italic>sec:datasets</italic> for the details). Then, their entropy decomposition was calculated and plotted in the ET (see Section <italic>sec:mms, sec:entropy-triangle</italic>). <xref ref-type="fig" rid="pone-0084217-g002">Figure 2</xref> presents the cases <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e037" xlink:type="simple"/></inline-formula> with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e038" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e039" xlink:type="simple"/></inline-formula> with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e040" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e041" xlink:type="simple"/></inline-formula> with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e042" xlink:type="simple"/></inline-formula>.</p>
<fig id="pone-0084217-g002" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0084217.g002</object-id><label>Figure 2</label><caption>
<title>(Color online) Entropy decomposition for square matrices of (A) <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e043" xlink:type="simple"/></inline-formula>, (B) <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e044" xlink:type="simple"/></inline-formula>, and (C) <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e045" xlink:type="simple"/></inline-formula> (decimated), representing confusion matrices for a classification task at different accuracy levels as described by the right color bar.</title>
<p>The interspersing of the plots representing matrices with different accuracies but similar entropies is evident at all levels for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e046" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e047" xlink:type="simple"/></inline-formula> but only for lower levels of accuracy for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e048" xlink:type="simple"/></inline-formula>. This entails that accuracy is not a good criterion to judge the flow of information from the input labels to the output labels of a classifier (see text).</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0084217.g002" position="float" xlink:type="simple"/></fig>
<p>A number of observations can be gleaned from this figure:</p>
<list list-type="bullet"><list-item>
<p><italic>Matrices of a particular accuracy level are interspersed with those of many other accuracy levels</italic>. This phenomenon is the more prevalent the lower the accuracy level, although the behavior differs for different <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e049" xlink:type="simple"/></inline-formula>. For <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e050" xlink:type="simple"/></inline-formula> interspersing ends for accuracies over <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e051" xlink:type="simple"/></inline-formula> while for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e052" xlink:type="simple"/></inline-formula> it spreads to the whole range <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e053" xlink:type="simple"/></inline-formula>.</p>
</list-item><list-item>
<p><italic>For every prescribed accuracy level, the normalized mutual information ranges in </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e054" xlink:type="simple"/></inline-formula>, that is, there are matrices with accuracy over <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e055" xlink:type="simple"/></inline-formula> transmitting little or no information. This is the case even for high-accuracy matrices, including those with accuracy <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e056" xlink:type="simple"/></inline-formula>.</p>
</list-item><list-item>
<p>Conversely, matrices with different accuracy may exhibit the same normalized mutual information, for instance, check at <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e057" xlink:type="simple"/></inline-formula>.</p>
</list-item><list-item>
<p>There is an accumulation of distributions with high entropy (low <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e058" xlink:type="simple"/></inline-formula> values, left side of ET), as predicted by theory <xref ref-type="bibr" rid="pone.0084217-Jaynes1">[22]</xref>.</p>
</list-item></list>
<p>We are driven to conclude that accuracy is not a trustworthy criterion to judge the degree to which a particular classification process transfers information from the input class distribution to the output decision class distribution.</p>
</sec><sec id="s2b">
<title>Perplexity and its propagation in multiclass classifiers</title>
<p>The question poses itself whether it is possible to conjoin accuracy and mutual information transfer in a single measure. To provide an affirmative answer to this we first state the hypothesis:</p>
<sec id="s2b1">
<title>Hypothesis 1</title>
<p>In the absence of information about the items distributed according to a uniform prior class distribution, a classifier is expected to guess correctly <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e059" xlink:type="simple"/></inline-formula> of the times.</p>
<p>We will show that the EMA amounts to a ‘pessimistic’ accuracy estimate according to this hypothesis. For the sake of generality, suppose that the cardinality of the set of atomic events of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e060" xlink:type="simple"/></inline-formula> is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e061" xlink:type="simple"/></inline-formula> and that of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e062" xlink:type="simple"/></inline-formula> is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e063" xlink:type="simple"/></inline-formula>. Classification tasks with uniform input class distributions are often called <italic>balanced</italic> or <italic>unskewed</italic>. Let us denote this uniform input distribution as <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e064" xlink:type="simple"/></inline-formula> and accordingly, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e065" xlink:type="simple"/></inline-formula> will represent a uniform distribution of the outputs. Now <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e066" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e067" xlink:type="simple"/></inline-formula> represent the entropies of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e068" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e069" xlink:type="simple"/></inline-formula> respectively. Then <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e070" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e071" xlink:type="simple"/></inline-formula>, so <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e072" xlink:type="simple"/></inline-formula> is a measure of the <italic>theoretical perplexity</italic> of a classifier in a balanced task, that is, the number of <italic>possible</italic> events.</p>
<p>By analogy, call <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e073" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e074" xlink:type="simple"/></inline-formula> the <italic>perplexities</italic> of variables <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e075" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e076" xlink:type="simple"/></inline-formula> respectively. They are in fact an estimation of the <italic>effective</italic>—as opposed to the <italic>possible</italic>—number of atomic events behind <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e077" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e078" xlink:type="simple"/></inline-formula>. Note that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e079" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e080" xlink:type="simple"/></inline-formula> and that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e081" xlink:type="simple"/></inline-formula> (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e082" xlink:type="simple"/></inline-formula>) precisely when <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e083" xlink:type="simple"/></inline-formula> (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e084" xlink:type="simple"/></inline-formula>). Similarly, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e085" xlink:type="simple"/></inline-formula> (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e086" xlink:type="simple"/></inline-formula>) when <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e087" xlink:type="simple"/></inline-formula> (resp. <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e088" xlink:type="simple"/></inline-formula>) resembles a Kronecker delta function—that is, the input (and output) distribution is utterly skewed towards one class.</p>
<p>If we now define the quotient <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e089" xlink:type="simple"/></inline-formula> (respectively, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e090" xlink:type="simple"/></inline-formula>) we can see that<disp-formula id="pone.0084217.e091"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e091" xlink:type="simple"/></disp-formula><disp-formula id="pone.0084217.e092"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e092" xlink:type="simple"/></disp-formula>where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e093" xlink:type="simple"/></inline-formula> ( <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e094" xlink:type="simple"/></inline-formula>.) We interpret this quantity as the decrement (increment) in perplexity due to the choice of input (output) marginals of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e095" xlink:type="simple"/></inline-formula>.</p>
<p>The most important concept in our discussion is the <italic>information transfer factor</italic> <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e096" xlink:type="simple"/></inline-formula>: if we introduce two new <italic>remaining perplexities</italic>, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e097" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e098" xlink:type="simple"/></inline-formula>, from the well-known formulae <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e099" xlink:type="simple"/></inline-formula> this crucial quantity can be understood as the perplexity variation of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e100" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e101" xlink:type="simple"/></inline-formula> produced by the subtraction/addition of their mutual information,<disp-formula id="pone.0084217.e102"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e102" xlink:type="simple"/></disp-formula><disp-formula id="pone.0084217.e103"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e103" xlink:type="simple"/></disp-formula>hence the name.</p>
<p>It is easy to see that we have completed two different, sequentially related, decompositions of the perplexity of the variables,<disp-formula id="pone.0084217.e104"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e104" xlink:type="simple"/><label>(1)</label></disp-formula><disp-formula id="pone.0084217.e105"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e105" xlink:type="simple"/></disp-formula></p>
<p>This proves that an alternative way of conceptualizing the flow of information from one variable to the other is in terms of increments or decrements of their perplexity instead of the flows of entropies, as depicted in <xref ref-type="fig" rid="pone-0084217-g003">Fig. 3</xref>. In fact, the following inequalities can easily be checked,<disp-formula id="pone.0084217.e106"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e106" xlink:type="simple"/><label>(2)</label></disp-formula></p>
<fig id="pone-0084217-g003" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0084217.g003</object-id><label>Figure 3</label><caption>
<title>(Color online) Entropy (above) and perplexity (below) decomposition chains for a joint distribution.</title>
<p>Left, perplexity reduction in the input (learning) chain; right, perplexity increase in the output chain, related to classifier specialization. The colors refer to those of Fig. 5.(B). The ordering of the boxes is a convention to reveal the prior and posterior natures of the perplexities of class distributions. </p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0084217.g003" position="float" xlink:type="simple"/></fig>
<p>Note that analogue decompositions for marginal entropies were introduced in <xref ref-type="bibr" rid="pone.0084217-ValverdeAlbacete1">[12]</xref>, and are here collected as <italic>sec:mms: sec:split-entr-triangle</italic>. We will see next how this conceptualization allows us to devise an alternative to accuracy where the decomposition of <xref ref-type="disp-formula" rid="pone.0084217.e104">equation (1</xref>) underlines the preeminence of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e107" xlink:type="simple"/></inline-formula> for assessing performance.</p>
</sec></sec><sec id="s2c">
<title>Two performance measures based on perplexity</title>
<p>Consider a confusion matrix for a classifier obtained from <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e108" xlink:type="simple"/></inline-formula> instances of classification pairs. The lowest accuracy is that of a classifier returning a uniform count matrix: the most balanced testing dataset will distribute <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e109" xlink:type="simple"/></inline-formula> to each class and a clueless classifier will further redistribute these uniformly to each output class as <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e110" xlink:type="simple"/></inline-formula> instances. Since the diagonal has <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e111" xlink:type="simple"/></inline-formula> cells, the diagonal sum is<disp-formula id="pone.0084217.e112"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e112" xlink:type="simple"/></disp-formula></p>
<p>It is bounded by <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e113" xlink:type="simple"/></inline-formula> and any value smaller than the lower bound is an sure indication that a permutation of the output tags will ensure higher classification accuracy, that is, a better mapping of input to output <italic>class names</italic>.</p>
<p>Consider the perplexity reduction chain of <xref ref-type="fig" rid="pone-0084217-g003">Fig. 3</xref>. To the extent that the number of input classes and their distribution is a given—whereas <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e114" xlink:type="simple"/></inline-formula> is a construct of the classifier—we want to concentrate on measuring how well the input class distribution was learned by the training process, that is, in the prior class distribution perplexity reduction of <xref ref-type="disp-formula" rid="pone.0084217.e104">equation (1</xref>). Regarding the classifier training algorithm, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e115" xlink:type="simple"/></inline-formula> is a given and cannot be modified, whereas <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e116" xlink:type="simple"/></inline-formula> quantifies the amount of <italic>successfully</italic> learned information. More importantly for our purposes, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e117" xlink:type="simple"/></inline-formula> is the amount of information the classifier <italic>failed</italic> to learn. Therefore the EMA appears naturally as a quality measure based in the remaining perplexity of the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e118" xlink:type="simple"/></inline-formula> variable<disp-formula id="pone.0084217.e119"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e119" xlink:type="simple"/><label>(3)</label></disp-formula></p>
<p>Since <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e120" xlink:type="simple"/></inline-formula> is the entropy of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e121" xlink:type="simple"/></inline-formula> ignored by the classifier, as per our hypothesis and in the absence of any other source of information <italic>this is the expected performance of the classifier with equivalent (possibly fractional), equally likely </italic><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e122" xlink:type="simple"/></inline-formula><italic> classes</italic>: the higher this number, the worse the classifier will be.</p>
<p>To illustrate this, notice that when the training process of the classifier has been able to capitalize on all mutual information to leave no remaining perplexity, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e123" xlink:type="simple"/></inline-formula>, whence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e124" xlink:type="simple"/></inline-formula>. Similarly, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e125" xlink:type="simple"/></inline-formula>, either because the classifier has utterly failed to capture any information between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e126" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e127" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e128" xlink:type="simple"/></inline-formula>, or because the entropy of the data was minimal, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e129" xlink:type="simple"/></inline-formula>.</p>
<p>Notice that when the entropy of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e130" xlink:type="simple"/></inline-formula> is not maximal <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e131" xlink:type="simple"/></inline-formula> then <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e132" xlink:type="simple"/></inline-formula> whence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e133" xlink:type="simple"/></inline-formula> and the EMA detects an <italic>artificial</italic> lower bound for (2), <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e134" xlink:type="simple"/></inline-formula>. The artifice here is that this increase does not depend on the training of the classifier but on the prior class distribution. This suggests including a correction into <xref ref-type="disp-formula" rid="pone.0084217.e119">equation (3</xref>) to account for the deviation from uniformity in the prior class distribution <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e135" xlink:type="simple"/></inline-formula> so that<disp-formula id="pone.0084217.e136"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e136" xlink:type="simple"/><label>(4)</label></disp-formula>with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e137" xlink:type="simple"/></inline-formula> when both <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e138" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e139" xlink:type="simple"/></inline-formula>, implying that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e140" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e141" xlink:type="simple"/></inline-formula>. Note that <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e142" xlink:type="simple"/></inline-formula> if and only if <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e143" xlink:type="simple"/></inline-formula>. Unlike the case of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e144" xlink:type="simple"/></inline-formula>, the eventuality that the data are not uniformly distributed is corrected on <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e145" xlink:type="simple"/></inline-formula>, as <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e146" xlink:type="simple"/></inline-formula> entails <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e147" xlink:type="simple"/></inline-formula>. Moreover, the further away from a uniform prior class distribution to the classifier, the worse its upper range bound will be. Eventually, for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e148" xlink:type="simple"/></inline-formula>—which implies <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e149" xlink:type="simple"/></inline-formula> by <xref ref-type="disp-formula" rid="pone.0084217.e106">equation (2</xref>) whence <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e150" xlink:type="simple"/></inline-formula>—we have, again, the worst possible value of the measure, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e151" xlink:type="simple"/></inline-formula>. Notice that in this accuracy-optimal case <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e152" xlink:type="simple"/></inline-formula>, but in an unhelpful way. Essentially, making the input data less random impacts the ability of the classifier to capitalize in mutual information to bind together input and output, and this is registered by the measure. The <italic>normalized information transfer factor</italic> can be rewritten as,</p>
<p><disp-formula id="pone.0084217.e153"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e153" xlink:type="simple"/><label>(5)</label></disp-formula>Note also that NIT factor does <italic>not</italic> depend directly on the input or output distribution. Conveniently, since the normalized information transfer factor is a monotonic function of normalized mutual information the relative height in the ET offers a visual tool to quickly inspect such effectiveness. Finally, when evaluating a set of systems in the same task, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e154" xlink:type="simple"/></inline-formula> is constant throughout the evaluation, so <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e155" xlink:type="simple"/></inline-formula>, and they offer the same ranking results, easily visualized in the ET.</p>
<p>For the reasons above, we posit the EMA in (3) to measure the performance of classification tasks, and the NIT factor in (4) or (5) to measure the effectiveness of the classifier learning process.</p>
</sec><sec id="s2d">
<title>Assessing classifiers with EMA and the NIT factor</title>
<p>In this Section we present an example of how to use the EMA and the NIT factor in automatic classifier evaluation tasks. We consider the case of the MEG mind reading challenge organized by the PASCAL (Pattern Analysis, Statistical modeling and ComputAtional Learning) network <xref ref-type="bibr" rid="pone.0084217-Klami1">[23]</xref>. Since accuracy was the “official” evaluation criterion, for comparison purposes <xref ref-type="fig" rid="pone-0084217-g004">Fig. 4</xref>.fig: (A) presents the results in the entropy triangle ordered by accuracy as reflected in the coloring of the points. System <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e156" xlink:type="simple"/></inline-formula> at <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e157" xlink:type="simple"/></inline-formula> was deemed the winner with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e158" xlink:type="simple"/></inline-formula> close behind at <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e159" xlink:type="simple"/></inline-formula>. In a detail of the dense region of harder competition in <xref ref-type="fig" rid="pone-0084217-g004">Fig. 4</xref>.(B) clusters <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e160" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e161" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e162" xlink:type="simple"/></inline-formula> are evident. We next suggest a procedure to analyze the classification performance of a population of classifiers:</p>
<fig id="pone-0084217-g004" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0084217.g004</object-id><label>Figure 4</label><caption>
<title>(Color online) Entropy triangle for the MEG mind Reading data ordered after accuracy (A) and a detail of the participants of higher accuracy (B).</title>
<p>The ranking following accuracy is at odds with the EMA and the NIT factor ranking based in mutual information (height, right scale of triangle). The detail in (B) shows that participant <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e174" xlink:type="simple"/></inline-formula>, closely followed by <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e175" xlink:type="simple"/></inline-formula> should have been ranked first after this criterion. </p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0084217.g004" position="float" xlink:type="simple"/></fig>
<list list-type="order"><list-item>
<p><bold>Use </bold><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e163" xlink:type="simple"/></inline-formula><bold> to assess the effective number of classes of the data.</bold> At <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e164" xlink:type="simple"/></inline-formula> down from <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e165" xlink:type="simple"/></inline-formula>, the task is quite balanced, guaranteeing that systems will find it harder to specialize as majority classifiers.</p>
</list-item><list-item>
<p><bold>Use EMA to rank classifiers.</bold> <xref ref-type="table" rid="pone-0084217-t001">Table 1</xref> presents the perplexities, accuracies, the EMA and the NIT factor for the confusion matrices of the classifiers that took part in the task. Ranking <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e166" xlink:type="simple"/></inline-formula> suggests itself, aligned with increasing mutual information (right axis). Indeed, after EMA, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e167" xlink:type="simple"/></inline-formula> should have been the winner of the competition, followed closely by <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e168" xlink:type="simple"/></inline-formula>.</p>
</list-item><list-item>
<p><bold>Use the ET to individually assess each classifier.</bold> From the ET diagram it is evident that those classifiers with highest mutual information and accuracy—the first seven classifiers—are not specialized while classifier <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e169" xlink:type="simple"/></inline-formula>, and, to a lesser extent, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e170" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e171" xlink:type="simple"/></inline-formula> are. The worst classifier is barely above random at <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e172" xlink:type="simple"/></inline-formula>.</p>
</list-item><list-item>
<p><bold>Use the NIT factor to assess whether the population of classifiers has solved the task.</bold> Overall, for the top ranked classifier we have <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e173" xlink:type="simple"/></inline-formula>, showing that the task has indeed not been effectively solved by the participants, either individually or collectively.</p>
</list-item></list>
<table-wrap id="pone-0084217-t001" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0084217.t001</object-id><label>Table 1</label><caption>
<title>Perplexities, accuracy (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e176" xlink:type="simple"/></inline-formula>), <italic>EMA</italic> (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e177" xlink:type="simple"/></inline-formula>) and <italic>NIT factor</italic> (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e178" xlink:type="simple"/></inline-formula>) for MEG Mind Reading confusion matrices ranked by accuracy.</title>
</caption><alternatives><graphic id="pone-0084217-t001-1" position="float" mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0084217.t001" xlink:type="simple"/>
<table><colgroup span="1"><col align="left" span="1"/><col align="center" span="1"/><col align="center" span="1"/><col align="center" span="1"/><col align="center" span="1"/><col align="center" span="1"/></colgroup>
<thead>
<tr>
<td align="left" rowspan="1" colspan="1">Exp.</td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e179" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e180" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e181" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e182" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e183" xlink:type="simple"/></inline-formula></td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e184" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">2.562</td>
<td align="left" rowspan="1" colspan="1">1.932</td>
<td align="left" rowspan="1" colspan="1">0.680</td>
<td align="left" rowspan="1" colspan="1">0.390</td>
<td align="left" rowspan="1" colspan="1">0.386</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e185" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">2.447</td>
<td align="left" rowspan="1" colspan="1">2.023</td>
<td align="left" rowspan="1" colspan="1">0.632</td>
<td align="left" rowspan="1" colspan="1">0.409</td>
<td align="left" rowspan="1" colspan="1">0.405</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e186" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">2.589</td>
<td align="left" rowspan="1" colspan="1">1.912</td>
<td align="left" rowspan="1" colspan="1">0.628</td>
<td align="left" rowspan="1" colspan="1">0.386</td>
<td align="left" rowspan="1" colspan="1">0.382</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e187" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">2.430</td>
<td align="left" rowspan="1" colspan="1">2.037</td>
<td align="left" rowspan="1" colspan="1">0.622</td>
<td align="left" rowspan="1" colspan="1">0.412</td>
<td align="left" rowspan="1" colspan="1">0.407</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e188" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">2.723</td>
<td align="left" rowspan="1" colspan="1">1.818</td>
<td align="left" rowspan="1" colspan="1">0.565</td>
<td align="left" rowspan="1" colspan="1">0.367</td>
<td align="left" rowspan="1" colspan="1">0.364</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e189" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">2.682</td>
<td align="left" rowspan="1" colspan="1">1.846</td>
<td align="left" rowspan="1" colspan="1">0.542</td>
<td align="left" rowspan="1" colspan="1">0.373</td>
<td align="left" rowspan="1" colspan="1">0.369</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e190" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">2.730</td>
<td align="left" rowspan="1" colspan="1">1.813</td>
<td align="left" rowspan="1" colspan="1">0.539</td>
<td align="left" rowspan="1" colspan="1">0.366</td>
<td align="left" rowspan="1" colspan="1">0.363</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e191" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">3.629</td>
<td align="left" rowspan="1" colspan="1">1.364</td>
<td align="left" rowspan="1" colspan="1">0.472</td>
<td align="left" rowspan="1" colspan="1">0.276</td>
<td align="left" rowspan="1" colspan="1">0.273</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e192" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">2.995</td>
<td align="left" rowspan="1" colspan="1">1.653</td>
<td align="left" rowspan="1" colspan="1">0.443</td>
<td align="left" rowspan="1" colspan="1">0.334</td>
<td align="left" rowspan="1" colspan="1">0.331</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1"><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e193" xlink:type="simple"/></inline-formula></td>
<td align="left" rowspan="1" colspan="1">4.801</td>
<td align="left" rowspan="1" colspan="1">1.031</td>
<td align="left" rowspan="1" colspan="1">0.242</td>
<td align="left" rowspan="1" colspan="1">0.208</td>
<td align="left" rowspan="1" colspan="1">0.206</td>
</tr>
</tbody>
</table>
</alternatives><table-wrap-foot><fn id="nt101"><label/><p>Class <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e194" xlink:type="simple"/></inline-formula> should have been ranked above the rest by EMA or NIT factor (in all cases <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e195" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e196" xlink:type="simple"/></inline-formula>).</p></fn></table-wrap-foot></table-wrap>
<p>The result of this process is an assessment of a (population of) classifiers, whereby one may discuss the advantages of EMA and NIT factor vis–vis other performance measures, for instance, accuracy. Further examples of using this procedure to evaluate classification tasks can be found in the <italic><xref ref-type="supplementary-material" rid="pone.0084217.s006">File S1</xref></italic>.</p>
<sec id="s2d1">
<title>EMA and NIT factor vs. Accuracy</title>
<p>The authors of the report on the MEG Mind Reading challenge attempted an analysis of the ranking results and specifically compare classifier <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e197" xlink:type="simple"/></inline-formula> to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e198" xlink:type="simple"/></inline-formula> since the heat map of the latter seems to be “cleaner” <xref ref-type="bibr" rid="pone.0084217-Klami1">[23]</xref> (see <xref ref-type="fig" rid="pone-0084217-g004">Fig. 4</xref> with the heat map of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e199" xlink:type="simple"/></inline-formula> (left) to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e200" xlink:type="simple"/></inline-formula> (right)). For them, classifier <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e201" xlink:type="simple"/></inline-formula> essentially came out first because it used the “learning capacity” of its technique to improve classification error while <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e202" xlink:type="simple"/></inline-formula> used the capacity to better distinguish the two categories of classes present in the task (with stimuli <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e203" xlink:type="simple"/></inline-formula> to <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e204" xlink:type="simple"/></inline-formula> belonging to a first category whilst <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e205" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e206" xlink:type="simple"/></inline-formula>, to another) but was worse at capturing the distinctions among the classes of the first category.</p>
<p>Our rejection of this judgement comes from believing that the goal of recovering class structure is as worthy as minimizing classification errors. The interpretability of the results of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e207" xlink:type="simple"/></inline-formula> is superior to those of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e208" xlink:type="simple"/></inline-formula> since it has better captured the nature of the underlying phenomenon. This means that the errors committed by <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e209" xlink:type="simple"/></inline-formula> are likely to be inside the same category of the correct response (given the nearly block diagonal structure of its heat map) while in the case of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e210" xlink:type="simple"/></inline-formula>, for example, the probability of having stimuli of the first category erroneously predicted as <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e211" xlink:type="simple"/></inline-formula> is very high.</p>
<p>The EMA and the NIT factor prove apt at considering the value of representing the underlying structure with their tight relation to perplexity. In fact, according to <xref ref-type="bibr" rid="pone.0084217-Klami1">[23]</xref>, while <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e212" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e213" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e214" xlink:type="simple"/></inline-formula> focused on solving the so-called <italic>domain adaptation problem</italic>—the mismatch in training and testing conditions—with advanced machine learning techniques, many of the other teams, including <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e215" xlink:type="simple"/></inline-formula>, addressed it by placing more weight on the labeled <italic>test</italic> samples provided along with the <italic>train</italic> samples, when validating the learned classifier, thus <italic>explicitly</italic> boosting test set accuracy.</p>
</sec></sec></sec><sec id="s3">
<title>Discussion</title>
<sec id="s3a">
<title>Measure definition</title>
<p>Perplexity has already been used as a performance measurement for language modeling where it refers to the expected average of alternatives a model has at every word history <xref ref-type="bibr" rid="pone.0084217-Jelinek1">[24]</xref>. It is also often used as an off-line method for speech recognition task evaluation following the intuition that a classifier using a lower-perplexity model will outperform a higher-perplexity one, all other things equal.</p>
<p>It cannot be stressed enough that since the EMA and the NIT factor concentrate in the prior class distribution and mutual information, it is harder for classifiers to boost their performance by manipulating the posterior class distribution through specialization: only the increase in information transfer through <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e216" xlink:type="simple"/></inline-formula> will improve the evaluation figure.</p>
<p>Considering robustness, the EMA, being a harsher, worst-case criterion, might be more deserving of trust than easygoing and unreliable accuracy to, for instance, guide decision making. It certainly has a more interpretable and less easily bendable criterion—specially if <italic>reporting</italic> the classification error is not the ultimate goal. Furthermore, in cases where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e217" xlink:type="simple"/></inline-formula>—for instance, when using a “reject” class—the EMA and the NIT factor are still defined, whereas accuracy is problematic, and not very much used.</p>
</sec><sec id="s3b">
<title>Classification task assessment</title>
<p>As seen in the MEG mind reading example, the EMA and the NIT factor are capable of determining whether a task has been effectively solved or not. But it cannot distinguish whether this is caused by technical limitations in the classifier selection process or because the task is inherently “hard”. Only the kind of iterated classification effort of community research that attempts many different classifier-building techniques on the same task can be effective for this purpose.</p>
<p>Nevertheless, the effective input perplexity <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e218" xlink:type="simple"/></inline-formula> can ensure that, methodologically at least, the task is “as hard as it should be” at <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e219" xlink:type="simple"/></inline-formula>. Furthermore, our developments show clearly that a failure to maintain prior class distribution uniformity in the design or capture of the task data entails that the expected mutual information—therefore the NIT factor —captured by any possible classifier that solves the task can never reach maximal levels. This is a strong guideline for prospective collectors of datasets, although data balancing strategies after data collection can also be used to achieve this goal <xref ref-type="bibr" rid="pone.0084217-GarciaMoral1">[19]</xref>.</p>
</sec><sec id="s3c">
<title>Measure comparison</title>
<p>Several other measures have sprouted to deal with the inadequacies of accuracy such as the Area-Under-the-(ROC)-Curve <xref ref-type="bibr" rid="pone.0084217-Swets1">[8]</xref>, <xref ref-type="bibr" rid="pone.0084217-Bradley1">[25]</xref>, the Variation of Information <xref ref-type="bibr" rid="pone.0084217-Meila1">[21]</xref>, the Relative Classifier Information <xref ref-type="bibr" rid="pone.0084217-Sindhwani1">[26]</xref>, the Confusion Entropy <xref ref-type="bibr" rid="pone.0084217-Jurman1">[10]</xref>, <xref ref-type="bibr" rid="pone.0084217-Wei2">[27]</xref> or Cohen's Kappa <xref ref-type="bibr" rid="pone.0084217-BenDavid1">[13]</xref>, but their use is not widespread, specially for the non-binary case, due to complexity of calculation, disparate purposes or each measures' own shortcomings. For instance, the AUC first needs to find a (multiclass) ROC representation of the task by obtaining multiple classifiers, possibly with the help of a parameter in the classifier learning process. The trading for good-vs-wrong decisions in terms of the parameter can then be judged from the Area-Under-the-ROC curve, which is then a measure <italic>on the learning method or model</italic>. In contrast, EMA would provide a different point in the ET for each classifier whence the best of these classifiers could be chosen. Complementarily, on the <italic>population of classifiers</italic>, a statistical description of the NIT factor could be used to assess the learning capabilities of the method.</p>
<p>In classification proper, to illustrate the disparity of the conclusions that can be reached with alternative performance measures, we have included in <italic><xref ref-type="supplementary-material" rid="pone.0084217.s006">File S1</xref></italic> a comparison of the classical Matthew Correlation Coefficient (MCC) <xref ref-type="bibr" rid="pone.0084217-Matthews1">[28]</xref> and the Confusion Entropy (CEN) <xref ref-type="bibr" rid="pone.0084217-Wei2">[27]</xref>—whose similarities are also explored in <xref ref-type="bibr" rid="pone.0084217-Jurman1">[10]</xref>–on three different classifications tasks: the MEG Mind Reading task already explored, the TASS sentiment analysis task <xref ref-type="bibr" rid="pone.0084217-ValverdeAlbacete2">[29]</xref>–both machine learning tasks—and the well-known Miller &amp; Nicely human perceptual capability exploration task <xref ref-type="bibr" rid="pone.0084217-Miller1">[5]</xref>.</p>
<p>For each task we provide the heat maps of the confusion matrices (<xref ref-type="supplementary-material" rid="pone.0084217.s001">Figs. S1</xref>, <xref ref-type="supplementary-material" rid="pone.0084217.s002">S2</xref> and <xref ref-type="supplementary-material" rid="pone.0084217.s004">S4</xref> in <italic><xref ref-type="supplementary-material" rid="pone.0084217.s006">File S1</xref></italic>) as customary. We also provide the tables detailing perplexities, EMA, NIT factor, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e220" xlink:type="simple"/></inline-formula> and MCC' related values (Tables S1, S2 and S3 in <italic><xref ref-type="supplementary-material" rid="pone.0084217.s006">File S1</xref></italic> ). The entries in the tables are ordered by accuracy. For the TASS and M&amp;N data we also supply the ET's with the color bar according to EMA, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e221" xlink:type="simple"/></inline-formula> and MCC' (<xref ref-type="supplementary-material" rid="pone.0084217.s003">Figs. S3</xref> and <xref ref-type="supplementary-material" rid="pone.0084217.s005">S5</xref>). Their comparison, detailed in the <italic><xref ref-type="supplementary-material" rid="pone.0084217.s006">File S1</xref></italic> Section, reveals that MCC' is highly correlated with accuracy in ranking results and shows similar shortcomings. Even though CEN performs a little better, it is highly biased towards majority classifiers providing over optimistic assessment for them. Notably, once the ET, EMA and the NIT factor have shed light on the problem, reassessment of prior evidences for either CEN or MCC prove them not to be so advantageous in evaluating classifiers.</p>
</sec></sec><sec id="s4" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="s4a">
<title>The entropy triangle</title>
<p>Consider two discrete random variables <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e222" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e223" xlink:type="simple"/></inline-formula> and their joint probability distribution <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e224" xlink:type="simple"/></inline-formula>. An entropy diagram somewhat more complete than what is normally used for the relations between their entropies was presented in <xref ref-type="bibr" rid="pone.0084217-ValverdeAlbacete1">[12]</xref> and is here depicted in <xref ref-type="fig" rid="pone-0084217-g005">Fig. 5(A)</xref>. We distinguish in it the familiar decomposition of the joint entropy <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e225" xlink:type="simple"/></inline-formula> as the two entropies <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e226" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e227" xlink:type="simple"/></inline-formula> whose intersection is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e228" xlink:type="simple"/></inline-formula>. But notice that the increment between <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e229" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e230" xlink:type="simple"/></inline-formula> is yet again <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e231" xlink:type="simple"/></inline-formula>, hence the expected mutual information appears <italic>twice</italic> in the diagram. Further, the interior of the outer rectangle represents <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e232" xlink:type="simple"/></inline-formula>—with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e233" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e234" xlink:type="simple"/></inline-formula> the uniform distribution on inputs and outputs—,the interior of the inner rectangle <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e235" xlink:type="simple"/></inline-formula>, and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e236" xlink:type="simple"/></inline-formula> is their difference. Finally, the <italic>variation of information</italic> <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e237" xlink:type="simple"/></inline-formula> was found to be an important quantity in <xref ref-type="bibr" rid="pone.0084217-Meila1">[21]</xref>. Putting together this information results in the <italic>balance equation for information related to a joint distribution</italic>,<disp-formula id="pone.0084217.e238"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e238" xlink:type="simple"/></disp-formula>which can be further normalized in <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e239" xlink:type="simple"/></inline-formula>,<disp-formula id="pone.0084217.e240"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e240" xlink:type="simple"/><label>(6)</label></disp-formula>and represented in a De Finetti or ternary diagram as the equation of the <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e241" xlink:type="simple"/></inline-formula>-simplex in normalized <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e242" xlink:type="simple"/></inline-formula> space, hence the name entropy triangle, <italic>ET</italic>.</p>
<fig id="pone-0084217-g005" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0084217.g005</object-id><label>Figure 5</label><caption>
<title>(Color online) Extended information diagrams of entropies related to a bivariate distribution: (A) conventional diagram, and (B) split diagram.</title>
<p>The bounding rectangle is the joint entropy of two uniform (thence independent) distributions <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e243" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e244" xlink:type="simple"/></inline-formula> of the same cardinality as <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e245" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e246" xlink:type="simple"/></inline-formula>. The expected mutual information <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e247" xlink:type="simple"/></inline-formula> appears <italic>twice</italic> in (A) and this makes the diagram split for each variable symmetrically in (B).</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0084217.g005" position="float" xlink:type="simple"/></fig>
<p>The position of the coordinates of a classifier on the Entropy Triangle characterizes its performance, and we use this characterization to visually assess it indicated in <xref ref-type="fig" rid="pone-0084217-g006">Fig. 6</xref>. Classifiers at the apex or close to it obtain the highest accuracy possible on balanced datasets and transmit a lot of mutual information, hence they are the <italic>best classifiers</italic> possible. Those at the left vertex or close to it are dealing with balanced data but doing a bad job of utilizing it: they are the <italic>worst classifiers</italic>. Those at the right vertex or close to it are dealing with very easy, unbalanced data and claiming very high accuracy, yet they are not learning anything from it: they are <italic>specialized (majority) classifier</italic>s and our intuition is that they are the kind of classifiers that generate the accuracy paradox <xref ref-type="bibr" rid="pone.0084217-Zhu1">[16]</xref>.</p>
<fig id="pone-0084217-g006" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0084217.g006</object-id><label>Figure 6</label><caption>
<title>Schematic Entropy Triangle showing interpretable zones and extreme cases of classifiers.</title>
<p>The annotations on the center of each side are meant to hold for that whole side.</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0084217.g006" position="float" xlink:type="simple"/></fig></sec><sec id="s4b">
<title>The split entropy triangle</title>
<p>Notice that in <xref ref-type="disp-formula" rid="pone.0084217.e240">equation (6</xref>), since both <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e248" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e249" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e250" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e251" xlink:type="simple"/></inline-formula> are independent as marginals of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e252" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e253" xlink:type="simple"/></inline-formula>, respectively, we may write:<disp-formula id="pone.0084217.e254"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e254" xlink:type="simple"/></disp-formula><disp-formula id="pone.0084217.e255"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e255" xlink:type="simple"/></disp-formula>what suggests writing separate balance equations for each variable,</p>
<p><disp-formula id="pone.0084217.e256"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e256" xlink:type="simple"/></disp-formula><disp-formula id="pone.0084217.e257"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0084217.e257" xlink:type="simple"/></disp-formula></p>
<p>The formulae above and the occurrence of twice the expected mutual information in <xref ref-type="disp-formula" rid="pone.0084217.e240">equation (6</xref>) suggests a different information diagram, depicted in <xref ref-type="fig" rid="pone-0084217-g005">Fig. 5(b)</xref>: both variables <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e258" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e259" xlink:type="simple"/></inline-formula> now appear somehow decoupled—in the sense that the areas representing them are disjoint—yet there is a strong coupling in that the expected mutual information appears in both <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e260" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e261" xlink:type="simple"/></inline-formula>. It is important to note that both decompositions can be represented in the same (split) entropy triangle as <xref ref-type="disp-formula" rid="pone.0084217.e240">equation (6</xref>) dictates. The technique is explained in <xref ref-type="bibr" rid="pone.0084217-ValverdeAlbacete1">[12]</xref>.</p>
</sec><sec id="s4c">
<title>Data</title>
<p>The space of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e262" xlink:type="simple"/></inline-formula> square confusion matrices, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e263" xlink:type="simple"/></inline-formula> of sizes <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e264" xlink:type="simple"/></inline-formula> and a given number of input samples, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e265" xlink:type="simple"/></inline-formula>, depicted in <xref ref-type="fig" rid="pone-0084217-g002">Fig. 2</xref> was obtained by first generating every possible partition of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e266" xlink:type="simple"/></inline-formula> with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e267" xlink:type="simple"/></inline-formula> parts as input distributions <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e268" xlink:type="simple"/></inline-formula>, allocating <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e269" xlink:type="simple"/></inline-formula> input samples in each of the input classes. In this way, the set of all possible input class distributions, from uniform <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e270" xlink:type="simple"/></inline-formula> to the most skewed <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e271" xlink:type="simple"/></inline-formula>, is obtained. Then, for each of the previous distributions, every possible weak composition of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e272" xlink:type="simple"/></inline-formula> with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e273" xlink:type="simple"/></inline-formula> parts is produced, yielding <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e274" xlink:type="simple"/></inline-formula> sets of all the possible distributions for each of the rows of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e275" xlink:type="simple"/></inline-formula>. Finally, the Cartesian product of those sets produces every possible combination of rows corresponding to the selection of one element in every one of the sets. Except from row permutations —that would only amount to a reordering of the input classes— this procedure guarantees the presence of every possible <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e276" xlink:type="simple"/></inline-formula>.</p>
<p>The MEG mind reading task aims at decoding the identity of a video stimulus based on magnetoencephalography (MEG) recordings done during naturalistic stimulation <xref ref-type="bibr" rid="pone.0084217-Klami1">[23]</xref>. In particular, subjects were exposed to video stimuli of different classes: a first category of <italic>short</italic> clips (6–26 s. long) with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e277" xlink:type="simple"/></inline-formula> being <italic>artificial</italic> stimuli (screen savers showing animated shapes or text), <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e278" xlink:type="simple"/></inline-formula> being <italic>natural</italic> stimuli (sceneries like mountains or oceans) and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e279" xlink:type="simple"/></inline-formula> being <italic>football</italic> stimuli (from —European— football matches) and a second category of <italic>long</italic> clips (approximately 10 min. long) with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e280" xlink:type="simple"/></inline-formula> being television series (from “Mr. Bean” in particular) and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e281" xlink:type="simple"/></inline-formula> being films (from Chaplin's “Modern times”). The goal was to classify unlabeled test examples into these classes based on the MEG signal alone. The competition took place in March, 2011 and 10 participants submitted their classifiers whose confusion matrices are analyzed in this paper. The data was provided upon request from the organizers of the competition.</p>
<p>The MATLAB(A registered trademark of The MathWorks, Inc.) code to draw the entropy triangles in <xref ref-type="fig" rid="pone-0084217-g002">Figures 2</xref> and <xref ref-type="fig" rid="pone-0084217-g004">4</xref> has been made available at: <ext-link ext-link-type="uri" xlink:href="http://www.mathworks.com/matlabcentral/fileexchange/30914" xlink:type="simple">http://www.mathworks.com/matlabcentral/fileexchange/30914</ext-link></p>
</sec></sec><sec id="s5">
<title>Supporting Information</title>
<supplementary-material id="pone.0084217.s001" mimetype="image/tiff" xlink:href="info:doi/10.1371/journal.pone.0084217.s001" position="float" xlink:type="simple"><label>Figure S1</label><caption>
<p><bold>Heat maps of the classifiers of the MEG mind reading competition </bold><xref ref-type="bibr" rid="pone.0084217-Klami1">[<bold>23</bold>]</xref><bold>.</bold> Rows correspond to stimulus <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e282" xlink:type="simple"/></inline-formula> and columns to the decision <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e283" xlink:type="simple"/></inline-formula> or response. Darker hues correlate with higher joint probability <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e284" xlink:type="simple"/></inline-formula>. The classifier denominations obey to their position in the ranking produced by accuracy.</p>
<p>(TIFF)</p>
</caption></supplementary-material><supplementary-material id="pone.0084217.s002" mimetype="image/tiff" xlink:href="info:doi/10.1371/journal.pone.0084217.s002" position="float" xlink:type="simple"><label>Figure S2</label><caption>
<p><bold>Heat maps of the classifiers of the TASS competition </bold><xref ref-type="bibr" rid="pone.0084217-ValverdeAlbacete2">[<bold>29</bold>]</xref><bold>.</bold> Rows correspond to stimulus <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e285" xlink:type="simple"/></inline-formula> and columns to the decision <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e286" xlink:type="simple"/></inline-formula> or response. Darker hues correlate with higher joint probability <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e287" xlink:type="simple"/></inline-formula>. The classifier denominations obey to their position in the ranking produced by accuracy <bold>A</bold> Color bar represents EMA <bold>B</bold> Color bar represents <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e288" xlink:type="simple"/></inline-formula> <bold>C</bold> Color bar represents <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e289" xlink:type="simple"/></inline-formula>.</p>
<p>(TIFF)</p>
</caption></supplementary-material><supplementary-material id="pone.0084217.s003" mimetype="image/tiff" xlink:href="info:doi/10.1371/journal.pone.0084217.s003" position="float" xlink:type="simple"><label>Figure S3</label><caption>
<p>(Color online) <bold>Entropy decomposition for the classifiers of the TASS competition (A) with the color bar representing EMA, (B) </bold><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e290" xlink:type="simple"/></inline-formula><bold>, and (C) </bold><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e291" xlink:type="simple"/></inline-formula>.</p>
<p>(TIFF)</p>
</caption></supplementary-material><supplementary-material id="pone.0084217.s004" mimetype="image/tiff" xlink:href="info:doi/10.1371/journal.pone.0084217.s004" position="float" xlink:type="simple"><label>Figure S4</label><caption>
<p><bold>Heatmaps of the classifiers of the TASS competition </bold><xref ref-type="bibr" rid="pone.0084217-ValverdeAlbacete2">[<bold>29</bold>]</xref><bold>.</bold> Rows correspond to stimulus <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e292" xlink:type="simple"/></inline-formula> and columns to the decision <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e293" xlink:type="simple"/></inline-formula> or response. Darker hues correlate with higher joint probability <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e294" xlink:type="simple"/></inline-formula>. The classifier denominations obey to their position in the ranking produced by accuracy <bold>A</bold> Color bar represents EMA <bold>B</bold> Color bar represents <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e295" xlink:type="simple"/></inline-formula>] withFigures <bold>C</bold> Color bar represents <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e296" xlink:type="simple"/></inline-formula>.</p>
<p>(TIFF)</p>
</caption></supplementary-material><supplementary-material id="pone.0084217.s005" mimetype="image/tiff" xlink:href="info:doi/10.1371/journal.pone.0084217.s005" position="float" xlink:type="simple"><label>Figure S5</label><caption>
<p>(Color online) <bold>Entropy decomposition for MN phonetic confusion matrices (A) with the color bar representing EMA, (B) </bold><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e297" xlink:type="simple"/></inline-formula><bold>, and (C) </bold><inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0084217.e298" xlink:type="simple"/></inline-formula><bold>.</bold></p>
<p>(TIFF)</p>
</caption></supplementary-material><supplementary-material id="pone.0084217.s006" mimetype="application/pdf" xlink:href="info:doi/10.1371/journal.pone.0084217.s006" position="float" xlink:type="simple"><label>File S1</label><caption>
<p><bold>Supporting Information.</bold> A comparison of the classical Matthew Correlation Coefficient (MCC) <xref ref-type="bibr" rid="pone.0084217-Matthews1">[28]</xref> and the Confusion Entropy (CEN) <xref ref-type="bibr" rid="pone.0084217-Wei2">[27]</xref>—whose similarities are also explored in <xref ref-type="bibr" rid="pone.0084217-Jurman1">[10]</xref>–on three different classifications tasks: the MEG Mind Reading task already explored, the TASS sentiment analysis task <xref ref-type="bibr" rid="pone.0084217-ValverdeAlbacete2">[29]</xref>–both machine learning tasks—and the well-known Miller &amp; Nicely human perceptual capability exploration task <xref ref-type="bibr" rid="pone.0084217-Miller1">[5]</xref>.</p>
<p>(PDF)</p>
</caption></supplementary-material></sec></body>
<back>
<ack>
<p>The authors would like to thank A. Klami for providing the MEG Mind Reading data, J. Villena for the TASS data, and both A. Sánchez, C. Bousoño and the anonymous reviewers for comments on previous versions of this paper.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="pone.0084217-Sokal1"><label>1</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sokal</surname><given-names>RR</given-names></name> (<year>1974</year>) <article-title>Classification: Purposes, principles, progress, prospects</article-title>. <source>Science</source> <volume>185</volume>: <fpage>1115</fpage>–<lpage>1123</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Huang1"><label>2</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Huang</surname><given-names>H</given-names></name>, <name name-style="western"><surname>Liu</surname><given-names>CC</given-names></name>, <name name-style="western"><surname>Zhou</surname><given-names>XJ</given-names></name> (<year>2010</year>) <article-title>Bayesian approach to transforming public gene expression repositories into disease diagnosis databases</article-title>. <source>Proceedings of the National Academy of Sciences of the United States of America</source> <volume>107</volume>: <fpage>6823</fpage>–<lpage>6828</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-West1"><label>3</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>West</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Blanchette</surname><given-names>C</given-names></name>, <name name-style="western"><surname>Dressman</surname><given-names>H</given-names></name>, <name name-style="western"><surname>Huang</surname><given-names>E</given-names></name>, <name name-style="western"><surname>Ishida</surname><given-names>S</given-names></name>, <etal>et al</etal>. (<year>2001</year>) <article-title>Predicting the clinical status of human breast cancer by using gene expression profiles</article-title>. <source>Proceedings of the National Academy of Sciences of the United States of America</source> <volume>98</volume>: <fpage>11462</fpage>–<lpage>11467</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Wei1"><label>4</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Wei</surname><given-names>X</given-names></name>, <name name-style="western"><surname>Li</surname><given-names>KC</given-names></name> (<year>2010</year>) <article-title>Exploring the within- and between-class correlation distributions for tumor classification</article-title>. <source>Proceedings of the National Academy of Sciences of the United States of America</source> <volume>107</volume>: <fpage>6737</fpage>–<lpage>6742</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Miller1"><label>5</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Miller</surname><given-names>GA</given-names></name>, <name name-style="western"><surname>Nicely</surname><given-names>PE</given-names></name> (<year>1955</year>) <article-title>An analysis of perceptual confusions among some English consonants</article-title>. <source>Journal of the Acoustical Society of America</source> <volume>27</volume>: <fpage>338</fpage>–<lpage>352</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Congalton1"><label>6</label>
<mixed-citation publication-type="other" xlink:type="simple">Congalton RG, Green K (1999) Assessing the Accuracy of Remotely Sensed Data: Principles and Practices. CRC Press, Inc.</mixed-citation>
</ref>
<ref id="pone.0084217-Jurafsky1"><label>7</label>
<mixed-citation publication-type="other" xlink:type="simple">Jurafsky D, Martin JH (2000) Speech and Language Processing. Prentice-Hall.</mixed-citation>
</ref>
<ref id="pone.0084217-Swets1"><label>8</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Swets</surname><given-names>JA</given-names></name> (<year>1988</year>) <article-title>Measuring the accuracy of diagnostic systems</article-title>. <source>Science</source> <volume>240</volume>: <fpage>1285</fpage>–<lpage>1293</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Powers1"><label>9</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Powers</surname><given-names>D</given-names></name> (<year>2011</year>) <article-title>Evaluation: From Precision, Recall and F-Measure to ROC., Informedness, Markedness &amp; Correlation</article-title>. <source>Journal of Machine Learning Technologies</source> <volume>2</volume>: <fpage>37</fpage>–<lpage>63</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Jurman1"><label>10</label>
<mixed-citation publication-type="other" xlink:type="simple">Jurman G, Riccadonna S, Furlanello C (2012) A comparison of MCC and CEN error measures in multi-class prediction. PLoS ONE 7.</mixed-citation>
</ref>
<ref id="pone.0084217-Fawcett1"><label>11</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Fawcett</surname><given-names>T</given-names></name> (<year>2006</year>) <article-title>An introduction to ROC analysis</article-title>. <source>Pattern Recognition Letters</source> <volume>27</volume>: <fpage>861</fpage>–<lpage>874</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-ValverdeAlbacete1"><label>12</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Valverde-Albacete</surname><given-names>FJ</given-names></name>, <name name-style="western"><surname>Peláez-Moreno</surname><given-names>C</given-names></name> (<year>2010</year>) <article-title>Two information-theoretic tools to assess the performance of multi-class classifiers</article-title>. <source>Pattern Recognition Letters</source> <volume>31</volume>: <fpage>1665</fpage>–<lpage>1671</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-BenDavid1"><label>13</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Ben-David</surname><given-names>A</given-names></name> (<year>2007</year>) <article-title>A lot of randomness is hiding in accuracy</article-title>. <source>Engineering Applications of Artificial Intelligence</source> <volume>20</volume>: <fpage>875</fpage>–<lpage>885</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Sokolova1"><label>14</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sokolova</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Lapalme</surname><given-names>G</given-names></name> (<year>2009</year>) <article-title>A systematic analysis of performance measures for classification tasks</article-title>. <source>Information Processing &amp; Management</source> <volume>45</volume>: <fpage>427</fpage>–<lpage>437</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Kononenko1"><label>15</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Kononenko</surname><given-names>I</given-names></name>, <name name-style="western"><surname>Bratko</surname><given-names>I</given-names></name> (<year>1991</year>) <article-title>Information-based evaluation criterion for classifier's performance</article-title>. <source>Machine Learning</source> <volume>6</volume>: <fpage>67</fpage>–<lpage>80</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Zhu1"><label>16</label>
<mixed-citation publication-type="other" xlink:type="simple">Zhu X, Davidson I (2007) Knowledge discovery and data mining: challenges and realities. Premier reference source. Information Science Reference.</mixed-citation>
</ref>
<ref id="pone.0084217-Thomas1"><label>17</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Thomas</surname><given-names>C</given-names></name>, <name name-style="western"><surname>Balakrishnan</surname><given-names>N</given-names></name> (<year>2008</year>) <article-title>Improvement in minority attack detection with skewness in network traffic</article-title>. <source>Proc SPIE Int Soc Opt Eng</source> <volume>6973</volume>: <fpage>69730N</fpage>–<lpage>69730N-12</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Fernandes1"><label>18</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Fernandes</surname><given-names>JA</given-names></name>, <name name-style="western"><surname>Irigoien</surname><given-names>X</given-names></name>, <name name-style="western"><surname>Goikoetxea</surname><given-names>N</given-names></name>, <name name-style="western"><surname>Lozano</surname><given-names>JA</given-names></name>, <name name-style="western"><surname>naki Inza</surname><given-names>I</given-names></name>, <etal>et al</etal>. (<year>2010</year>) <article-title>Fish recruitment prediction, using robust supervised classification methods</article-title>. <source>Ecological Modelling</source> <volume>221</volume>: <fpage>338</fpage>–<lpage>352</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-GarciaMoral1"><label>19</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Garcia-Moral</surname><given-names>A</given-names></name>, <name name-style="western"><surname>Solera-Urena</surname><given-names>R</given-names></name>, <name name-style="western"><surname>Pelaez-Moreno</surname><given-names>C</given-names></name>, <name name-style="western"><surname>Diaz-de Maria</surname><given-names>F</given-names></name> (<year>2011</year>) <article-title>Data balancing for efficient training of hybrid ann/hmm automatic speech recognition systems</article-title>. <source>Audio, Speech, and Language Processing, IEEE Transactions on</source> <volume>19</volume>: <fpage>468</fpage>–<lpage>481</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Fano1"><label>20</label>
<mixed-citation publication-type="other" xlink:type="simple">Fano RM (1961) Transmission of Information: A Statistical Theory of Communication. The MIT Press, 400 pp.</mixed-citation>
</ref>
<ref id="pone.0084217-Meila1"><label>21</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Meila</surname><given-names>M</given-names></name> (<year>2007</year>) <article-title>Comparing clusterings|an information based distance</article-title>. <source>Journal of Multivariate Analysis</source> <volume>28</volume>: <fpage>875</fpage>–<lpage>893</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Jaynes1"><label>22</label>
<mixed-citation publication-type="other" xlink:type="simple">Jaynes E (1983) Concentration of distributions at entropy maxima. In: Rosenkrantz R, editor, Papers on Probability, Statistics, and Statistical Physics, D. Reidel Publishing.</mixed-citation>
</ref>
<ref id="pone.0084217-Klami1"><label>23</label>
<mixed-citation publication-type="other" xlink:type="simple">Klami A, Ramkumar P, Virtanen S, Parkkonen L, Hari R, <etal>et al</etal>.. (2011) ICANN/PASCAL2 challenge: MEG mind reading – overview and results. In: Klami A, editor, Proceedings of ICANN/PASCAL2 Challenge: MEG Mind Reading. Espoo, Aalto University Publication series SCIENCE + TECHNOLOGY 29/2011, pp. 3–19.</mixed-citation>
</ref>
<ref id="pone.0084217-Jelinek1"><label>24</label>
<mixed-citation publication-type="other" xlink:type="simple">Jelinek F (1997) Statistical Methods for Speech Recognition. Cambridge, Ma; London, UK: The MIT Press.</mixed-citation>
</ref>
<ref id="pone.0084217-Bradley1"><label>25</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Bradley</surname><given-names>AP</given-names></name> (<year>1997</year>) <article-title>The use of the area under the ROC curve in the evaluation of machine learning algorithms</article-title>. <source>Pattern Recognition</source> <volume>30</volume>: <fpage>1145</fpage>–<lpage>1159</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Sindhwani1"><label>26</label>
<mixed-citation publication-type="other" xlink:type="simple">Sindhwani V, Bhattacharya P, Rakshit S (2001) Information theoretic feature crediting in multiclass support vector machines. In: Proceedings of the First SIAM International Conference on Data Mining. pp.5–7.</mixed-citation>
</ref>
<ref id="pone.0084217-Wei2"><label>27</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Wei</surname><given-names>JM</given-names></name>, <name name-style="western"><surname>Yuan</surname><given-names>XJ</given-names></name>, <name name-style="western"><surname>Hu</surname><given-names>QH</given-names></name>, <name name-style="western"><surname>Wang</surname><given-names>SQ</given-names></name> (<year>2010</year>) <article-title>A novel measure for evaluating classifiers</article-title>. <source>Expert Systems with Applications</source> <volume>37</volume>: <fpage>3799</fpage>–<lpage>3809</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-Matthews1"><label>28</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Matthews</surname><given-names>B</given-names></name> (<year>1975</year>) <article-title>Comparison of the predicted and observed secondary structure of {T4} phage lysozyme</article-title>. <source>Biochimica et Biophysica Acta (BBA) - Protein Structure</source> <volume>405</volume>: <fpage>442</fpage>–<lpage>451</lpage>.</mixed-citation>
</ref>
<ref id="pone.0084217-ValverdeAlbacete2"><label>29</label>
<mixed-citation publication-type="other" xlink:type="simple">Valverde-Albacete FJ, Carrillo-de Albornoz J, Peláez-Moreno C (2013) A proposal for new evaluation metrics and result visualization technique for sentiment analysis tasks. In: Information Access Evaluation. Multilinguality, Multimodality, and Visualization, Springer Berlin Heidelberg, volume 8138 of <italic>Lecture Notes in Computer Science</italic>. pp. 41–52.</mixed-citation>
</ref>
</ref-list></back>
</article>