<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">PLoS ONE</journal-id>
<journal-id journal-id-type="publisher-id">plos</journal-id>
<journal-id journal-id-type="pmc">plosone</journal-id>
<journal-title-group>
<journal-title>PLOS ONE</journal-title>
</journal-title-group>
<issn pub-type="epub">1932-6203</issn>
<publisher>
<publisher-name>Public Library of Science</publisher-name>
<publisher-loc>San Francisco, CA USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">PONE-D-15-02411</article-id>
<article-id pub-id-type="doi">10.1371/journal.pone.0139475</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Research Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Zipf’s Law: Balancing Signal Usage Cost and Communication Efficiency</article-title>
<alt-title alt-title-type="running-head">Zipf’s Law: Balancing Signal Usage Cost and Communication Efficiency</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" xlink:type="simple">
<name name-style="western">
<surname>Salge</surname> <given-names>Christoph</given-names></name>
<xref ref-type="aff" rid="aff001"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple">
<name name-style="western">
<surname>Ay</surname> <given-names>Nihat</given-names></name>
<xref ref-type="aff" rid="aff002"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff003"><sup>3</sup></xref>
<xref ref-type="aff" rid="aff004"><sup>4</sup></xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple">
<name name-style="western">
<surname>Polani</surname> <given-names>Daniel</given-names></name>
<xref ref-type="aff" rid="aff001"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff005"><sup>5</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes" xlink:type="simple">
<name name-style="western">
<surname>Prokopenko</surname> <given-names>Mikhail</given-names></name>
<xref ref-type="aff" rid="aff005"><sup>5</sup></xref>
<xref ref-type="aff" rid="aff006"><sup>6</sup></xref>
<xref ref-type="corresp" rid="cor001">*</xref>
</contrib>
</contrib-group>
<aff id="aff001">
<label>1</label>
<addr-line>Department of Computer Science, University of Hertfordshire, Hatfield, United Kingdom</addr-line>
</aff>
<aff id="aff002">
<label>2</label>
<addr-line>Max Planck Institute for Mathematics in the Sciences, Leipzig, Germany</addr-line>
</aff>
<aff id="aff003">
<label>3</label>
<addr-line>Santa Fe Institute, Santa Fe, United States of America</addr-line>
</aff>
<aff id="aff004">
<label>4</label>
<addr-line>Department of Mathematics and Computer Science, Leipzig University, Leipzig, Germany</addr-line>
</aff>
<aff id="aff005">
<label>5</label>
<addr-line>Complex Systems Research Group, Faculty of Engineering and IT, The University of Sydney, Sydney, Australia</addr-line>
</aff>
<aff id="aff006">
<label>6</label>
<addr-line>Department of Computing, Macquarie University, Sydney, Australia</addr-line>
</aff>
<contrib-group>
<contrib contrib-type="editor" xlink:type="simple">
<name name-style="western">
<surname>Smalheiser</surname> <given-names>Neil R.</given-names></name>
<role>Editor</role>
<xref ref-type="aff" rid="edit1"/>
</contrib>
</contrib-group>
<aff id="edit1">
<addr-line>University of Illinois-Chicago, UNITED STATES</addr-line>
</aff>
<author-notes>
<fn fn-type="conflict" id="coi001">
<p>The authors have declared that no competing interests exist.</p>
</fn>
<fn fn-type="con" id="contrib001">
<p>Conceived and designed the experiments: CS NA DP MP. Performed the experiments: CS. Analyzed the data: CS NA DP MP. Contributed reagents/materials/analysis tools: NA MP. Wrote the paper: CS NA DP MP.</p>
</fn>
<corresp id="cor001">* E-mail: <email xlink:type="simple">mikhail.prokopenko@sydney.edu.au</email></corresp>
</author-notes>
<pub-date pub-type="collection">
<year>2015</year>
</pub-date>
<pub-date pub-type="epub">
<day>1</day>
<month>10</month>
<year>2015</year>
</pub-date>
<volume>10</volume>
<issue>10</issue>
<elocation-id>e0139475</elocation-id>
<history>
<date date-type="received">
<day>18</day>
<month>1</month>
<year>2015</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>9</month>
<year>2015</year>
</date>
</history>
<permissions>
<copyright-year>2015</copyright-year>
<copyright-holder>Salge et al</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/" xlink:type="simple">
<license-p>This is an open access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/" xlink:type="simple">Creative Commons Attribution License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="info:doi/10.1371/journal.pone.0139475" xlink:type="simple"/>
<abstract>
<p>We propose a model that explains the reliable emergence of power laws (e.g., Zipf’s law) during the development of different human languages. The model incorporates the principle of least effort in communications, minimizing a combination of the information-theoretic communication inefficiency and direct signal cost. We prove a general relationship, for all optimal languages, between the signal cost distribution and the resulting distribution of signals. Zipf’s law then emerges for logarithmic signal cost distributions, which is the cost distribution expected for words constructed from letters or phonemes.</p>
</abstract>
<funding-group>
<funding-statement>CS and DP were supported by the European Commission as part of the CORBYS (Cognitive Control Framework for Robotic Systems) project under contract FP7 ICT-270219.</funding-statement>
</funding-group>
<counts>
<fig-count count="1"/>
<table-count count="0"/>
<page-count count="14"/>
</counts>
<custom-meta-group>
<custom-meta id="data-availability" xlink:type="simple">
<meta-name>Data Availability</meta-name>
<meta-value>All relevant data are within the paper.</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
</front>
<body>
<sec id="sec001" sec-type="intro">
<title>Introduction</title>
<p>Zipf’s law [<xref ref-type="bibr" rid="pone.0139475.ref001">1</xref>] for natural languages states that the frequency <italic>p</italic>(<italic>s</italic>) of a given word <italic>s</italic> in a large enough corpus of a (natural) language is inversely proportional to the word’s frequency rank. Zipf’s law postulates a power-law distribution for languages with a specific power law exponent <italic>β</italic>, so if <italic>s</italic><sub><italic>t</italic></sub> is the <italic>t</italic>-th most common word, then its frequency is proportional to
<disp-formula id="pone.0139475.e001"><alternatives><graphic id="pone.0139475.e001g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e001"/><mml:math id="M1" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>t</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>∼</mml:mo> <mml:mfrac><mml:mn>1</mml:mn> <mml:msup><mml:mi>t</mml:mi> <mml:mi>β</mml:mi></mml:msup></mml:mfrac> <mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(1)</label></disp-formula>
with <italic>β</italic> ≈ 1. Empirical data suggests that the power law holds across a variety of natural languages [<xref ref-type="bibr" rid="pone.0139475.ref002">2</xref>], but the exponent <italic>β</italic> can vary, depending on the language and the context, with a usual value of <italic>β</italic> ≈ 2 [<xref ref-type="bibr" rid="pone.0139475.ref003">3</xref>]. While the adherence to this “law” in different languages suggests a underlying common principle or mechanism, a generally accepted explanation for this phenomenon is still lacking [<xref ref-type="bibr" rid="pone.0139475.ref004">4</xref>].</p>
<p>Several papers [<xref ref-type="bibr" rid="pone.0139475.ref005">5</xref>–<xref ref-type="bibr" rid="pone.0139475.ref007">7</xref>] suggest that random texts already display a power law distribution sufficient to explain Zipf’s law, but a detailed analysis [<xref ref-type="bibr" rid="pone.0139475.ref008">8</xref>] with different statistical tests rejects this hypothesis and argues, that there is a “meaningful” mechanism at play, which causes this distribution across different natural languages.</p>
<p>If we reject the idea that Zipfian distribution are produced as a result of a process that randomly produces words, then the next logical step is to ask what models can produce such distributions and agrees with our basic assumptions about language? Mandelbrot [<xref ref-type="bibr" rid="pone.0139475.ref009">9</xref>] models language as a process of producing symbols, where each different symbol (word) has a specific cost. He argues that this cost grows logarithmically for more expensive symbols. He then considers the information of this process, and proves that a Zipfian distribution of the symbols produces the maximal information per cost ratio. Similar, more recent models [<xref ref-type="bibr" rid="pone.0139475.ref010">10</xref>] prove that power laws result from minimizing a logarithmic cost functions while maximising a process’s entropy (or self-information) [<xref ref-type="bibr" rid="pone.0139475.ref011">11</xref>]. But all these cost functions look at languages as a single random process, only optimizing the output distribution and ignoring any relationship between used words and intended meaning. This makes it a questionable model for human language (similar to the models with random text) as it does not account for communication efficiency, i.e., the model is not sensitive to how much information the words contain about the referenced concepts, nor does it offer any explanation on how certain words come to be assigned to certain meanings.</p>
<p>An alternative model by Cancho and Solé [<xref ref-type="bibr" rid="pone.0139475.ref012">12</xref>] follows the original idea of Zipf [<xref ref-type="bibr" rid="pone.0139475.ref001">1</xref>], by modelling the evolution of language based on the principle of least effort, where the assignment of words to concepts is optimized to minimize a weighted sum of speaker and listener effort. While simulations of the model produce distributions which qualitatively resemble power laws, a detailed mathematical investigation [<xref ref-type="bibr" rid="pone.0139475.ref004">4</xref>] reveals that the optimal solution of this model is, in fact, not following a power law; thus, the power law characteristics of the simulation results seems to be an artefact of the particular optimization model utilized.</p>
<p>Thus, to our knowledge, the question of how to achieve power laws in human language from the least effort principle is still not satisfactorily solved. Nevertheless, the idea from [<xref ref-type="bibr" rid="pone.0139475.ref001">1</xref>, <xref ref-type="bibr" rid="pone.0139475.ref012">12</xref>] to explain power laws as the result of an evolutionary optimization process that minimizes some form of language usage cost remains attractive. In this vein, we present an alternative model for the least effort principle in language: we minimize a cost function consisting of communication inefficiency and an inherent cost for each signal (word). To avoid past pitfalls of statistical analysis when looking for power laws [<xref ref-type="bibr" rid="pone.0139475.ref013">13</xref>], we offer mathematical proof that any optimal solution for our cost function necessarily realizes a power law distribution, as long as the underlying cost function for the signals increases logarithmically (if the signals are ordered according to cost rank). The result generalizes beyond this as we can state a general relationship between the cost structure of the individual signals and the resulting optimal distribution of the language signals.</p>
<p>We should also point out that a power-law often is not the best fit to real data [<xref ref-type="bibr" rid="pone.0139475.ref014">14</xref>]. However, the motivation of our study differs from that of [<xref ref-type="bibr" rid="pone.0139475.ref014">14</xref>] which attempted to find a mechanism, i.e., Random Group Formation (RGF), that fits and, crucially, predicts the data very well—instead, we attempt to find a model formalizing the least effort principle as a mechanism generating power laws.</p>
<p>Another important consideration is that there in general may be multiple mechanisms generating power laws, and one cannot <italic>post hoc</italic> reconstruct necessarily which mechanism resulted in the observed power law. We believe, however, that it is nevertheless useful to develop a mathematically rigorous version of such a mechanism (i.e., the least effort principle) applicable to languages in particular, as it would provide additional explanatory capacity in analyzing structures and patterns observed in languages [<xref ref-type="bibr" rid="pone.0139475.ref015">15</xref>, <xref ref-type="bibr" rid="pone.0139475.ref016">16</xref>].</p>
<p>The resulting insights may be of interest beyond the confines of power-law structures and offer an opportunity to study optimality conditions in other types of self-organizing coding systems, for instance in the case of the genetic code [<xref ref-type="bibr" rid="pone.0139475.ref017">17</xref>]. The suggested formalization covers a general class of optimal solutions balancing cost and efficiency, with power laws appearing as a special case. Furthermore, the proposed derivation highlights a connection between scaling in languages and thermodynamics, as the scaling exponent of the resulting power law is given by the corresponding inverse temperature (which in general relates the information-theoretic or statistical-mechanical interpretation of a system through its entropy and the system’s thermodynamics associated with its energy).</p>
</sec>
<sec id="sec002">
<title>1 Model</title>
<p>We will use a model, similar to that used by Ferrer i Cancho and Solé [<xref ref-type="bibr" rid="pone.0139475.ref012">12</xref>], which considers languages as an assignment of symbols to objects, and then optimizes this assignment function in regard to some form of combined speaker and listener effort. The language emerging from our model is also based on the optimality principle of least effort in communication, but uses a different cost function.</p>
<p>The model has a set of <italic>n</italic> signals <italic>S</italic> and a set of <italic>m</italic> objects <italic>R</italic>. Signals are used to reference objects, and a language is defined by how the speaker assigns signals to objects, i.e. by the relation between signals and objects. The relation between <italic>S</italic> and <italic>R</italic> in this model can be expressed by a binary matrix <italic>A</italic>, where an element <italic>a</italic><sub><italic>i</italic>,<italic>j</italic></sub> = 1 if and only if signal <italic>s</italic><sub><italic>i</italic></sub> refers to object <italic>r</italic><sub><italic>j</italic></sub>.</p>
<p>This model allows one to represent both <italic>polysemy</italic> (that is, the capacity for a signal to have multiple meanings by referring to multiple objects), and <italic>synonymy</italic>, where multiple signals refer to the same object. The relevant probabilities are then defined as follows:
<disp-formula id="pone.0139475.e002"><alternatives><graphic id="pone.0139475.e002g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e002"/><mml:math id="M2" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mfrac><mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:msub><mml:mi>ω</mml:mi> <mml:mi>j</mml:mi></mml:msub></mml:mfrac></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(2)</label></disp-formula>
where <italic>ω</italic><sub><italic>j</italic></sub> is the number of synonyms for object <italic>r</italic><sub><italic>j</italic></sub>, that is <italic>ω</italic><sub><italic>j</italic></sub> = ∑<sub><italic>i</italic></sub> <italic>a</italic><sub><italic>i</italic>,<italic>j</italic></sub>. Thus, the probability of using a synonym is equally distributed over all synonyms referring to a particular object. Importantly, it is also assumed that <inline-formula id="pone.0139475.e003"><alternatives><graphic id="pone.0139475.e003g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e003"/><mml:math id="M3" display="inline" overflow="scroll"><mml:mrow><mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo> <mml:mo>=</mml:mo> <mml:mfrac><mml:mn>1</mml:mn> <mml:mi>m</mml:mi></mml:mfrac></mml:mrow></mml:math></alternatives></inline-formula> is uniformly distributed over the objects, leading to a joint distribution:
<disp-formula id="pone.0139475.e004"><alternatives><graphic id="pone.0139475.e004g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e004"/><mml:math id="M4" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mfrac><mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mi>m</mml:mi> <mml:msub><mml:mi>ω</mml:mi> <mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfrac> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(3)</label></disp-formula>
In the previous model [<xref ref-type="bibr" rid="pone.0139475.ref012">12</xref>] each language has a cost based on a weighted combination of speaker and listener effort. The effort for the listener should be low if the received signal <italic>s</italic><sub><italic>i</italic></sub> leaves little ambiguity as to what object <italic>r</italic><sub><italic>j</italic></sub> is referenced, so there is little chance that the listener misunderstands what the speaker wanted to say. In the model of Ferrer i Cancho and Solé [<xref ref-type="bibr" rid="pone.0139475.ref012">12</xref>], the cost for listening to a specific signal <italic>s</italic><sub><italic>i</italic></sub> is expressed by the conditional entropy:
<disp-formula id="pone.0139475.e005"><alternatives><graphic id="pone.0139475.e005g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e005"/><mml:math id="M5" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo></mml:mrow> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>≡</mml:mo> <mml:mo>-</mml:mo> <mml:munderover><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>j</mml:mi> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn></mml:mrow> <mml:mi>m</mml:mi></mml:munderover> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>m</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(4)</label></disp-formula>
The overall effort for the listener is then dependent on the probability of each signal and the effort to decode it, that is
<disp-formula id="pone.0139475.e006"><alternatives><graphic id="pone.0139475.e006g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e006"/><mml:math id="M6" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo> <mml:mi>S</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>≡</mml:mo> <mml:munderover><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn></mml:mrow> <mml:mi>n</mml:mi></mml:munderover> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo></mml:mrow> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(5)</label></disp-formula>
Ferrer i Cancho and Solé argue that the listener effort is minimal when this entropy is minimal, in which case there is a deterministic mapping between signals and objects.</p>
<p>The effort for the speaker is expressed by the entropy <italic>H</italic><sub><italic>S</italic></sub>, which is, as the term in <xref ref-type="disp-formula" rid="pone.0139475.e006">Eq (5)</xref>, bound between 0 and 1, via the log with respect to <italic>n</italic>:
<disp-formula id="pone.0139475.e007"><alternatives><graphic id="pone.0139475.e007g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e007"/><mml:math id="M7" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>H</mml:mi> <mml:mi>S</mml:mi></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>≡</mml:mo> <mml:mo>-</mml:mo> <mml:munderover><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn></mml:mrow> <mml:mi>n</mml:mi></mml:munderover> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(6)</label></disp-formula>
Ferrer i Cancho and Solé then combine the listener’s and speaker’s efforts within the cost function Ω<sub><italic>λ</italic></sub> as follows:
<disp-formula id="pone.0139475.e008"><alternatives><graphic id="pone.0139475.e008g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e008"/><mml:math id="M8" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi></mml:msub> <mml:mo>=</mml:mo> <mml:mi>λ</mml:mi> <mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo> <mml:mi>S</mml:mi></mml:mrow></mml:msub> <mml:mo>+</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mi>H</mml:mi> <mml:mi>S</mml:mi></mml:msub> <mml:mspace width="4pt"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(7)</label></disp-formula>
with 0 ≤ <italic>λ</italic> ≤ 1.</p>
<p>It can be shown that the cost function Ω<sub><italic>λ</italic></sub> given by <xref ref-type="disp-formula" rid="pone.0139475.e008">Eq (7)</xref> is a specific case of a more general <italic>energy</italic> function that a communication system must minimize [<xref ref-type="bibr" rid="pone.0139475.ref004">4</xref>, <xref ref-type="bibr" rid="pone.0139475.ref018">18</xref>]
<disp-formula id="pone.0139475.e009"><alternatives><graphic id="pone.0139475.e009g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e009"/><mml:math id="M9" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mrow><mml:mi>λ</mml:mi></mml:mrow> <mml:mn>0</mml:mn></mml:msubsup> <mml:mo>=</mml:mo> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>I</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>S</mml:mi> <mml:mo>;</mml:mo> <mml:mi>R</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>+</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mi>H</mml:mi> <mml:mi>S</mml:mi></mml:msub> <mml:mspace width="4pt"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(8)</label></disp-formula>
where the mutual information <italic>I</italic>(<italic>S</italic>;<italic>R</italic>) = <italic>H</italic><sub><italic>R</italic></sub> − <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> captures the communication efficiency, i.e. how much information the signals contain about the objects. This energy function better accounts for subtle communication efforts [<xref ref-type="bibr" rid="pone.0139475.ref019">19</xref>], since <italic>H</italic><sub><italic>S</italic></sub> is arguably both a source of effort for the speaker and the listener because the word frequency affects not only word production but also recognition of spoken and written words [<xref ref-type="bibr" rid="pone.0139475.ref016">16</xref>]. The component <italic>I</italic>(<italic>S</italic>;<italic>R</italic>) also implicitly accounts for both <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> (a measure of the speaker’s effort of coding objects) and <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> (i.e., a measure of the listener’s effort of decoding signals). It is easy to see that
<disp-formula id="pone.0139475.e010"><alternatives><graphic id="pone.0139475.e010g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e010"/><mml:math id="M10" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mrow><mml:mi>λ</mml:mi></mml:mrow> <mml:mn>0</mml:mn></mml:msubsup> <mml:mo>=</mml:mo> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:msub><mml:mi>H</mml:mi> <mml:mi>R</mml:mi></mml:msub> <mml:mo>+</mml:mo> <mml:mi>λ</mml:mi> <mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo> <mml:mi>S</mml:mi></mml:mrow></mml:msub> <mml:mo>+</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mi>H</mml:mi> <mml:mi>S</mml:mi></mml:msub> <mml:mo>=</mml:mo> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:msub><mml:mi>H</mml:mi> <mml:mi>R</mml:mi></mml:msub> <mml:mo>+</mml:mo> <mml:msub><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi></mml:msub> <mml:mspace width="4pt"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(9)</label></disp-formula>
and so when the entropy <italic>H</italic><sub><italic>R</italic></sub> is constant, e.g. under the uniformity condition <inline-formula id="pone.0139475.e011"><alternatives><graphic id="pone.0139475.e011g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e011"/><mml:math id="M11" display="inline" overflow="scroll"><mml:mrow><mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo> <mml:mo>=</mml:mo> <mml:mfrac><mml:mn>1</mml:mn> <mml:mi>m</mml:mi></mml:mfrac></mml:mrow></mml:math></alternatives></inline-formula>, the more generic energy function <inline-formula id="pone.0139475.e012"><alternatives><graphic id="pone.0139475.e012g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e012"/><mml:math id="M12" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mn>0</mml:mn></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> reduces to the specific Ω<sub><italic>λ</italic></sub>.</p>
<p>We propose instead another cost function that not only produces optimal languages exhibiting power laws, but also retains the clear intuition of generic energy functions which typically reflect the global quality of a solution. Firstly, we represent the communication inefficiency by the information distance, the Rokhlin metric, <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> + <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> [<xref ref-type="bibr" rid="pone.0139475.ref020">20</xref>, <xref ref-type="bibr" rid="pone.0139475.ref021">21</xref>]. This distance is often more sensitive than − <italic>I</italic>(<italic>S</italic>;<italic>R</italic>) in measuring the “disagreements” between variables, especially in the case when one information source is contained within another [<xref ref-type="bibr" rid="pone.0139475.ref022">22</xref>].</p>
<p>Secondly, we define the signal usage effort by introducing an explicit cost function <italic>c</italic>(<italic>s</italic><sub><italic>i</italic></sub>), which assigns each signal a specific cost. The signal usage cost for a language is then the weighted average of this signal specific cost:
<disp-formula id="pone.0139475.e013"><alternatives><graphic id="pone.0139475.e013g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e013"/><mml:math id="M13" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:munderover><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn></mml:mrow> <mml:mi>n</mml:mi></mml:munderover> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mi>c</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(10)</label></disp-formula>
This is motivated by the basic idea that words have an intrinsic cost associated with using (speaking, writing, hearing, reading) them. To illustrate, a version of English where each use of the word “I” is replaced with “Antidisestablishmentarianism” and vice versa should not have the same signal usage cost as normal English. The optimal solution considering the signal usage cost alone would be to reference every object with the cheapest signal.</p>
<p>The overall cost function for a language <inline-formula id="pone.0139475.e014"><alternatives><graphic id="pone.0139475.e014g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e014"/><mml:math id="M14" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> is the energy function trading off the communicative inefficiency with the signal usage cost, with 0 &lt; <italic>λ</italic> ≤ 1 trading off the efforts as follows:
<disp-formula id="pone.0139475.e015"><alternatives><graphic id="pone.0139475.e015g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e015"/><mml:math id="M15" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mrow><mml:mi>λ</mml:mi></mml:mrow> <mml:mi>c</mml:mi></mml:msubsup> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mi>λ</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>S</mml:mi> <mml:mo>|</mml:mo> <mml:mi>R</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>+</mml:mo> <mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo> <mml:mi>S</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>+</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:munderover><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn></mml:mrow> <mml:mi>n</mml:mi></mml:munderover> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mi>c</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(11)</label></disp-formula>
where <italic>p</italic> = <italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>, <italic>r</italic><sub><italic>j</italic></sub>) is the joint probability. A language can be optimized for different values of <italic>λ</italic>, weighting the respective costs. The extreme case (<italic>λ</italic> = 0) with only the signal usage cost defining the energy function is excluded, while the opposite extreme (<italic>λ</italic> = 1) focusing on the communication inefficiency is considered. Following the principle of least effort, we aim to determine the properties of those languages that have minimal cost according to <inline-formula id="pone.0139475.e016"><alternatives><graphic id="pone.0139475.e016g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e016"/><mml:math id="M16" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula>.</p>
</sec>
<sec id="sec003" sec-type="results">
<title>2 Results</title>
<p>First of all, we establish that all local minimizers, and hence all global minimizers, of the cost function <xref ref-type="disp-formula" rid="pone.0139475.e015">(11)</xref> are solutions without synonyms. Formally, we obtain the following result.</p>
<p><bold>Theorem 1.</bold> <italic>Each local minimizer of the function</italic> <disp-formula id="pone.0139475.e017">
<alternatives>
<graphic id="pone.0139475.e017g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e017"/>
<mml:math id="M17" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:mo>𝓒</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>→</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>ℝ</mml:mo>
<mml:mo>,</mml:mo>
<mml:mspace width="2.em"/>
<mml:mi>p</mml:mi>
<mml:mspace width="0.277778em"/>
<mml:mo>↦</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:msubsup>
<mml:mo>Ω</mml:mo>
<mml:mi>λ</mml:mi>
<mml:mi>c</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
</disp-formula> <italic>where</italic> <disp-formula id="pone.0139475.e018">
<alternatives>
<graphic id="pone.0139475.e018g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e018"/>
<mml:math id="M18" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:mo>𝓒</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>:</mml:mo>
<mml:mo>=</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>{</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>∈</mml:mo>
<mml:mo>𝓟</mml:mo>
<mml:mo>(</mml:mo>
<mml:mi>S</mml:mi>
<mml:mo>×</mml:mo>
<mml:mi>R</mml:mi>
<mml:mo>)</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>:</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mspace width="0.277778em"/>
<mml:mo>=</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:msub>
<mml:mo>∑</mml:mo>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>m</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mspace width="4.pt"/>
<mml:mtext mathvariant="italic">for</mml:mtext>
<mml:mspace width="4.pt"/>
<mml:mtext mathvariant="italic">all</mml:mtext>
<mml:mspace width="4.pt"/>
<mml:mi>j</mml:mi>
<mml:mo>}</mml:mo>
<mml:mspace width="4pt"/>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
</disp-formula> <italic>and</italic> <inline-formula id="pone.0139475.e019"><alternatives><graphic id="pone.0139475.e019g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e019"/><mml:math id="M19" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup> <mml:mo stretchy="false">(</mml:mo> <mml:mi>p</mml:mi> <mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></alternatives></inline-formula> <italic>is specified by the</italic> <xref ref-type="disp-formula" rid="pone.0139475.e015">Eq (11)</xref>, 0 &lt; <italic>λ</italic> ≤ 1, <italic>can be represented as a function</italic> <italic>f</italic> : <italic>R</italic> → <italic>S</italic> <italic>such that</italic> <disp-formula id="pone.0139475.e020">
<alternatives>
<graphic id="pone.0139475.e020g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e020"/>
<mml:math id="M20" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mo>{</mml:mo>
<mml:mtable>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mspace width="4.pt"/>
<mml:mspace width="4.pt"/>
<mml:mrow><mml:mi>i</mml:mi><mml:mi>f</mml:mi></mml:mrow>
<mml:mspace width="4.pt"/>
<mml:mspace width="4.pt"/>
<mml:mspace width="4.pt"/>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>;</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mspace width="4.pt"/>
<mml:mspace width="4.pt"/>
<mml:mrow><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mo>.</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
<mml:mo/>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
<label>(12)</label>
</disp-formula>
The proof is given in Appendix 1. Note that each solution, i.e. each distribution <italic>p</italic> in expression <xref ref-type="disp-formula" rid="pone.0139475.e004">(3)</xref>, corresponds to a matrix <italic>A</italic> (henceforth called <italic>minimizer matrix</italic>) which is given in terms of function <italic>f</italic> as follows:
<disp-formula id="pone.0139475.e021"><alternatives><graphic id="pone.0139475.e021g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e021"/><mml:math id="M21" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>=</mml:mo> <mml:mo>{</mml:mo> <mml:mtable><mml:mtr><mml:mtd columnalign="left"><mml:mn>1</mml:mn></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mspace width="4.pt"/><mml:mspace width="4.pt"/><mml:mtext>if</mml:mtext> <mml:mspace width="4.pt"/><mml:mspace width="4.pt"/><mml:mspace width="4.pt"/><mml:mrow><mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>=</mml:mo> <mml:mi>f</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mrow> <mml:mo>;</mml:mo></mml:mrow></mml:mtd></mml:mtr> <mml:mtr><mml:mtd columnalign="left"><mml:mn>0</mml:mn></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mspace width="4.pt"/><mml:mspace width="4.pt"/><mml:mtext>otherwise.</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable> <mml:mo/></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(13)</label></disp-formula>
The main outcome of this observation is that the analytical minimization of the suggested cost function results in solutions without synonyms—since any function <italic>f</italic> precludes multiple signals <italic>s</italic> referring to the same object <italic>r</italic>. That is, each column in the minimizer matrix has precisely one non-zero element. Polysemy is allowed within the solutions.</p>
<p>We need the following lemma as an intermediate step towards deriving the analytical relationship between the specific word cost <italic>c</italic>(<italic>s</italic>) and the resulting distribution <italic>p</italic>(<italic>s</italic>).</p>
<p><bold>Lemma 2.</bold> <italic>For each solution</italic> <italic>p</italic> <italic>minimizing the function</italic> <inline-formula id="pone.0139475.e022"><alternatives><graphic id="pone.0139475.e022g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e022"/><mml:math id="M22" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula>, <disp-formula id="pone.0139475.e023">
<alternatives>
<graphic id="pone.0139475.e023g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e023"/>
<mml:math id="M23" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mo>|</mml:mo>
<mml:mi>S</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mo form="prefix">log</mml:mo>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mi>S</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mspace width="4pt"/>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
<label>(14)</label>
</disp-formula>
The proof follows from the joint entropy representations
<disp-formula id="pone.0139475.e024"><alternatives><graphic id="pone.0139475.e024g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e024"/><mml:math id="M24" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>S</mml:mi> <mml:mo>,</mml:mo> <mml:mi>R</mml:mi></mml:mrow></mml:msub> <mml:mo>=</mml:mo> <mml:mfrac><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo> <mml:mi>S</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mn>1</mml:mn> <mml:mo>+</mml:mo> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>m</mml:mi></mml:msub> <mml:mi>n</mml:mi></mml:mrow></mml:mfrac> <mml:mo>+</mml:mo> <mml:mfrac><mml:msub><mml:mi>H</mml:mi> <mml:mi>S</mml:mi></mml:msub> <mml:mrow><mml:mn>1</mml:mn> <mml:mo>+</mml:mo> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>m</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(15)</label></disp-formula> <disp-formula id="pone.0139475.e025"><alternatives><graphic id="pone.0139475.e025g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e025"/><mml:math id="M25" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mspace width="2.em"/><mml:mo>=</mml:mo> <mml:mfrac><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>S</mml:mi> <mml:mo>|</mml:mo> <mml:mi>R</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mn>1</mml:mn> <mml:mo>+</mml:mo> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>m</mml:mi></mml:mrow></mml:mfrac> <mml:mo>+</mml:mo> <mml:mfrac><mml:msub><mml:mi>H</mml:mi> <mml:mi>R</mml:mi></mml:msub> <mml:mrow><mml:mn>1</mml:mn> <mml:mo>+</mml:mo> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>m</mml:mi></mml:msub> <mml:mi>n</mml:mi></mml:mrow></mml:mfrac> <mml:mspace width="4pt"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(16)</label></disp-formula>
noting that for each minimal solution <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> = 0, while <italic>H</italic><sub><italic>R</italic></sub> = 1 under the uniformity constraint <inline-formula id="pone.0139475.e026"><alternatives><graphic id="pone.0139475.e026g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e026"/><mml:math id="M26" display="inline" overflow="scroll"><mml:mrow><mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo> <mml:mo>=</mml:mo> <mml:mfrac><mml:mn>1</mml:mn> <mml:mi>m</mml:mi></mml:mfrac></mml:mrow></mml:math></alternatives></inline-formula>.</p>
<p><bold>Corollary 3.</bold> <italic>If n = m, H<sub>R∣S</sub> + H<sub>S</sub></italic> = 1.</p>
<p>Using this lemma, and noting that each such solution represented as a function <italic>f</italic> : <italic>R</italic> → <italic>S</italic> has the property <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> = 0, we reduce the <xref ref-type="disp-formula" rid="pone.0139475.e015">Eq (11)</xref> to
<disp-formula id="pone.0139475.e027"><alternatives><graphic id="pone.0139475.e027g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e027"/><mml:math id="M27" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mrow><mml:mi>λ</mml:mi></mml:mrow> <mml:mi>c</mml:mi></mml:msubsup> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mfrac><mml:mn>1</mml:mn> <mml:mrow><mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>m</mml:mi></mml:mrow></mml:mfrac> <mml:msub><mml:mi>H</mml:mi> <mml:mi>S</mml:mi></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>)</mml:mo> <mml:mo>+</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>∑</mml:mo> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mi>c</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(17)</label></disp-formula> <disp-formula id="pone.0139475.e028"><alternatives><graphic id="pone.0139475.e028g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e028"/><mml:math id="M28" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mspace width="1.em"/><mml:mo>=</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>+</mml:mo> <mml:mfrac><mml:mi>λ</mml:mi> <mml:mrow><mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>m</mml:mi></mml:mrow></mml:mfrac> <mml:mo>∑</mml:mo> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>+</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>∑</mml:mo> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mi>c</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(18)</label></disp-formula></p>
<p>Varying with respect to <italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>), under the constraint ∑<italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>) = 1, yields the extremality condition
<disp-formula id="pone.0139475.e029"><alternatives><graphic id="pone.0139475.e029g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e029"/><mml:math id="M29" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mfrac><mml:mi>λ</mml:mi> <mml:mrow><mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>m</mml:mi></mml:mrow></mml:mfrac> <mml:mo>(</mml:mo> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>+</mml:mo> <mml:mn>1</mml:mn> <mml:mo>)</mml:mo> <mml:mo>+</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mi>c</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>-</mml:mo> <mml:msup><mml:mi>κ</mml:mi> <mml:mo>′</mml:mo></mml:msup> <mml:mo>=</mml:mo> <mml:mn>0</mml:mn> <mml:mspace width="4pt"/></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(19)</label></disp-formula>
for some Lagrange multiplier <italic>κ</italic>′. The minimum is achieved when
<disp-formula id="pone.0139475.e030"><alternatives><graphic id="pone.0139475.e030g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e030"/><mml:math id="M30" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mi>κ</mml:mi> <mml:msup><mml:mi>e</mml:mi> <mml:mrow><mml:mo>-</mml:mo> <mml:mi>β</mml:mi> <mml:mi>c</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:msup> <mml:mspace width="4pt"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(20)</label></disp-formula>
where
<disp-formula id="pone.0139475.e031"><alternatives><graphic id="pone.0139475.e031g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e031"/><mml:math id="M31" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>β</mml:mi> <mml:mo>=</mml:mo> <mml:mfrac><mml:mrow><mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi></mml:mrow> <mml:mi>λ</mml:mi></mml:mfrac> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub> <mml:mi>m</mml:mi> <mml:mspace width="4pt"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(21)</label></disp-formula> <disp-formula id="pone.0139475.e032"><alternatives><graphic id="pone.0139475.e032g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e032"/><mml:math id="M32" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>κ</mml:mi> <mml:mo>=</mml:mo> <mml:mfrac><mml:mn>1</mml:mn> <mml:mrow><mml:mo>∑</mml:mo> <mml:msup><mml:mi>e</mml:mi> <mml:mrow><mml:mo>-</mml:mo> <mml:mi>β</mml:mi> <mml:mi>c</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfrac> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(22)</label></disp-formula>
In addition, we require
<disp-formula id="pone.0139475.e033"><alternatives><graphic id="pone.0139475.e033g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e033"/><mml:math id="M33" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>c</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mo form="prefix">ln</mml:mo> <mml:msub><mml:mi>m</mml:mi> <mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(23)</label></disp-formula>
for some integer <italic>m</italic><sub><italic>i</italic></sub> such that ∑<italic>m</italic><sub><italic>i</italic></sub> = <italic>m</italic>. The last condition ensures that the minimal solutions <italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>) correspond to functions <italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>, <italic>r</italic><sub><italic>j</italic></sub>) (i.e., minimizer matrices without synonyms). In other words, the marginal probability <xref ref-type="disp-formula" rid="pone.0139475.e030">(20)</xref> without the condition <xref ref-type="disp-formula" rid="pone.0139475.e033">(23)</xref> may not concur with the probability <italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>, <italic>r</italic><sub><italic>j</italic></sub>) that represents a minimizer matrix under the uniformity constraint <inline-formula id="pone.0139475.e034"><alternatives><graphic id="pone.0139475.e034g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e034"/><mml:math id="M34" display="inline" overflow="scroll"><mml:mrow><mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo> <mml:mo>=</mml:mo> <mml:mfrac><mml:mn>1</mml:mn> <mml:mi>m</mml:mi></mml:mfrac></mml:mrow></mml:math></alternatives></inline-formula>.</p>
<p>Under the condition <xref ref-type="disp-formula" rid="pone.0139475.e033">(23)</xref>, we have <inline-formula id="pone.0139475.e035"><alternatives><graphic id="pone.0139475.e035g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e035"/><mml:math id="M35" display="inline" overflow="scroll"><mml:mrow><mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo> <mml:mo>=</mml:mo> <mml:mi>κ</mml:mi> <mml:msubsup><mml:mi>m</mml:mi> <mml:mi>i</mml:mi> <mml:mrow><mml:mo>−</mml:mo> <mml:mi>β</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula>, while <inline-formula id="pone.0139475.e036"><alternatives><graphic id="pone.0139475.e036g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e036"/><mml:math id="M36" display="inline" overflow="scroll"><mml:mrow><mml:mi>κ</mml:mi> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn> <mml:mo>/</mml:mo> <mml:mo>∑</mml:mo> <mml:msubsup><mml:mi>m</mml:mi> <mml:mi>i</mml:mi> <mml:mrow><mml:mo>−</mml:mo> <mml:mi>β</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula>. In general, one may relax the condition <xref ref-type="disp-formula" rid="pone.0139475.e033">(23)</xref>, specifying instead an upper-bounded error of approximating the minimal solution by any <italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>) = <italic>κe</italic><sup>−<italic>βc</italic>(<italic>s</italic><sub><italic>i</italic></sub>)</sup> which would then allow for arbitrary cost functions <italic>c</italic>(<italic>s</italic>).</p>
<p>Interestingly, the optimal marginal probability distribution <xref ref-type="disp-formula" rid="pone.0139475.e030">(20)</xref> is the Gibbs measure with the energy <italic>c</italic>(<italic>s</italic><sub><italic>i</italic></sub>), while the parameter <italic>β</italic> is, thermodynamically, the inverse temperature. It is well-known that the Gibbs measure is the unique measure maximizing the entropy for a given expected energy, and appears in many solutions outside of thermodynamics [<xref ref-type="bibr" rid="pone.0139475.ref023">23</xref>–<xref ref-type="bibr" rid="pone.0139475.ref025">25</xref>].</p>
<p>Let us now consider some special cases. For the case of equal effort, i.e. <italic>λ</italic> = 0.5, and <italic>n</italic> = <italic>m</italic>, the solution simplifies to <italic>β</italic> = 1 and <inline-formula id="pone.0139475.e037"><alternatives><graphic id="pone.0139475.e037g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e037"/><mml:math id="M37" display="inline" overflow="scroll"><mml:mrow><mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo> <mml:mo>=</mml:mo> <mml:mi>κ</mml:mi> <mml:msubsup><mml:mi>m</mml:mi> <mml:mi>i</mml:mi> <mml:mrow><mml:mo>−</mml:mo> <mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula>, where <inline-formula id="pone.0139475.e038"><alternatives><graphic id="pone.0139475.e038g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e038"/><mml:math id="M38" display="inline" overflow="scroll"><mml:mrow><mml:mi>κ</mml:mi> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn> <mml:mo>/</mml:mo> <mml:mrow><mml:mo>∑</mml:mo> <mml:msubsup><mml:mi>m</mml:mi> <mml:mi>i</mml:mi> <mml:mrow><mml:mo>−</mml:mo> <mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mrow></mml:math></alternatives></inline-formula>.</p>
<p>Another important special case is given by the cost function <italic>c</italic>(<italic>s</italic><sub><italic>i</italic></sub>) = ln <italic>ρ</italic><sub><italic>i</italic></sub>/<italic>N</italic>, where <italic>ρ</italic><sub><italic>i</italic></sub> is the rank of symbol <italic>s</italic><sub><italic>i</italic></sub>, and <italic>N</italic> is a normalization constant equal to <inline-formula id="pone.0139475.e039"><alternatives><graphic id="pone.0139475.e039g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e039"/><mml:math id="M39" display="inline" overflow="scroll"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>n</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:mi>n</mml:mi> <mml:mo>+</mml:mo> <mml:mn>1</mml:mn> <mml:mo stretchy="false">)</mml:mo></mml:mrow> <mml:mrow><mml:mn>2</mml:mn> <mml:mi>m</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:math></alternatives></inline-formula> (so that ∑<italic>ρ</italic><sub><italic>i</italic></sub>/<italic>N</italic> = <italic>m</italic>). In this case, the optimal solution is attained when
<disp-formula id="pone.0139475.e040"><alternatives><graphic id="pone.0139475.e040g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e040"/><mml:math id="M40" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mfrac><mml:mrow><mml:mi>κ</mml:mi> <mml:msup><mml:mi>N</mml:mi> <mml:mi>β</mml:mi></mml:msup></mml:mrow> <mml:msubsup><mml:mi>ρ</mml:mi> <mml:mi>i</mml:mi> <mml:mi>β</mml:mi></mml:msubsup></mml:mfrac></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(24)</label></disp-formula>
with
<disp-formula id="pone.0139475.e041"><alternatives><graphic id="pone.0139475.e041g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e041"/><mml:math id="M41" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>κ</mml:mi> <mml:mo>=</mml:mo> <mml:mfrac><mml:mn>1</mml:mn> <mml:mrow><mml:msup><mml:mi>N</mml:mi> <mml:mi>β</mml:mi></mml:msup> <mml:mo>∑</mml:mo> <mml:msubsup><mml:mi>ρ</mml:mi> <mml:mi>j</mml:mi> <mml:mrow><mml:mo>-</mml:mo> <mml:mi>β</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(25)</label></disp-formula>
This means that a power law with the exponent <italic>β</italic>, specified by <xref ref-type="disp-formula" rid="pone.0139475.e031">Eq (21)</xref>, is the optimal solution in regard to our cost function <xref ref-type="disp-formula" rid="pone.0139475.e015">(11)</xref> if the signal usage cost increases logarithmically. In this case, the exponent <italic>β</italic> depends on the system’s size (<italic>n</italic> and <italic>m</italic>) and the efforts’ trade-off <italic>λ</italic>. Importantly, this derivation shows a connection between scaling in languages and thermodynamics: if the signal usage cost increases logarithmically, then the scaling exponent of the resulting power law is given by the corresponding inverse temperature.</p>
<p>Zipf’s law (a power law with exponent <italic>β</italic> = 1) is then nothing but a special case for systems that satisfy <inline-formula id="pone.0139475.e042"><alternatives><graphic id="pone.0139475.e042g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e042"/><mml:math id="M42" display="inline" overflow="scroll"><mml:mrow><mml:msub><mml:mtext>log</mml:mtext> <mml:mi>n</mml:mi></mml:msub><mml:mspace width="0.277778em"/> <mml:mi>m</mml:mi> <mml:mo>=</mml:mo> <mml:mfrac><mml:mi>λ</mml:mi> <mml:mrow><mml:mn>1</mml:mn> <mml:mo>−</mml:mo> <mml:mi>λ</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:math></alternatives></inline-formula>. For instance, for square matrices, Zipf’s law results from the optimal languages which satisfy equal efforts, i.e., <italic>λ</italic> = 0.5. The importance of equal cost was emphasized in earlier works [<xref ref-type="bibr" rid="pone.0139475.ref004">4</xref>, <xref ref-type="bibr" rid="pone.0139475.ref026">26</xref>]. The exponent defined by <xref ref-type="disp-formula" rid="pone.0139475.e031">Eq (21)</xref> changes with the system size (<italic>n</italic> or <italic>m</italic>), and so the resulting power law “adapts” to linguistic dynamics and language evolution in general.</p>
<p>The assumption that the cost function is precisely logarithmic results in an exact power law. If, on the other hand, the cost function deviates from being precisely logarithmic, then the resulting dependency would only approximate a power law—this imprecision may in fact account for different degrees of success in fitting power laws to real data.</p>
<p>In summary, the derived relationship expresses the optimal probability <italic>p</italic>(<italic>s</italic>) in terms of the usage cost <italic>c</italic>(<italic>s</italic>), yielding Zipf’s law when this cost is logarithmically distributed over the symbols.</p>
</sec>
<sec id="sec004" sec-type="conclusions">
<title>3 Discussion</title>
<p>To explain the emergence of power laws for signal selection, we need to explain why the cost function of the signals would increase logarithmically, if the signals are ordered by their cost rank. This can be motivated, across a number of languages, by assuming that signals are in fact words, which are made up of letters from a finite alphabet; or in regard to spoken language, are made of from a finite set of phonemes. Compare [<xref ref-type="bibr" rid="pone.0139475.ref027">27</xref>], in which Nowak and Krakauer demonstrate how the error limits of communication with a finite list of phonemes can be overcome by combining phonemes into words.</p>
<p>Lets assume that each letter (or phoneme) has an inherent cost which is approximate to a unit letter cost. Furthermore, assume that the cost of a word roughly equals the sum of its letter costs. A language with an alphabet of size <italic>a</italic> then has <italic>a</italic> unique one letter words which the approximate cost of one, <italic>a</italic><sup>2</sup> two letter words with an approximate cost of two, <italic>a</italic><sup>3</sup> three letter words with a cost of three, etcetera. If we rank these words by their cost, then their cost will increase approximately logarithmically with their cost rank. To illustrate, <xref ref-type="fig" rid="pone.0139475.g001">Fig 1</xref> is a plot of the 1000 cheapest unique words formed with a ten letter alphabet (with no word length restriction), where each letter has a random cost between 1.0 and 2.0. The first few words deviate from the logarithmic cost function, as their cost only depends on the letter cost itself, but the latter words closely follow a logarithmic function. A similar derivation of the logarithmic cost function from first principles can be found in the model of Mandelbrot [<xref ref-type="bibr" rid="pone.0139475.ref009">9</xref>].</p>
<fig id="pone.0139475.g001" position="float">
<object-id pub-id-type="doi">10.1371/journal.pone.0139475.g001</object-id>
<label>Fig 1</label>
<caption>
<title>A log-plot of the 1000 cheapest words created from a 10 letter alphabet, ordered by their cost rank.</title>
<p>Word cost is a sum of individual letter cost, and letter cost is between 1.0 and 2.0 units.</p>
</caption>
<graphic mimetype="image" xlink:type="simple" position="float" xlink:href="info:doi/10.1371/journal.pone.0139475.g001"/>
</fig>
<p>This signal usage cost can be interpreted in different ways. In spoken language it might simply be the time needed to utter a word, which makes it a cost both for the listener and the speaker. In written language it might be the effort to write a word, or the bandwidth needed to transmit it, in which case it is a speaker cost. On the other hand, if one is reading a written text, then the length of the words might translate into “listener” cost again. In general, the average signal usage cost corresponds to the effort of using a specific language to communicate for all involved parties. This differs from the original least effort idea, which balances listener and speaker effort [<xref ref-type="bibr" rid="pone.0139475.ref001">1</xref>]. In our model we balance the general effort of using the language with the communication efficiency, which creates a similar tension, as described in [<xref ref-type="bibr" rid="pone.0139475.ref012">12</xref>], between using a language that only uses one signal, and a language that references every object with its own signal. If only communication efficiency was relevant, then each object would have its own signal. Conversely, if only cost mattered, then all objects would be referenced by the same cheapest signal. Balancing these two components with a weighting factor <italic>λ</italic> yields power laws, where <italic>β</italic> varies with changes in the weighting factor. This is in contrast to the model in [<xref ref-type="bibr" rid="pone.0139475.ref012">12</xref>], where power laws were only found in a phase transition along the weighting factor. Also, in [<xref ref-type="bibr" rid="pone.0139475.ref003">3</xref>] Cancho discusses how some variants of language (military, children) have <italic>β</italic> values that deviate from the <italic>β</italic> value of their base language, which could indicate that the effort of language production or communication efficiency is weighted differently in these cases, resulting in different optimal solutions, which are power laws with other values for <italic>β</italic>.</p>
<p>We noted earlier that there are other options to produce power laws, which are insensitive to the relationship between objects and signals. Baek et al. [<xref ref-type="bibr" rid="pone.0139475.ref014">14</xref>] obtain a power law by minimizing the cost function <italic>I</italic><sub><italic>cost</italic></sub> = −<italic>H</italic><sub><italic>S</italic></sub> + 〈log <italic>s</italic>〉 + log <italic>N</italic>, where 〈log <italic>s</italic>〉 = ∑<italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>)log(<italic>s</italic><sub><italic>i</italic></sub>), and log(<italic>s</italic><sub><italic>i</italic></sub>) is interpreted as the logarithm of the index of <italic>s</italic><sub><italic>i</italic></sub> (specifically, its rank). Their argument that this cost function follows from a more general cost function <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> = −<italic>I</italic>(<italic>S</italic>;<italic>R</italic>) + <italic>H</italic><sub><italic>R</italic></sub>, where <italic>H</italic><sub><italic>R</italic></sub> is constant, is undermined by their unconventional definition of conditional probability (cf. Appendix A [<xref ref-type="bibr" rid="pone.0139475.ref014">14</xref>]). Specifically, this probability is defined as <inline-formula id="pone.0139475.e043"><alternatives><graphic id="pone.0139475.e043g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e043"/><mml:math id="M43" display="inline" overflow="scroll"><mml:mrow><mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:mi>r</mml:mi> <mml:mo stretchy="false">∣</mml:mo> <mml:mi>s</mml:mi> <mml:mo stretchy="false">)</mml:mo> <mml:mo>=</mml:mo> <mml:mfrac><mml:msub><mml:mi>δ</mml:mi> <mml:mrow><mml:msup><mml:mi>s</mml:mi> <mml:mo>′</mml:mo></mml:msup> <mml:mo stretchy="false">(</mml:mo> <mml:mi>r</mml:mi> <mml:mo stretchy="false">)</mml:mo> <mml:mo>,</mml:mo> <mml:mi>s</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mi>s</mml:mi> <mml:mi>N</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:mi>s</mml:mi> <mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></alternatives></inline-formula>, where <italic>N</italic>(<italic>s</italic>) is the number of objects to which signal <italic>s</italic> refers. This definition not only requires some additional assumptions in order to make <italic>p</italic>(<italic>r</italic>∣<italic>s</italic>) a conditional probability, but also implicitly embeds the “cost” of symbol <italic>s</italic> within the conditional probability <italic>p</italic>(<italic>r</italic>∣<italic>s</italic>), by dividing it by <italic>s</italic>. Thus, we are left with the cost function <italic>I</italic><sub><italic>cost</italic></sub> <italic>per se</italic>, not rigorously derived from a generic principle, and this cost function ignores joint probabilities and the communication efficiency in particular.</p>
<p>A very similar cost function was offered by Visser [<xref ref-type="bibr" rid="pone.0139475.ref010">10</xref>], who suggested to maximize <italic>H</italic><sub><italic>S</italic></sub> subject to a constraint 〈log <italic>s</italic>〉 = <italic>χ</italic>, for some constant <italic>χ</italic>. Again, this maximization produces a power law, and again we may note that the cost function and the constraint used in the derivation do not capture communication efficiency or trade-offs between speaker and listener, omitting joint probabilities as well.</p>
<p>Finally, we would like to point out that the cost function −<italic>H</italic><sub><italic>S</italic></sub> + 〈log <italic>s</italic>〉 is equivalent to the cost function <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> − <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> + 〈log <italic>s</italic>〉, under constant <italic>H</italic><sub><italic>R</italic></sub>. This expression reveals another important drawback of minimizing −<italic>H</italic><sub><italic>S</italic></sub> + 〈log <italic>s</italic>〉 directly: while minimizing <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> reduces the ambiguity of polysemy, minimizing −<italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> explicitly “rewards” the ambiguity of synonyms. In other words, languages obtained by minimizing such a cost directly do exhibit a power law, but mostly at the expense of potentially unnecessary synonyms.</p>
<p>There may be a number of reasons for the avoidance of synonyms in real languages. While an analysis of synonymy dynamics in child languages or aphasiacs is outside of scope of this paper, it is worth pointing out that some studies have suggested that the learning of new words by children is driven by synonymy avoidance [<xref ref-type="bibr" rid="pone.0139475.ref028">28</xref>]. As the vocabulary and the word use are growing in children (with meaning overextensions decreasing over time), reducing the effort for the listener becomes more important [<xref ref-type="bibr" rid="pone.0139475.ref029">29</xref>]. Several principles underlying lexicon acquisition by children, identified by Clark [<xref ref-type="bibr" rid="pone.0139475.ref030">30</xref>], emphasize the dynamics of synonymy reduction. For example, the principle of <italic>conventionality</italic> and <italic>contrast</italic> (“speakers take every difference in form to mark a difference in meaning”) combine in providing some precedence to semantic overlaps, leading children to eventually accept the parents’ (more conventional) word for a semantically overlapping concept. The principle of <italic>transparency</italic> explains how a preference to use a more transparent word helps to reduce ambiguity in the lexicon. It has also been recently shown that the exponent of Zipf’s law (when rank is the random variable) tends to decrease over time in children [<xref ref-type="bibr" rid="pone.0139475.ref031">31</xref>]. The study correlated this evolution of the exponent with the reduction of a simple indicator of syntactic complexity given by the mean length of utterances (MLU), and concluded that this supports the hypothesis that the inter-related exponent of Zipf’s law and linguistic complexity tend to decrease in parallel.</p>
<p>Regarding synonyms it should also be noted, that while they exist, their number is usually comparatively low. If we are looking at a natural language, which might have ca. 100.000 words, we will not find a concept that has 95.000 synonyms. Most concepts have synonyms in the single digits, if they have any. The models that look at just the output distribution could produce languages with such an excessive number of synonyms. In our model the ideal solution has no synonyms, but the existing languages, which are constantly adapting, could be seen as close approximations, where out of 100.000 possible synonyms, most concepts have only very few synonyms, if any. As noted earlier, while precise logarithmic cost functions would produce perfect power-law distributions, natural languages do not fit Zipf’s law exactly but only approximately.</p>
<p>These observations support our conjecture that, as languages mature, the communicative efficiency and the balance between speaker’s and listener’s efforts become a more significant driver, and so the simplistic cost function −<italic>H</italic><sub><italic>S</italic></sub> + 〈log <italic>s</italic>〉 can no longer be justified.</p>
<p>In contrast, the cost function proposed in this paper <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> + <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> + 〈log <italic>s</italic>〉 reduces to −<italic>H</italic><sub><italic>S</italic></sub> + 〈log <italic>s</italic>〉 only <italic>after</italic> minimizing over the joint probabilities <italic>p</italic>(<italic>s</italic>, <italic>r</italic>). Importantly, it captures communication (in)efficiency and average signal usage explicitly, balancing out different aspects of the communication trade-offs and representing the concept of least effort in a principled way. The resulting solutions do not contain synonyms, which disappear at the step of minimizing over <italic>p</italic>(<italic>s</italic>, <italic>r</italic>), and so correspond to “perfect”, maximally efficient and balanced, languages. The fact that even these languages exhibit power (Zipf’s) laws is a manifestation of the continuity of scale-freedom in structuring of languages, along the refinement of cost functions representing the least effort principle: as long as the language develops closely to the optima of the prevailing cost function, power laws will be adaptively maintained.</p>
<p>In conclusion, our paper addresses the long-held conjecture that the principle of least effort provides a plausible mechanism for generating power laws. In deriving such a formalization, we interpret the effort in suitable information-theoretic terms and prove that its global minimum produces Zipf’s law. Our formalization enables a derivation of languages which are optimal with respect to both the communication inefficiency and direct signal cost. The proposed combination of these two factors within a generic cost function is an intuitive and powerful method to capture the trade-offs intrinsic to least-effort communication.</p>
</sec>
<sec id="sec005">
<title>4 Appendix</title>
<p><bold>Theorem 1.</bold> <italic>Each local minimizer of the function</italic> <disp-formula id="pone.0139475.e044">
<alternatives>
<graphic id="pone.0139475.e044g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e044"/>
<mml:math id="M44" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:mo>𝓒</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>→</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>ℝ</mml:mo>
<mml:mo>,</mml:mo>
<mml:mspace width="2.em"/>
<mml:mi>p</mml:mi>
<mml:mspace width="0.277778em"/>
<mml:mo>↦</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:msubsup>
<mml:mo>Ω</mml:mo>
<mml:mi>λ</mml:mi>
<mml:mi>c</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
</disp-formula> <italic>where</italic> <disp-formula id="pone.0139475.e045">
<alternatives>
<graphic id="pone.0139475.e045g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e045"/>
<mml:math id="M45" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:mo>𝓒</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>:</mml:mo>
<mml:mo>=</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>{</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>∈</mml:mo>
<mml:mo>𝓟</mml:mo>
<mml:mo>(</mml:mo>
<mml:mi>S</mml:mi>
<mml:mo>×</mml:mo>
<mml:mi>R</mml:mi>
<mml:mo>)</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>:</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mspace width="0.277778em"/>
<mml:mo>=</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:msub>
<mml:mo>∑</mml:mo>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>m</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mspace width="4.pt"/>
<mml:mtext mathvariant="italic">for</mml:mtext>
<mml:mspace width="4.pt"/>
<mml:mtext mathvariant="italic">all</mml:mtext>
<mml:mspace width="4.pt"/>
<mml:mi>j</mml:mi>
<mml:mo>}</mml:mo>
<mml:mspace width="4pt"/>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
</disp-formula> <italic>and</italic> <inline-formula id="pone.0139475.e046">
<alternatives>
<graphic id="pone.0139475.e046g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e046"/>
<mml:math id="M46" display="inline" overflow="scroll">
<mml:mrow>
<mml:msubsup>
<mml:mo>Ω</mml:mo>
<mml:mi>λ</mml:mi>
<mml:mi>c</mml:mi>
</mml:msubsup>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</alternatives>
</inline-formula> <italic>is specified by the</italic> <xref ref-type="disp-formula" rid="pone.0139475.e015">Eq (11)</xref>, 0 &lt; <italic>λ</italic> ≤ 1, <italic>can be represented as a function</italic> <italic>f</italic> : <italic>R</italic> → <italic>S</italic> <italic>such that</italic> <disp-formula id="pone.0139475.e047"><alternatives><graphic id="pone.0139475.e047g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e047"/><mml:math id="M47" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>=</mml:mo> <mml:mo>{</mml:mo> <mml:mtable><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mn>1</mml:mn> <mml:mo>/</mml:mo> <mml:mi>m</mml:mi></mml:mrow></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mspace width="4.pt"/><mml:mspace width="4.pt"/><mml:mrow><mml:mi>i</mml:mi><mml:mi>f</mml:mi></mml:mrow> <mml:mspace width="4.pt"/><mml:mspace width="4.pt"/><mml:mspace width="4.pt"/><mml:mrow><mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>=</mml:mo> <mml:mi>f</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mrow> <mml:mo>;</mml:mo></mml:mrow></mml:mtd></mml:mtr> <mml:mtr><mml:mtd columnalign="left"><mml:mn>0</mml:mn></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mspace width="4.pt"/><mml:mspace width="4.pt"/><mml:mrow><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable> <mml:mo/></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives> <label>(26)</label></disp-formula></p>
<p>In order to prove this theorem, we establish a few preliminary propositions (these results are obtained by Nihat Ay).</p>
<sec id="sec006">
<title>4.1 Extreme points</title>
<p>The extreme points of <inline-formula id="pone.0139475.e073"><alternatives><graphic id="pone.0139475.e073g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e073"/><mml:math id="M73" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> are specified by the following proposition.</p>
<p><bold>Proposition 2.</bold> <italic>The set <inline-formula id="pone.0139475.e074"><alternatives><graphic id="pone.0139475.e074g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e074"/><mml:math id="M74" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> has the extreme points</italic> <disp-formula id="pone.0139475.e048">
<alternatives>
<graphic id="pone.0139475.e048g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e048"/>
<mml:math id="M48" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:mtext>Ext</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mo>𝓒</mml:mo>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mspace width="0.277778em"/>
<mml:mo>=</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>{</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>∈</mml:mo>
<mml:mo>𝓟</mml:mo>
<mml:mo>(</mml:mo>
<mml:mi>S</mml:mi>
<mml:mo>×</mml:mo>
<mml:mi>R</mml:mi>
<mml:mo>)</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>:</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>m</mml:mi>
</mml:mfrac>
<mml:mspace width="0.166667em"/>
<mml:msub>
<mml:mi>δ</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mspace width="4.pt"/>
<mml:mo>}</mml:mo>
<mml:mspace width="4pt"/>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
</disp-formula>
</p>
<p><italic>where</italic> <italic>f</italic> <italic>is a function</italic> <italic>R</italic> → <italic>S</italic>.</p>
<p><italic><bold>Proof.</bold></italic> Consider the convex set
<disp-formula id="pone.0139475.e049"><alternatives><graphic id="pone.0139475.e049g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e049"/><mml:math id="M49" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mo>𝓣</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>=</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>{</mml:mo> <mml:mi>A</mml:mi> <mml:mo>=</mml:mo> <mml:msub><mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>|</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>∈</mml:mo> <mml:msup><mml:mrow><mml:mo>ℝ</mml:mo></mml:mrow> <mml:mrow><mml:mi>m</mml:mi> <mml:mo>·</mml:mo> <mml:mi>n</mml:mi></mml:mrow></mml:msup> <mml:mspace width="0.277778em"/><mml:mo>:</mml:mo> <mml:mspace width="0.277778em"/><mml:mrow><mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>|</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>≥</mml:mo> <mml:mn>0</mml:mn></mml:mrow> <mml:mspace width="4.pt"/><mml:mtext>for</mml:mtext> <mml:mspace width="4.pt"/><mml:mtext>all</mml:mtext> <mml:mspace width="4.pt"/><mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow> <mml:mtext>,</mml:mtext> <mml:mo/></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula> <disp-formula id="pone.0139475.e050"><alternatives><graphic id="pone.0139475.e050g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e050"/><mml:math id="M50" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mspace width="2.em"/><mml:mo/><mml:mspace width="4.pt"/><mml:mtext>and</mml:mtext> <mml:mspace width="4.pt"/><mml:mrow><mml:msub><mml:mo>∑</mml:mo> <mml:mi>i</mml:mi></mml:msub> <mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>|</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn></mml:mrow> <mml:mspace width="4.pt"/><mml:mtext>for</mml:mtext> <mml:mspace width="4.pt"/><mml:mtext>all</mml:mtext> <mml:mspace width="4.pt"/><mml:mi>j</mml:mi> <mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
of transition matrices. The extreme points of <inline-formula id="pone.0139475.e086"><alternatives><graphic id="pone.0139475.e086g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e086"/><mml:math id="M86" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> are given by functions <italic>f</italic> : <italic>j</italic> ↦ <italic>i</italic>. More precisely, each extreme point has the structure
<disp-formula id="pone.0139475.e051"><alternatives><graphic id="pone.0139475.e051g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e051"/><mml:math id="M51" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>|</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mspace width="0.277778em"/><mml:mo>=</mml:mo> <mml:mspace width="0.277778em"/><mml:msub><mml:mi>δ</mml:mi> <mml:mrow><mml:mi>f</mml:mi> <mml:mo>(</mml:mo> <mml:mi>j</mml:mi> <mml:mo>)</mml:mo></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>i</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
Now consider the map <italic>φ</italic> : <inline-formula id="pone.0139475.e075"><alternatives><graphic id="pone.0139475.e075g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e075"/><mml:math id="M75" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mo>→</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> that maps each matrix <italic>A</italic> = (<italic>a</italic><sub><italic>i</italic>∣<italic>j</italic></sub>)<sub><italic>i</italic>,<italic>j</italic></sub> to the probability vector
<disp-formula id="pone.0139475.e052"><alternatives><graphic id="pone.0139475.e052g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e052"/><mml:math id="M52" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="0.277778em"/><mml:mo>:</mml:mo> <mml:mo>=</mml:mo><mml:mspace width="0.277778em"/> <mml:mfrac><mml:mn>1</mml:mn> <mml:mi>m</mml:mi></mml:mfrac> <mml:mspace width="0.166667em"/><mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>|</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>,</mml:mo> <mml:mspace width="2.em"/><mml:mtext>for</mml:mtext> <mml:mspace width="4.pt"/><mml:mtext>all</mml:mtext> <mml:mspace width="4pt"/><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi> <mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
This map is bijective and satisfies <italic>φ</italic>((1 − <italic>t</italic>) <italic>A</italic> + <italic>t</italic> <italic>B</italic>) = (1 − <italic>t</italic>) <italic>φ</italic>(<italic>A</italic>) + <italic>t</italic> <italic>φ</italic>(<italic>B</italic>). Therefore, the extreme points of <inline-formula id="pone.0139475.e076"><alternatives><graphic id="pone.0139475.e076g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e076"/><mml:math id="M76" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> can be identified with the extreme points of <inline-formula id="pone.0139475.e077"><alternatives><graphic id="pone.0139475.e077g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e077"/><mml:math id="M77" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>.</p>
</sec>
<sec id="sec007">
<title>4.2 Concavity</title>
<p>Consider the set <italic>S</italic> = {<italic>s</italic><sub>1</sub>, …, <italic>s</italic><sub><italic>n</italic></sub>} of signals with <italic>n</italic> elements and the set <italic>R</italic> = {<italic>r</italic><sub>1</sub>, …, <italic>r</italic><sub><italic>m</italic></sub>} of <italic>m</italic> objects, and denote with <inline-formula id="pone.0139475.e088"><alternatives><graphic id="pone.0139475.e088g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e088"/><mml:math id="M88" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">P</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>(<italic>S</italic> × <italic>R</italic>) the set of all probability vectors <italic>p</italic>(<italic>s</italic><sub><italic>i</italic></sub>, <italic>r</italic><sub><italic>j</italic></sub>), 1 ≤ <italic>i</italic> ≤ <italic>n</italic>, 1 ≤ <italic>j</italic> ≤ <italic>m</italic>. We define the following functions on <inline-formula id="pone.0139475.e089"><alternatives><graphic id="pone.0139475.e089g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e089"/><mml:math id="M89" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">P</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>(<italic>S</italic> × <italic>R</italic>):
<disp-formula id="pone.0139475.e053"><alternatives><graphic id="pone.0139475.e053g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e053"/><mml:math id="M53" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>S</mml:mi> <mml:mo>|</mml:mo> <mml:mi>R</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="0.277778em"/><mml:mo>:</mml:mo> <mml:mo>=</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mi>j</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:munder><mml:mo>∑</mml:mo> <mml:mi>i</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub><mml:mspace width="0.277778em"/> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
and
<disp-formula id="pone.0139475.e054"><alternatives><graphic id="pone.0139475.e054g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e054"/><mml:math id="M54" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo> <mml:mi>S</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="0.277778em"/><mml:mo>:</mml:mo> <mml:mo>=</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mi>i</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:munder><mml:mo>∑</mml:mo> <mml:mi>j</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>m</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula></p>
<p><bold>Proposition 3.</bold> <italic>All three functions</italic> <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub>, <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub>, <italic>and</italic> <disp-formula id="pone.0139475.e055">
<alternatives>
<graphic id="pone.0139475.e055g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e055"/>
<mml:math id="M55" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:mrow>
<mml:mo>〈</mml:mo>
<mml:mi>c</mml:mi>
<mml:mo>〉</mml:mo>
</mml:mrow>
<mml:mspace width="0.277778em"/>
<mml:mo>:</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mspace width="0.277778em"/>
<mml:mi>p</mml:mi>
<mml:mo>↦</mml:mo>
<mml:munder>
<mml:mo>∑</mml:mo>
<mml:mi>i</mml:mi>
</mml:munder>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mspace width="0.166667em"/>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
</disp-formula> <italic>that are involved in the definition of</italic> <inline-formula id="pone.0139475.e056"><alternatives><graphic id="pone.0139475.e056g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e056"/><mml:math id="M56" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> <italic>are concave in</italic> <italic>p</italic>. <italic>Furthermore, the restriction of</italic> <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> <italic>to the set <inline-formula id="pone.0139475.e078"><alternatives><graphic id="pone.0139475.e078g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e078"/><mml:math id="M78" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> is strictly concave.</italic></p>
<p><italic><bold>Proof.</bold></italic> The statements follow from well-known convexity properties of the entropy and the relative entropy.</p>
<p><bold>(1)</bold> <italic>Concavity of <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub></italic>: We rewrite the function <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> as
<disp-formula id="pone.0139475.e057"><alternatives><graphic id="pone.0139475.e057g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e057"/><mml:math id="M57" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo> <mml:mi>S</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd> <mml:mtd><mml:mo>=</mml:mo></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mi>i</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:munder><mml:mo>∑</mml:mo> <mml:mi>j</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>m</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr> <mml:mtr><mml:mtd/><mml:mtd><mml:mo>=</mml:mo></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>m</mml:mi></mml:msub> <mml:mfrac><mml:mrow><mml:mi>p</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mrow><mml:msub><mml:mo>∑</mml:mo> <mml:mi>j</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr> <mml:mtr><mml:mtd/><mml:mtd><mml:mo>=</mml:mo></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>m</mml:mi></mml:msub> <mml:mfrac><mml:mrow><mml:mi>p</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mrow><mml:mi>m</mml:mi> <mml:mspace width="0.166667em"/><mml:mfrac><mml:mn>1</mml:mn> <mml:mi>m</mml:mi></mml:mfrac> <mml:msub><mml:mo>∑</mml:mo> <mml:mi>j</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr> <mml:mtr><mml:mtd/><mml:mtd><mml:mo>=</mml:mo></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>m</mml:mi></mml:msub> <mml:mfrac><mml:mrow><mml:mi>p</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mrow><mml:mfrac><mml:mn>1</mml:mn> <mml:mi>m</mml:mi></mml:mfrac> <mml:msub><mml:mo>∑</mml:mo> <mml:mi>j</mml:mi></mml:msub> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac> <mml:mspace width="0.277778em"/><mml:mo>+</mml:mo> <mml:mspace width="0.277778em"/><mml:mn>1</mml:mn> <mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
The concavity of <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> now follows from the joint convexity of the relative entropy <inline-formula id="pone.0139475.e058"><alternatives><graphic id="pone.0139475.e058g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e058"/><mml:math id="M58" display="inline" overflow="scroll"><mml:mrow><mml:mo stretchy="false">(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>,</mml:mo> <mml:mi>q</mml:mi> <mml:mo stretchy="false">)</mml:mo> <mml:mo>↦</mml:mo> <mml:mi>D</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:mi>p</mml:mi> <mml:mo stretchy="false">‖</mml:mo> <mml:mi>q</mml:mi> <mml:mo stretchy="false">)</mml:mo> <mml:mo>=</mml:mo> <mml:msub><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo><mml:mspace width="0.277778em"/> <mml:msub><mml:mtext>log</mml:mtext> <mml:mi>m</mml:mi></mml:msub> <mml:mfrac><mml:mrow><mml:mi>p</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo></mml:mrow> <mml:mrow><mml:mi>q</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></alternatives></inline-formula>.</p>
<p><bold>(2)</bold> <italic>Concavity of</italic> <italic>H<sub>S∣R</sub></italic>: The concavity of <italic>H</italic><sub><italic>S</italic>∣<italic>R</italic></sub> follows by the same arguments as in (1). We now prove the strict concavity of its restriction to <inline-formula id="pone.0139475.e079"><alternatives><graphic id="pone.0139475.e079g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e079"/><mml:math id="M79" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>.
<disp-formula id="pone.0139475.e059"><alternatives><graphic id="pone.0139475.e059g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e059"/><mml:math id="M59" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>S</mml:mi> <mml:mo>|</mml:mo> <mml:mi>R</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd> <mml:mtd><mml:mo>=</mml:mo></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mi>j</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:munder><mml:mo>∑</mml:mo> <mml:mi>i</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub><mml:mspace width="0.277778em"/> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>|</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr> <mml:mtr><mml:mtd/><mml:mtd><mml:mo>=</mml:mo></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub><mml:mspace width="0.277778em"/> <mml:mfrac><mml:mrow><mml:mi>p</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mrow><mml:mi>p</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr> <mml:mtr><mml:mtd/><mml:mtd><mml:mo>=</mml:mo></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub><mml:mspace width="0.277778em"/> <mml:mfrac><mml:mrow><mml:mi>p</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mfrac><mml:mn>1</mml:mn> <mml:mi>m</mml:mi></mml:mfrac></mml:mfrac></mml:mrow></mml:mtd></mml:mtr> <mml:mtr><mml:mtd/><mml:mtd><mml:mo>=</mml:mo></mml:mtd> <mml:mtd columnalign="left"><mml:mrow><mml:mo>-</mml:mo> <mml:munder><mml:mo>∑</mml:mo> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub><mml:mspace width="0.277778em"/> <mml:mrow><mml:mi>p</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>,</mml:mo> <mml:msub><mml:mi>r</mml:mi> <mml:mi>j</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="0.277778em"/><mml:mo>-</mml:mo> <mml:mspace width="0.277778em"/><mml:msub><mml:mo form="prefix">log</mml:mo> <mml:mi>n</mml:mi></mml:msub><mml:mspace width="0.277778em"/> <mml:mi>m</mml:mi> <mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
The strict concavity of <italic>H</italic><sub><italic>R</italic>∣<italic>S</italic></sub> now follows from the strict concavity of the Shannon entropy.</p>
<p><bold>(2)</bold> <italic>Concavity of 〈<italic>c</italic>〉</italic>: This simply follows from the fact that 〈<italic>c</italic>〉 is an affine function and therefore concave and convex at the same time.</p>
<p>With a number 0 &lt; <italic>λ</italic> ≤ 1, we now consider the function
<disp-formula id="pone.0139475.e060"><alternatives><graphic id="pone.0139475.e060g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e060"/><mml:math id="M60" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="0.277778em"/><mml:mo>=</mml:mo> <mml:mspace width="0.277778em"/><mml:mi>λ</mml:mi> <mml:mo>(</mml:mo> <mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>R</mml:mi> <mml:mo>|</mml:mo> <mml:mi>S</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>+</mml:mo> <mml:msub><mml:mi>H</mml:mi> <mml:mrow><mml:mi>S</mml:mi> <mml:mo>|</mml:mo> <mml:mi>R</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:mo>(</mml:mo> <mml:mi>p</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>)</mml:mo> <mml:mo>+</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:mn>1</mml:mn> <mml:mo>-</mml:mo> <mml:mi>λ</mml:mi> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="0.166667em"/><mml:munder><mml:mo>∑</mml:mo> <mml:mi>i</mml:mi></mml:munder> <mml:mi>p</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="0.166667em"/><mml:mi>c</mml:mi> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>s</mml:mi> <mml:mi>i</mml:mi></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
From Proposition 3, it immediately follows that <inline-formula id="pone.0139475.e061"><alternatives><graphic id="pone.0139475.e061g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e061"/><mml:math id="M61" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> also has corresponding concavity properties.</p>
<p><bold>Corollary 4.</bold> <italic>For</italic> 0 ≤ <italic>λ</italic> ≤ 1, <italic>the function</italic> <inline-formula id="pone.0139475.e062"><alternatives><graphic id="pone.0139475.e062g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e062"/><mml:math id="M62" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> <italic>is concave in</italic> <italic>p</italic>, <italic>and, if</italic> <italic>λ</italic> &gt; 0, <italic>its restriction to the convex set</italic> <inline-formula id="pone.0139475.e080"><alternatives><graphic id="pone.0139475.e080g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e080"/><mml:math id="M80" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> <italic>is strictly concave.</italic></p>
</sec>
<sec id="sec008">
<title>4.3 Minimizers</title>
<p>We have the following direct implication of Corollary 4.</p>
<p><bold>Corollary 5.</bold> <italic>Let</italic> 0 &lt; <italic>λ</italic> ≤ 1 <italic>and let</italic> <italic>p</italic> <italic>be a local minimizer of the map</italic> <disp-formula id="pone.0139475.e063">
<alternatives>
<graphic id="pone.0139475.e063g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e063"/>
<mml:math id="M63" display="block" overflow="scroll">
<mml:mtable displaystyle="true">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mrow>
<mml:mo>𝓒</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>→</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:mo>ℝ</mml:mo>
<mml:mo>,</mml:mo>
<mml:mspace width="2.em"/>
<mml:mi>p</mml:mi>
<mml:mspace width="0.277778em"/>
<mml:mo>↦</mml:mo>
<mml:mspace width="0.277778em"/>
<mml:msubsup>
<mml:mo>Ω</mml:mo>
<mml:mi>λ</mml:mi>
<mml:mi>c</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</alternatives>
</disp-formula> <italic>Then</italic> <italic>p</italic> <italic>is an extreme point of</italic> <inline-formula id="pone.0139475.e081"><alternatives><graphic id="pone.0139475.e081g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e081"/><mml:math id="M81" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>.</p>
<p><italic><bold>Proof.</bold></italic> This directly follows from the strict concavity of this function.</p>
<p>Together with Proposition 2, this implies Theorem 1, our main result on minimizers of the restriction of <inline-formula id="pone.0139475.e064"><alternatives><graphic id="pone.0139475.e064g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e064"/><mml:math id="M64" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> to the convex set <inline-formula id="pone.0139475.e082"><alternatives><graphic id="pone.0139475.e082g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e082"/><mml:math id="M82" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>.</p>
<p>We finish this analysis by addressing the problem of minimizing <inline-formula id="pone.0139475.e065"><alternatives><graphic id="pone.0139475.e065g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e065"/><mml:math id="M65" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> on a discrete set. In order to do so, consider the set of 0/1-matrices that have at least one “1”-entry in each column:
<disp-formula id="pone.0139475.e066"><alternatives><graphic id="pone.0139475.e066g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e066"/><mml:math id="M66" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mo>𝓢</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>:</mml:mo> <mml:mo>=</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>{</mml:mo> <mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mo>∈</mml:mo> <mml:msup><mml:mrow><mml:mo>{</mml:mo> <mml:mn>0</mml:mn> <mml:mo>,</mml:mo> <mml:mn>1</mml:mn> <mml:mo>}</mml:mo></mml:mrow> <mml:mrow><mml:mi>n</mml:mi> <mml:mo>·</mml:mo> <mml:mi>m</mml:mi></mml:mrow></mml:msup> <mml:mspace width="0.277778em"/><mml:mo>:</mml:mo> <mml:mspace width="0.277778em"/><mml:munder><mml:mo>∑</mml:mo> <mml:mi>i</mml:mi></mml:munder> <mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>≥</mml:mo> <mml:mn>1</mml:mn> <mml:mspace width="0.277778em"/><mml:mtext>for</mml:mtext> <mml:mspace width="4.pt"/><mml:mtext>all</mml:mtext> <mml:mspace width="4.pt"/><mml:mi>j</mml:mi> <mml:mo>}</mml:mo> <mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
This set can naturally be embedded into the set <inline-formula id="pone.0139475.e087"><alternatives><graphic id="pone.0139475.e087g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e087"/><mml:math id="M87" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow></mml:math></alternatives></inline-formula>, which we have considered in the proof of Proposition 2:
<disp-formula id="pone.0139475.e067"><alternatives><graphic id="pone.0139475.e067g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e067"/><mml:math id="M67" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mo mathvariant="italic">ı</mml:mo> <mml:mo>:</mml:mo> <mml:mspace width="0.277778em"/><mml:mspace width="0.277778em"/><mml:mo>𝓢</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>↪</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>𝓣</mml:mo> <mml:mo>,</mml:mo> <mml:mspace width="1.em"/><mml:msub><mml:mrow><mml:mo>(</mml:mo> <mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mo>)</mml:mo></mml:mrow> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mspace width="0.277778em"/><mml:mo>↦</mml:mo> <mml:mspace width="0.277778em"/><mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>|</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mspace width="0.277778em"/><mml:mo>:</mml:mo> <mml:mo>=</mml:mo><mml:mspace width="0.277778em"/> <mml:mfrac><mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub> <mml:mrow><mml:msub><mml:mo>∑</mml:mo> <mml:mi>i</mml:mi></mml:msub><mml:mspace width="0.277778em"/> <mml:msub><mml:mi>a</mml:mi> <mml:mrow><mml:mi>i</mml:mi> <mml:mo>,</mml:mo> <mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac> <mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
Together with the map <italic>φ</italic> : <inline-formula id="pone.0139475.e083"><alternatives><graphic id="pone.0139475.e083g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e083"/><mml:math id="M83" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mo>→</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> we have the injective composition <italic>φ</italic> ∘ <italic>ı</italic>. From Proposition 2 it follows that the extreme points of <inline-formula id="pone.0139475.e084"><alternatives><graphic id="pone.0139475.e084g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e084"/><mml:math id="M84" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> are in the image of <italic>φ</italic> ∘ <italic>ı</italic>. Furthermore, Corollary 5 implies that all local, and therefore also all global, minimizers of <inline-formula id="pone.0139475.e068"><alternatives><graphic id="pone.0139475.e068g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e068"/><mml:math id="M68" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> are in the image of <italic>φ</italic> ∘ <italic>ı</italic>. The previous work of Ferrer i Cancho and Sole [<xref ref-type="bibr" rid="pone.0139475.ref012">12</xref>] refers to the minimization of a function on the discrete set 𝓢:
<disp-formula id="pone.0139475.e069"><alternatives><graphic id="pone.0139475.e069g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e069"/><mml:math id="M69" display="block" overflow="scroll"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mo>Ω</mml:mo> <mml:mo>˜</mml:mo></mml:mover> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup> <mml:mspace width="0.277778em"/><mml:mo>:</mml:mo> <mml:mo>=</mml:mo><mml:mspace width="0.277778em"/> <mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup> <mml:mo>∘</mml:mo> <mml:mi>φ</mml:mi> <mml:mo>∘</mml:mo> <mml:mo mathvariant="italic">ı</mml:mo> <mml:mo>:</mml:mo> <mml:mspace width="0.277778em"/><mml:mspace width="0.277778em"/><mml:mo>𝓢</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>→</mml:mo> <mml:mspace width="0.277778em"/><mml:mo>ℝ</mml:mo> <mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></alternatives></disp-formula>
It is not obvious how to relate local minimizers of this function, with an appropriate notion of locality in 𝓢, to local minimizers of <inline-formula id="pone.0139475.e070"><alternatives><graphic id="pone.0139475.e070g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e070"/><mml:math id="M70" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula>. However, we have the following obvious relation between global minimizers.</p>
<p><bold>Corollary 6.</bold> <italic>A point</italic> <italic>p</italic> ∈ <inline-formula id="pone.0139475.e085"><alternatives><graphic id="pone.0139475.e085g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e085"/><mml:math id="M85" display="inline" overflow="scroll"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:mrow></mml:math></alternatives></inline-formula> <italic>is a global minimizer of</italic> <inline-formula id="pone.0139475.e071"><alternatives><graphic id="pone.0139475.e071g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e071"/><mml:math id="M71" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mo>Ω</mml:mo> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula> <italic>if and only if it is in the image of</italic> <italic>φ</italic> ∘ <italic>ı</italic> <italic>and</italic> (<italic>φ</italic> ∘ <italic>ı</italic>)<sup>−1</sup>(<italic>p</italic>) <italic>globally minimizes</italic> <inline-formula id="pone.0139475.e072"><alternatives><graphic id="pone.0139475.e072g" mimetype="image" xlink:type="simple" position="anchor" xlink:href="info:doi/10.1371/journal.pone.0139475.e072"/><mml:math id="M72" display="inline" overflow="scroll"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mo>Ω</mml:mo> <mml:mo>˜</mml:mo></mml:mover> <mml:mi>λ</mml:mi> <mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:math></alternatives></inline-formula>.</p>
</sec>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="pone.0139475.ref001">
<label>1</label>
<mixed-citation xlink:type="simple" publication-type="other">Zipf GK (1949) Human behavior and the principle of least effort.</mixed-citation>
</ref>
<ref id="pone.0139475.ref002">
<label>2</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Balasubrahmanyan</surname> <given-names>V</given-names></name>, <name name-style="western"><surname>Naranan</surname> <given-names>S</given-names></name> (<year>1996</year>) <article-title>Quantitative linguistics and complex system studies*</article-title>. <source>Journal of Quantitative Linguistics</source> <volume>3</volume>: <fpage>177</fpage>–<lpage>228</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1080/09296179608599629" xlink:type="simple">10.1080/09296179608599629</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref003">
<label>3</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Ferrer i Cancho</surname> <given-names>R</given-names></name> (<year>2005</year>) <article-title>The variation of Zipf’s law in human language</article-title>. <source>The European Physical Journal B-Condensed Matter and Complex Systems</source> <volume>44</volume>: <fpage>249</fpage>–<lpage>257</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1140/epjb/e2005-00121-8" xlink:type="simple">10.1140/epjb/e2005-00121-8</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref004">
<label>4</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Prokopenko</surname> <given-names>M</given-names></name>, <name name-style="western"><surname>Ay</surname> <given-names>N</given-names></name>, <name name-style="western"><surname>Obst</surname> <given-names>O</given-names></name>, <name name-style="western"><surname>Polani</surname> <given-names>D</given-names></name> (<year>2010</year>) <article-title>Phase transitions in least-effort communications</article-title>. <source>Journal of Statistical Mechanics: Theory and Experiment</source> <volume>2010</volume>: <fpage>P11025</fpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1088/1742-5468/2010/11/P11025" xlink:type="simple">10.1088/1742-5468/2010/11/P11025</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref005">
<label>5</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Li</surname> <given-names>W</given-names></name> (<year>1992</year>) <article-title>Random texts exhibit Zipf’s-law-like word frequency distribution</article-title>. <source>Information Theory, IEEE Transactions on</source> <volume>38</volume>: <fpage>1842</fpage>–<lpage>1845</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1109/18.165464" xlink:type="simple">10.1109/18.165464</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref006">
<label>6</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Miller</surname> <given-names>GA</given-names></name> (<year>1957</year>) <article-title>Some effects of intermittent silence</article-title>. <source>The American Journal of Psychology</source> <volume>70</volume>: <fpage>311</fpage>–<lpage>314</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.2307/1419346" xlink:type="simple">10.2307/1419346</ext-link></comment> <object-id pub-id-type="pmid">13424784</object-id></mixed-citation>
</ref>
<ref id="pone.0139475.ref007">
<label>7</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Suzuki</surname> <given-names>R</given-names></name>, <name name-style="western"><surname>Buck</surname> <given-names>JR</given-names></name>, <name name-style="western"><surname>Tyack</surname> <given-names>PL</given-names></name> (<year>2005</year>) <article-title>The use of Zipf’s law in animal communication analysis</article-title>. <source>Animal Behaviour</source> <volume>69</volume>: <fpage>F9</fpage>–<lpage>F17</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.anbehav.2004.08.004" xlink:type="simple">10.1016/j.anbehav.2004.08.004</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref008">
<label>8</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Ferrer i Cancho</surname> <given-names>R</given-names></name>, <name name-style="western"><surname>Elvevåg</surname> <given-names>B</given-names></name> (<year>2010</year>) <article-title>Random texts do not exhibit the real Zipf’s law-like rank distribution</article-title>. <source>PLoS One</source> <volume>5</volume>: <fpage>e9411</fpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1371/journal.pone.0009411" xlink:type="simple">10.1371/journal.pone.0009411</ext-link></comment> <object-id pub-id-type="pmid">20231884</object-id></mixed-citation>
</ref>
<ref id="pone.0139475.ref009">
<label>9</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Mandelbrot</surname> <given-names>B</given-names></name> (<year>1953</year>) <article-title>An informational theory of the statistical structure of language</article-title>. <source>Communication theory</source> <volume>84</volume>: <fpage>486</fpage>–<lpage>502</lpage>.</mixed-citation>
</ref>
<ref id="pone.0139475.ref010">
<label>10</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Visser</surname> <given-names>M</given-names></name> (<year>2013</year>) <article-title>Zipf’s law, power laws and maximum entropy</article-title>. <source>New Journal of Physics</source> <volume>15</volume>: <fpage>043021</fpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1088/1367-2630/15/4/043021" xlink:type="simple">10.1088/1367-2630/15/4/043021</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref011">
<label>11</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Shannon</surname> <given-names>CE</given-names></name> (<year>1948</year>) <article-title>A Mathematical Theory of Communication</article-title>. <source>Bell System Technical Journal</source> <volume>27</volume>: <fpage>379</fpage>–<lpage>423</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1002/j.1538-7305.1948.tb01338.x" xlink:type="simple">10.1002/j.1538-7305.1948.tb01338.x</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref012">
<label>12</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Ferrer i Cancho</surname> <given-names>R</given-names></name>, <name name-style="western"><surname>Solé</surname> <given-names>RV</given-names></name> (<year>2003</year>) <article-title>Least effort and the origins of scaling in human language</article-title>. <source>Proceedings of the National Academy of Sciences</source> <volume>100</volume>: <fpage>788</fpage>–<lpage>791</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1073/pnas.0335980100" xlink:type="simple">10.1073/pnas.0335980100</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref013">
<label>13</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Clauset</surname> <given-names>A</given-names></name>, <name name-style="western"><surname>Shalizi</surname> <given-names>CR</given-names></name>, <name name-style="western"><surname>Newman</surname> <given-names>ME</given-names></name> (<year>2009</year>) <article-title>Power-law distributions in empirical data</article-title>. <source>SIAM review</source> <volume>51</volume>: <fpage>661</fpage>–<lpage>703</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1137/070710111" xlink:type="simple">10.1137/070710111</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref014">
<label>14</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Baek</surname> <given-names>SK</given-names></name>, <name name-style="western"><surname>Bernhardsson</surname> <given-names>S</given-names></name>, <name name-style="western"><surname>Minnhagen</surname> <given-names>P</given-names></name> (<year>2011</year>) <article-title>Zipf’s law unzipped</article-title>. <source>New Journal of Physics</source> <volume>13</volume>: <fpage>043004</fpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1088/1367-2630/13/4/043004" xlink:type="simple">10.1088/1367-2630/13/4/043004</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref015">
<label>15</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Ferrer i Cancho</surname> <given-names>R</given-names></name> (<year>2005</year>) <article-title>Decoding least effort and scaling in signal frequency distributions</article-title>. <source>Physica A: Statistical Mechanics and its Applications</source> <volume>345</volume>: <fpage>275</fpage>–<lpage>284</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1016/j.physa.2004.06.158" xlink:type="simple">10.1016/j.physa.2004.06.158</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref016">
<label>16</label>
<mixed-citation xlink:type="simple" publication-type="other">Ferrer i Cancho R (2014) Optimization models of natural communication. arXiv preprint arXiv:14122486.</mixed-citation>
</ref>
<ref id="pone.0139475.ref017">
<label>17</label>
<mixed-citation xlink:type="simple" publication-type="other">Obst O, Polani D, Prokopenko M (2011) Origins of scaling in genetic code. In: Advances in Artificial Life. Darwin Meets von Neumann. Proceedings, 10th European Conference, ECAL 2009, Budapest, Hungary, September 13-16, 2009, Springer. pp. 85–93.</mixed-citation>
</ref>
<ref id="pone.0139475.ref018">
<label>18</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Ferrer i Cancho</surname> <given-names>R</given-names></name> (<year>2005</year>) <article-title>Zipf’s law from a communicative phase transition</article-title>. <source>The European Physical Journal B-Condensed Matter and Complex Systems</source> <volume>47</volume>: <fpage>449</fpage>–<lpage>457</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1140/epjb/e2005-00340-y" xlink:type="simple">10.1140/epjb/e2005-00340-y</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref019">
<label>19</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Ferrer i Cancho</surname> <given-names>R</given-names></name>, <name name-style="western"><surname>Díaz-Guilera</surname> <given-names>A</given-names></name> (<year>2007</year>) <article-title>The global minima of the communicative energy of natural communication systems</article-title>. <source>Journal of Statistical Mechanics: Theory and Experiment</source> <volume>2007</volume>: <fpage>P06009</fpage>.</mixed-citation>
</ref>
<ref id="pone.0139475.ref020">
<label>20</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Rokhlin</surname> <given-names>VA</given-names></name> (<year>1967</year>) <article-title>Lectures on the entropy theory of measure-preserving transformations</article-title>. <source>Russian Mathematical Surveys</source> <volume>22</volume>: <fpage>1</fpage>–<lpage>52</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1070/RM1967v022n05ABEH001224" xlink:type="simple">10.1070/RM1967v022n05ABEH001224</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref021">
<label>21</label>
<mixed-citation xlink:type="simple" publication-type="book">
<name name-style="western"><surname>Crutchfield</surname> <given-names>JP</given-names></name> (<year>1990</year>) <chapter-title>Information and its Metric</chapter-title>. In: <name name-style="western"><surname>Lam</surname> <given-names>L</given-names></name>, <name name-style="western"><surname>Morris</surname> <given-names>HC</given-names></name>, editors, <source>Nonlinear Structures in Physical Systems—Pattern Formation, Chaos and Waves</source>, <publisher-name>Springer Verlag</publisher-name>. pp. <fpage>119</fpage>–<lpage>130</lpage>.</mixed-citation>
</ref>
<ref id="pone.0139475.ref022">
<label>22</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Prokopenko</surname> <given-names>M</given-names></name>, <name name-style="western"><surname>Polani</surname> <given-names>D</given-names></name>, <name name-style="western"><surname>Chadwick</surname> <given-names>M</given-names></name> (<year>2009</year>) <article-title>Stigmergic gene transfer and emergence of universal coding</article-title>. <source>HFSP Journal</source> <volume>3</volume>: <fpage>317</fpage>–<lpage>327</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.2976/1.3175813" xlink:type="simple">10.2976/1.3175813</ext-link></comment> <object-id pub-id-type="pmid">20357889</object-id></mixed-citation>
</ref>
<ref id="pone.0139475.ref023">
<label>23</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Niven</surname> <given-names>RK</given-names></name> (<year>2009</year>) <article-title>Steady state of a dissipative flow-controlled system and the maximum entropy production principle</article-title>. <source>Physical Review E</source> <volume>80</volume>: <fpage>021113</fpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1103/PhysRevE.80.021113" xlink:type="simple">10.1103/PhysRevE.80.021113</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref024">
<label>24</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Niven</surname> <given-names>RK</given-names></name> (<year>2010</year>) <article-title>Minimization of a free-energy-like potential for non-equilibrium flow systems at steady state</article-title>. <source>Philosophical Transactions of the Royal Society B: Biological Sciences</source> <volume>365</volume>: <fpage>1323</fpage>–<lpage>1331</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1098/rstb.2009.0296" xlink:type="simple">10.1098/rstb.2009.0296</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref025">
<label>25</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Prokopenko</surname> <given-names>M</given-names></name>, <name name-style="western"><surname>Lizier</surname> <given-names>JT</given-names></name>, <name name-style="western"><surname>Obst</surname> <given-names>O</given-names></name>, <name name-style="western"><surname>Wang</surname> <given-names>XR</given-names></name> (<year>2011</year>) <article-title>Relating Fisher information to order parameters</article-title>. <source>Physical Review E</source> <volume>84</volume>: <fpage>041116</fpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1103/PhysRevE.84.041116" xlink:type="simple">10.1103/PhysRevE.84.041116</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref026">
<label>26</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Dickman</surname> <given-names>R</given-names></name>, <name name-style="western"><surname>Moloney</surname> <given-names>NR</given-names></name>, <name name-style="western"><surname>Altmann</surname> <given-names>EG</given-names></name> (<year>2012</year>) <article-title>Analysis of an information-theoretic model for communication</article-title>. <source>Journal of Statistical Mechanics: Theory and Experiment</source> <volume>2012</volume>: <fpage>P12022</fpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1088/1742-5468/2012/12/P12022" xlink:type="simple">10.1088/1742-5468/2012/12/P12022</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref027">
<label>27</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Nowak</surname> <given-names>MA</given-names></name>, <name name-style="western"><surname>Krakauer</surname> <given-names>DC</given-names></name> (<year>1999</year>) <article-title>The evolution of language</article-title>. <source>Proceedings of the National Academy of Sciences</source> <volume>96</volume>: <fpage>8028</fpage>–<lpage>8033</lpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1073/pnas.96.14.8028" xlink:type="simple">10.1073/pnas.96.14.8028</ext-link></comment></mixed-citation>
</ref>
<ref id="pone.0139475.ref028">
<label>28</label>
<mixed-citation xlink:type="simple" publication-type="other">Ferrer i Cancho R (2013) The optimality of attaching unlinked labels to unlinked meanings. arXiv preprint arXiv:13105884.</mixed-citation>
</ref>
<ref id="pone.0139475.ref029">
<label>29</label>
<mixed-citation xlink:type="simple" publication-type="other">Saxton M (2010) Child language: Acquisition and development. Sage.</mixed-citation>
</ref>
<ref id="pone.0139475.ref030">
<label>30</label>
<mixed-citation xlink:type="simple" publication-type="book">
<name name-style="western"><surname>Clark</surname> <given-names>EV</given-names></name> (<year>1995</year>) <source>The lexicon in acquisition, volume 65 of <italic>Cambridge Studies in Linguistics</italic></source>. <publisher-name>Cambridge University Press</publisher-name>.</mixed-citation>
</ref>
<ref id="pone.0139475.ref031">
<label>31</label>
<mixed-citation xlink:type="simple" publication-type="journal">
<name name-style="western"><surname>Baixeries</surname> <given-names>J</given-names></name>, <name name-style="western"><surname>Elvevåg</surname> <given-names>B</given-names></name>, <name name-style="western"><surname>Ferrer i Cancho</surname> <given-names>R</given-names></name> (<year>2013</year>) <article-title>The evolution of the exponent of zipf’s law in language ontogeny</article-title>. <source>PloS ONE</source> <volume>8</volume>: <fpage>e53227</fpage>. <comment>doi: <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1371/journal.pone.0053227" xlink:type="simple">10.1371/journal.pone.0053227</ext-link></comment> <object-id pub-id-type="pmid">23516390</object-id></mixed-citation>
</ref>
</ref-list>
</back>
</article>