<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">PLoS ONE</journal-id>
<journal-id journal-id-type="publisher-id">plos</journal-id>
<journal-id journal-id-type="pmc">plosone</journal-id><journal-title-group>
<journal-title>PLoS ONE</journal-title></journal-title-group>
<issn pub-type="epub">1932-6203</issn>
<publisher>
<publisher-name>Public Library of Science</publisher-name>
<publisher-loc>San Francisco, USA</publisher-loc></publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">PONE-D-13-52093</article-id>
<article-id pub-id-type="doi">10.1371/journal.pone.0098914</article-id>
<article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer and information sciences</subject><subj-group><subject>Information technology</subject></subj-group></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Social sciences</subject><subj-group><subject>Sociology</subject><subj-group><subject>Social systems</subject></subj-group></subj-group></subj-group></article-categories>
<title-group>
<article-title>Leveraging Position Bias to Improve Peer Recommendation</article-title>
<alt-title alt-title-type="running-head">Leveraging Position Bias to Improve Peer Recommendation</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Lerman</surname><given-names>Kristina</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Hogg</surname><given-names>Tad</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib>
</contrib-group>
<aff id="aff1"><label>1</label><addr-line>USC Information Sciences Institute, Marina Del Rey, California, United States of America</addr-line></aff>
<aff id="aff2"><label>2</label><addr-line>Institute for Molecular Manufacturing, Palo Alto, California, United States of America</addr-line></aff>
<contrib-group>
<contrib contrib-type="editor" xlink:type="simple"><name name-style="western"><surname>Suleman</surname><given-names>Hussein</given-names></name>
<role>Editor</role>
<xref ref-type="aff" rid="edit1"/></contrib>
</contrib-group>
<aff id="edit1"><addr-line>University of Cape Town, South Africa</addr-line></aff>
<author-notes>
<corresp id="cor1">* E-mail: <email xlink:type="simple">lerman@isi.edu</email></corresp>
<fn fn-type="conflict"><p>The authors have declared that no competing interests exist.</p></fn>
<fn fn-type="con"><p>Conceived and designed the experiments: KL TH. Performed the experiments: KL. Analyzed the data: TH. Wrote the paper: KL TH.</p></fn>
</author-notes>
<pub-date pub-type="collection"><year>2014</year></pub-date>
<pub-date pub-type="epub"><day>11</day><month>6</month><year>2014</year></pub-date>
<volume>9</volume>
<issue>6</issue>
<elocation-id>e98914</elocation-id>
<history>
<date date-type="received"><day>10</day><month>12</month><year>2013</year></date>
<date date-type="accepted"><day>8</day><month>5</month><year>2014</year></date>
</history>
<permissions>
<copyright-year>2014</copyright-year>
<copyright-holder>Lerman, Hogg</copyright-holder><license xlink:type="simple"><license-p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/" xlink:type="simple">Creative Commons Attribution License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</license-p></license></permissions>
<abstract>
<p>With the advent of social media and peer production, the amount of new online content has grown dramatically. To identify interesting items in the vast stream of new content, providers must rely on peer recommendation to aggregate opinions of their many users. Due to human cognitive biases, the presentation order strongly affects how people allocate attention to the available content. Moreover, we can manipulate attention through the presentation order of items to change the way peer recommendation works. We experimentally evaluate this effect using Amazon Mechanical Turk. We find that different policies for ordering content can steer user attention so as to improve the outcomes of peer recommendation.</p>
</abstract>
<funding-group><funding-statement>This work was supported in part by AFOSR (contract FA9550-10-1-0569) and by DARPA (contract W911NF-12-1-0034). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</funding-statement></funding-group><counts><page-count count="8"/></counts></article-meta>
</front>
<body><sec id="s1">
<title>Introduction</title>
<p>The growing volume of content created in online social media and other peer production systems is making it increasingly difficult to identify interesting items. On YouTube alone, over 100 hours of video are uploaded every minute. Which of the many videos are worth watching? Likewise, which of the thousands of new daily articles and comments on the social news web site Reddit are worth reading?</p>
<p>The challenge facing content providers, such as YouTube and Reddit, is identifying items their user communities will find interesting from among the vast numbers of newly created items. If a better item comes along, content providers need to identify it in a timely manner. Providers have addressed this challenge via peer recommendation. Social news aggregators Digg and Reddit, for example, ask users to recommend interesting news items and prominently feature those with the most recommendations. Flickr and Yelp aggregate their users’ opinions to identify top photos and restaurants respectively. By exposing information about the preferences of others, providers hope to leverage collective intelligence <xref ref-type="bibr" rid="pone.0098914-Surowiecki1">[1]</xref> to accelerate the discovery of interesting content. In practice, however, peer recommendation often produces “winner-take-all” and “irrational herding” behaviors in which similar items receive widely different numbers of recommendations <xref ref-type="bibr" rid="pone.0098914-Salganik1">[2]</xref>, <xref ref-type="bibr" rid="pone.0098914-Muchnik1">[3]</xref>. Moreover, collective judgements obtained through peer recommendation are biased <xref ref-type="bibr" rid="pone.0098914-Lampe1">[4]</xref>, <xref ref-type="bibr" rid="pone.0098914-Lorenz1">[5]</xref> and inconsistent, with the same items ending up with very different recommendations under virtually the same conditions <xref ref-type="bibr" rid="pone.0098914-Salganik1">[2]</xref>.</p>
<p>While many strategies for aggregating opinions are possible, not all of them are equally effective in peer recommendation. We investigate some popular strategies and evaluate their ability to identify interesting content. We show that some strategies uncover the underlying population preferences for content more quickly and accurately than others. Our approach exploits <italic>position bias</italic>: people pay more attention to items at the top of a web page or a list of items than those below them <xref ref-type="bibr" rid="pone.0098914-Payne1">[6]</xref>, <xref ref-type="bibr" rid="pone.0098914-Buscher1">[7]</xref>. A consequence of this bias is the strong effect of presentation order on choices people make. For instance, presentation order affects which items in a list of search results users click on <xref ref-type="bibr" rid="pone.0098914-Joachims1">[8]</xref>–<xref ref-type="bibr" rid="pone.0098914-Yue1">[10]</xref>, and the answer they select when responding to a multiple choice question <xref ref-type="bibr" rid="pone.0098914-Payne1">[6]</xref>, <xref ref-type="bibr" rid="pone.0098914-Blunch1">[11]</xref>. Thus, a content provider can change how much attention items receive simply by changing their presentation order.</p>
<p>Studying peer recommendation is difficult due to confounding effects. These include heterogeneity of content quality, its changing relevance (novelty), commonality of user preferences (homophily), and social influence (when showing users a summary of prior users’ behavior). Another important effect is history-dependence, which can be due to having different content available at different times or the web site changing the order of presented items based on prior users’ responses. We disentangle some of these effects through randomized experiments on Amazon Mechanical Turk (Mturk), a marketplace for work <xref ref-type="bibr" rid="pone.0098914-Kittur1">[12]</xref> which is also an increasingly popular experimental platform for behavioral research <xref ref-type="bibr" rid="pone.0098914-Bohannon1">[13]</xref>–<xref ref-type="bibr" rid="pone.0098914-Crump1">[15]</xref>. The experiments allow us to determine how some of the strategies used by content providers for ordering items affect the outcomes of peer recommendation. We experimentally evaluate the effect of position bias, in contrast to previous studies of social influence <xref ref-type="bibr" rid="pone.0098914-Salganik1">[2]</xref>, <xref ref-type="bibr" rid="pone.0098914-Muchnik1">[3]</xref>, <xref ref-type="bibr" rid="pone.0098914-Lorenz1">[5]</xref>, <xref ref-type="bibr" rid="pone.0098914-Salganik2">[16]</xref>. By leveraging position bias, we can systematically direct user attention so as to improve peer recommendation. Specifically, we demonstrate that ordering items by recency of recommendation generates better estimates of underlying population preferences than ordering them by their aggregate popularity.</p>
<p>Our experiments showed people a list of science stories and asked them to recommend, or vote for, ones they found interesting. We tested five strategies for ordering content, which we refer to as “visibility policies”. The <italic>random</italic> policy presented the stories in a random order, with a new ordering generated for each participant. The <italic>popularity</italic> policy ordered stories by their popularity, i.e., in decreasing order of the number of recommendations they had already received. The <italic>activity</italic> policy ordered stories in chronological order of the latest recommendation they received, with the most recently recommended story at the top of the list. Finally, the <italic>fixed</italic> policy showed all stories in the same order to every study participant, and the <italic>reverse</italic> policy simply inverted that order. There was no adaptive ordering of content in the last two policies. Each study participant was assigned to one of these policies. We refer to participants who successfully completed the task as “users” in our study.</p>
<p>These orderings are common in social media and peer recommendation applications that exploit collective intelligence. For example, the default presentation of news stories shown in Digg’s front page (circa 2009) was by the time of promotion, which corresponds to a fixed ordering, since every user sees the stories in the same order. Digg users could also sort stories by popularity, i.e., by the number of recommendations they received during the last day or week. A Twitter stream, on the other hand, is ordered by activity, because each new retweet of an item (which we treat as a recommendation) appears at the top of a follower’s stream.</p>
<p>We demonstrate that the choice of ordering policy strongly affects the outcome of peer recommendation. We evaluate these outcomes with respect to the following goals: 1) accurately estimate population preferences for content, 2) rapidly and 3) consistently produce the estimates, and 4) focus user attention on highly interesting content. Specifically, we show that ordering items by activity produces more accurate and less variable estimates than ordering items by popularity, a widely-used policy in peer recommendation for aggregating user opinions. On the other hand, popularity-based ordering more effectively focuses attention on more interesting content.</p>
</sec><sec id="s2">
<title>Results</title>
<p>This section presents the results of our experiments. The methods section describes the experiment procedures in detail.</p>
<sec id="s2a">
<title>Story Appeal</title>
<p>Item “quality” varies significantly, although it is difficult to define or measure <xref ref-type="bibr" rid="pone.0098914-Salganik1">[2]</xref>. Instead of “quality” we use story appeal, which we define operationally as the likelihood a user who sees a story votes for (recommends) it. We assume that appeal is stable in time, which generally holds for the science stories in our experiments. While our definition of appeal conflates factors related to a story with preferences and motivations of users, it captures the notion that some content is inherently more appealing or interesting to a community. In general, this conditional probability is difficult to measure because it requires knowing both whether a user saw and voted on a story. While votes are readily recorded, views are not readily available, e.g., requiring eye tracking or, for a less precise measure, whether particular content was delivered to the user’s browser. Nevertheless, controlled experiments can measure the average appeal of a story to a user population <xref ref-type="bibr" rid="pone.0098914-Salganik1">[2]</xref> by, for example, randomizing over possible confounding effects such as the order of the story. After enough people had seen each story, the number of votes they receive will reflect how interesting or appealing people find them.</p>
<p>The random policy in our experiments provides the control for estimating appeal. Specifically, we define the appeal <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e002" xlink:type="simple"/></inline-formula> of a story <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e003" xlink:type="simple"/></inline-formula> to a population of users as the fraction of users in a sufficiently large sample from that population who vote for <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e004" xlink:type="simple"/></inline-formula>. The random policy averages over positions, so <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e005" xlink:type="simple"/></inline-formula> captures the underlying population preferences for stories. <xref ref-type="fig" rid="pone-0098914-g001">Fig. 1</xref> shows that appeal is broadly distributed, varying by about a factor of four among stories.</p>
<fig id="pone-0098914-g001" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.g001</object-id><label>Figure 1</label><caption>
<title>Distribution of story appeal <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e001" xlink:type="simple"/></inline-formula>, i.e., probabilities users vote on each story under the random order policy.</title>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.g001" position="float" xlink:type="simple"/></fig></sec><sec id="s2b">
<title>Position Bias</title>
<p>The probabilities for votes on each story (i.e., its appeal) allow estimating the number of votes we would expect at each position in the random policy. Specifically, suppose stories <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e006" xlink:type="simple"/></inline-formula> are shown to successive users at position <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e007" xlink:type="simple"/></inline-formula>. The expected number of votes for these stories is <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e008" xlink:type="simple"/></inline-formula>. With <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e009" xlink:type="simple"/></inline-formula> the actual number of votes for these stories, the ratio <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e010" xlink:type="simple"/></inline-formula> is the relative increase or decrease in votes for that position compared to average, i.e., position bias. <xref ref-type="fig" rid="pone-0098914-g002">Fig. 2</xref> shows these ratios.</p>
<fig id="pone-0098914-g002" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.g002</object-id><label>Figure 2</label><caption>
<title>Position bias: variation in votes based on position.</title>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.g002" position="float" xlink:type="simple"/></fig>
<p>Position bias is quite pronounced: a story at the top of a list gets about five times as much attention as a story lower in the list. This behavior is similar to how users respond to web search results <xref ref-type="bibr" rid="pone.0098914-Joachims1">[8]</xref>–<xref ref-type="bibr" rid="pone.0098914-Yue1">[10]</xref>, content in social media (e.g., Digg <xref ref-type="bibr" rid="pone.0098914-Hogg1">[17]</xref>), and online cultural marketplaces <xref ref-type="bibr" rid="pone.0098914-Salganik2">[16]</xref>, <xref ref-type="bibr" rid="pone.0098914-Krumme1">[18]</xref>. The moderate increase in votes at the end of the list was observed by Salganik et al. <xref ref-type="bibr" rid="pone.0098914-Salganik2">[16]</xref>, who attributed it to ‘contrarians’, who navigate the list starting from the end. Another possibility is this behavior results from strategic decisions made by participants to give an impression that they had inspected all stories.</p>
</sec><sec id="s2c">
<title>Votes and Appeal</title>
<p><xref ref-type="fig" rid="pone-0098914-g003">Fig. 3</xref> shows the variation in votes on stories, compared to the random policy. The activity policy, by continually moving recommended stories to the top of the list, divides user votes roughly in proportion to their appeal. The popularity policy is much more variable, both among stories with similar appeal and between repeated experiments. The fixed policy focuses user attention on the same stories, leading to a large deviation from their appeal. Similarly, all users in the reverse policy see the stories in the same order, which also leads to a large deviation. Specifically, the fixed and reverse policies have correlations between votes and appeal of <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e011" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e012" xlink:type="simple"/></inline-formula>, respectively. Both parallel worlds for the activity policy have larger correlations, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e013" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e014" xlink:type="simple"/></inline-formula>, while the popularity policy is intermediate between activity and fixed, with correlations <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e015" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e016" xlink:type="simple"/></inline-formula> in the parallel worlds experiments. These correlations are statistically significant, with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e017" xlink:type="simple"/></inline-formula>-values less than <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e018" xlink:type="simple"/></inline-formula> in all cases according to the Spearman rank test for zero correlation. The activity policy leads to, on average, higher correlation between votes and appeal than the other policies. Since an item’s popularity is often used as a proxy for how appealing it is to a user population, the activity policy is better for evaluating items.</p>
<fig id="pone-0098914-g003" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.g003</object-id><label>Figure 3</label><caption>
<title>Fraction of users voting for a story vs its appeal <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e019" xlink:type="simple"/></inline-formula> under different policies for ordering stories.</title>
<p>The lines are the expected number of votes per user based on the random policy.</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.g003" position="float" xlink:type="simple"/></fig>
<p>Next we examine how quickly the policies estimate appeal. While both popularity and activity policies quickly converge to their estimates, the popularity policy may be slow to respond to changing user interests. This is because after the first 50 or so users, the popularity policy becomes a (nearly) fixed ordering, with stories near the top of the list accumulating votes more rapidly than other stories, making it difficult for a new, more appealing story to reach the top position. One measure of the responsiveness of a policy is how rapidly the number of votes approaches that expected from the stories’ appeal. <xref ref-type="fig" rid="pone-0098914-g004">Fig. 4</xref> shows this behavior. Repeated experiments with each policy give consistent behavior. Activity converges more rapidly, and to a higher correlation with appeal, than popularity. The final values of the correlations correspond to those for all votes, discussed with <xref ref-type="fig" rid="pone-0098914-g003">Fig. 3</xref>.</p>
<fig id="pone-0098914-g004" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.g004</object-id><label>Figure 4</label><caption>
<title>Correlation between number of votes each story receives and its appeal as a function of number of users voting.</title>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.g004" position="float" xlink:type="simple"/></fig></sec><sec id="s2d">
<title>Inequality of Outcomes</title>
<p>Variations in the distribution of attention produced by different orderings lead to large differences in the number of votes stories receive, i.e., their popularity. Since stories differ in appeal, when attention is distributed uniformly (as in the random policy) we expect votes to vary in proportion to their appeal. Orderings that direct user attention toward the same stories will result in greater inequality of popularity.</p>
<p>We quantify the variation in popularity of stories by the Gini coefficient, a measure of statistical dispersion:<disp-formula id="pone.0098914.e020"><graphic position="anchor" xlink:href="info:doi/10.1371/journal.pone.0098914.e020" xlink:type="simple"/><label>(1)</label></disp-formula>where <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e021" xlink:type="simple"/></inline-formula> is the number of stories and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e022" xlink:type="simple"/></inline-formula> is the fraction of all votes that story <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e023" xlink:type="simple"/></inline-formula> receives, so <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e024" xlink:type="simple"/></inline-formula>. In our experiments, <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e025" xlink:type="simple"/></inline-formula>.</p>
<p><xref ref-type="fig" rid="pone-0098914-g005">Fig. 5</xref> shows the values of the Gini coefficient in our experiments. In the random policy, the fraction <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e026" xlink:type="simple"/></inline-formula> is, by definition, the appeal <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e027" xlink:type="simple"/></inline-formula> for that story. Thus the value for the random policy indicates the inequality expected solely from the variation in story appeal. The activity policy results in slightly more inequality than would be expected from the inherent differences in story appeal. On the other hand, a policy that shows stories in a fixed order focuses attention on the same most visible stories, leading to a large inequality in the distribution of votes. This is the case for the fixed and reverse policies. This observation also explains the large inequality in the popularity policy because its story order essentially stops changing after 50 users make recommendations. Thus, for subsequent users its position bias is similar to that of a fixed policy. As a consistency check, the two parallel worlds for each of the activity and popularity policies give the same Gini coefficients. Nevertheless, the particular stories receiving the most votes differ between the two worlds, especially for the popularity policy.</p>
<fig id="pone-0098914-g005" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.g005</object-id><label>Figure 5</label><caption>
<title>Gini coefficient showing inequality of the total votes received by items in different policies.</title>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.g005" position="float" xlink:type="simple"/></fig>
<p>For the policies without history dependence (i.e., random, fixed and reverse), we can assess the significance of the different Gini coefficients with a permutation test. Specifically, to compare two policies under the null hypothesis that they do not differ in how user choices contribute to inequality, we randomly permute the users in those experiments between the policies, while keeping the same <italic>number</italic> of users assigned to each policy. From this permutation, we compute the resulting difference in Gini coefficients. Repeating this evaluation many times gives an estimate of how the difference would vary if user behavior was the same in the two policies. Comparing this variation with the actual difference in Gini coefficient between those two policies indicates how likely that observed difference could arise under the null hypothesis. We use this method to compare each pair of the three policies (random, fixed and reverse), using 100 permutations for each pair. In all cases, the observed difference in Gini coefficient is larger than the differences from all these permutations, indicating the differences are significant with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e028" xlink:type="simple"/></inline-formula>-value below <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e029" xlink:type="simple"/></inline-formula>.</p>
<p>This permutation test does not apply to the policies with history dependence (i.e., activity and popularity), since the presentation of stories depends on the actions of previous users. Instead, repeating the experiments (i.e., parallel worlds) gives independent estimates of the Gini coefficient for these policies. The small differences in Gini coefficients between the parallel worlds for each policy suggests the popularity policy leads to greater inequality than the activity policy.</p>
</sec><sec id="s2e">
<title>Predictability of Outcomes</title>
<p><xref ref-type="fig" rid="pone-0098914-g003">Fig. 3</xref> shows votes under the popularity policy have larger variation than those under the activity policy, particularly for high-appeal stories. Moreover, comparing the outcomes of parallel worlds experiments shows much larger consistency between worlds for the activity policy. For instance, the top quartile of stories (those with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e030" xlink:type="simple"/></inline-formula>) have correlations <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e031" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e032" xlink:type="simple"/></inline-formula> between votes in the parallel worlds for popularity and activity policies, respectively. This large a difference in correlation is unlikely to arise if in fact these two policies had the same correlations between parallel worlds (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e033" xlink:type="simple"/></inline-formula>-value <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e034" xlink:type="simple"/></inline-formula> with the Spearman rank test). Moreover, the pattern of votes in the two parallel worlds for popularity is consistent with no correlation between the worlds (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e035" xlink:type="simple"/></inline-formula>-value 0.2 with Spearman rank test). On the other hand, zero correlation is unlikely for the activity policy (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e036" xlink:type="simple"/></inline-formula>-value <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e037" xlink:type="simple"/></inline-formula>). Thus outcomes are more predictable for the activity policy: a given high-appeal story is more likely to get a similar number of votes if repeated with a new group of users. Popularity, on the other hand, is less consistent due to the amplification of the effects of early votes through its “rich get richer” behavior.</p>
<p>In contrast with stories in the top quartile, these two policies have no significant difference in correlation for the less appealing stories: those stories receive similar, low numbers of votes in both parallel worlds for each policy.</p>
</sec><sec id="s2f">
<title>Focusing Attention on Appealing Items</title>
<p>How well do the visibility policies focus user attention on appealing stories? This is an important measure of user experience in peer recommendation systems: showing users appealing stories indicates to those users the site has interesting content, making it more likely the users will return to the site <xref ref-type="bibr" rid="pone.0098914-Brandtzaeg1">[19]</xref>.</p>
<p>Web users typically view only a fraction of the available content, starting from the top of the list of items. Thus one measure of user experience is the appeal of the stories they are most likely to view, i.e., those near the top of the list. We quantify this aspect of user experience by how well the policy delivers high-appeal stories to early positions in the list of stories shown to a user. As a specific example, we examine the first 20 positions and measure the fraction of those positions containing stories whose appeal is among the top 20% of stories (as measured in the random policy).</p>
<p><xref ref-type="fig" rid="pone-0098914-g006">Fig. 6</xref> shows the resulting distributions for users assigned to the activity and popularity policies. By this measure of user experience, the activity policy has lower average fraction and is more variable among users than the popularity policy. In other words, users assigned to the activity policy tend to see fewer top stories than users assigned to the popularity policy. Moreover, the high variability under the activity policy means a significant fraction of users are likely to see very few top stories. For comparison, users assigned to the random policy will likely see about 20% of the top stories, which is even less than the activity policy.</p>
<fig id="pone-0098914-g006" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.g006</object-id><label>Figure 6</label><caption>
<title>Distribution of fraction of the first 20 stories shown to a user that are among the most-appealing 20% of stories.</title>
<p>Under the popularity policy, for most users at least 40% of the initial stories are among the most-appealing stories, whereas under the activity policy, most users see fewer than 40%. These histograms do not include the first 50 users in each experiment, to avoid the initialization phase of the policies.</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.g006" position="float" xlink:type="simple"/></fig></sec></sec><sec id="s3">
<title>Discussion</title>
<p>Our findings demonstrate that the ordering of items significantly affects the outcome of peer recommendation. The differences in outcomes stem from human cognitive biases, specifically the position bias that results in people paying more attention to items appearing near the top of the list. These items have high visibility, since it takes little effort to discover them. The more effort required to find an item, the less attention it will receive. While this bias cannot be altered, we can control which items people pay attention to simply by changing their position in the list of items.</p>
<p>Visibility policies differ in how well they fulfill the goals of peer recommendation described in the introduction. Clearly, random policy is best for unbiased estimates of preferences. However, since a small fraction of user-generated content is interesting, users will mainly see uninteresting content under the random policy. As a consequence, they may then form an impression that the site does not provide anything of interest and fail to return. Unlike the random policy, the popularity policy does not accurately estimate preferences, since small early differences in popularity may be amplified via a “rich get richer” effect. As a result, item ordering quickly becomes fixed, which leads to greater inequality and less consistency. On the other hand, the popularity policy emphasizes highly appealing content for users better than the random policy does. In contrast, the recency condition of the activity policy leads to more robust estimates of underlying population preferences than ordering by popularity. It was second only to the random policy in how well the observed popularity correlated with user interests in items, and also produced less variable, more predictable outcomes. While the activity policy was not as effective as the popularity policy at focusing user attention on appealing content, it was better than the random policy. The activity policy is also a good choice for time critical domains, where novelty is a factor, since continuously moving items to the top of the list can rapidly bring newer items to users’ attention. In summary, the choice of ordering allows steering peer recommendation toward a desired goal, such as accurately estimating appeal or highlighting interesting content for users visiting the web site.</p>
<p>Beyond peer recommendation, position bias also affects the performance of social media, discussion forums, online markets, and crowdfunding sites. Specifically, the amount of attention a message receives in social media is largely determined by its position in the user’s stream, and this affects the ease with which the message spreads <xref ref-type="bibr" rid="pone.0098914-Hogg1">[17]</xref>, <xref ref-type="bibr" rid="pone.0098914-Hodas1">[20]</xref>. By directing user attention to certain messages, a social media site can selectively enhance their spread. Crowdsourcing applications which require users to select tasks or items from a list can similarly manipulate individuals’ attention to drive human computation in a particular direction. In online discussion forums, user attention can be directed so as to improve the performance of distributed moderation. Current moderation schemes can give messages unfairly low scores, because early negative scores reduce their visibility and prevent them from receiving the attention needed for a fair evaluation <xref ref-type="bibr" rid="pone.0098914-Lampe1">[4]</xref>.</p>
<p>Quantitative understanding of position bias is important from the design perspective, as it allows for more accurate and robust estimation of how interesting some content is to a user population. For instance, a web site could estimate content appeal from the responses of an initial cohort of users, and then place content with the highest estimated appeal in the most visible positions to improve user experience. The web site could also adjust its presentation method dynamically to adapt to changing user preferences and content novelty.</p>
<p>Our study did not directly examine social influence, since users were not shown the number of votes stories received. For influence to occur, a social signal has to be present, but even then, the individual first has to discover the item before he or she can be affected by this signal. Hence, an item’s visibility, which affects how easily it can be discovered, plays a big role in how popular it will become.</p>
<p>Our experiments are similar in design to those of Salganik et al. <xref ref-type="bibr" rid="pone.0098914-Salganik1">[2]</xref>, <xref ref-type="bibr" rid="pone.0098914-Salganik3">[21]</xref>, which examined why some cultural artifacts become vastly more popular than others, and why their popularity is largely unpredictable. The studies asked participants to rate songs by unknown bands. Songs were presented either in random order (<italic>cf</italic> random policy) or sorted by popularity (<italic>cf</italic> popularity policy). Salganik et al. found that sorting by popularity resulted in more unpredictability and greater inequality of popularity. Moreover, providing a signal of popularity, by showing participants how popular songs are, further increased inequality and unpredictability. They attributed both effects to social influence. In contrast, our study suggests that inequality and unpredictability of popularity could arise even in the absence of social influence, since biases in perception lead users to pay more attention to items near the top of the list. If those items are already the most popular ones, this creates a “rich get richer” effect that amplifies their popularity. A re-examination of Salganik et al.’s experimental data <xref ref-type="bibr" rid="pone.0098914-Krumme1">[18]</xref> showed that a song’s position in the list can explain much of its near-term popularity. This is encouraging, as it suggests that knowing an item’s visibility can help predict its future success.</p>
</sec><sec id="s4" sec-type="methods">
<title>Methods</title>
<p>University of Southern California’s Institutional Review Board (IRB) reviewed the experiment design and classified it as “non-human subjects research.” Our experiments were published as tasks (HITs) on Amazon Mechanical Turk, which allowed us to recruit study participants from a large pool of workers. Workers who accepted the task were shown the following instructions: “We are conducting a study of the role of social media in promoting science. Please click ‘Start’ button and recommend articles from the list below that you think report important scientific topics. When you finish, you will be asked a few questions about the articles you recommended. (Please remember, once you finish the job, system won’t allow you to do it again).” They were paid $0.12 for completing the experiment and each person was allowed to do the experiment only once. The pay rate was set low to make the task less attractive to workers attempting to game Mturk and is comparable to similar tasks in other research studies <xref ref-type="bibr" rid="pone.0098914-Kittur1">[12]</xref>, <xref ref-type="bibr" rid="pone.0098914-Mason1">[14]</xref>. Although we paid people to vote, we assume their behavior is similar to that in recommendation systems. This assumption is validated by the growing body of work using Mturk for behavioral research <xref ref-type="bibr" rid="pone.0098914-Bohannon1">[13]</xref>–<xref ref-type="bibr" rid="pone.0098914-Crump1">[15]</xref>.</p>
<p>We showed the participants a list of one hundred science stories, drawn from the Science section of the New York Times and science-related press releases from major universities (sciencenewsdaily.com). Stories were delivered to the browser in a single page, as illustrated in <xref ref-type="fig" rid="pone-0098914-g007">Fig. 7</xref>. The list was sufficiently long to require them to scroll to see all stories. Each story contained a title, a short description, and a link to a page where the person could read the full story. Participants could choose to recommend a story based on the short description or click on the link to view the full story. We recorded all actions, including recommendations and URL clicks, and the position of all stories shown to each participant. When a person recommended a story, the recommend button changed color to indicate that story was recommended. The experiment did not allow participants to undo their recommendations: subsequent clicks of the recommend button brought up a message box reminding participants to recommend a story only once. Although participants were not told ahead of time how many stories to recommend, if they tried to finish the task before making five actions (either recommendations or URL clicks), a message box prompted them to make five recommendations.</p>
<fig id="pone-0098914-g007" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.g007</object-id><label>Figure 7</label><caption>
<title>Screenshot of a web page shown to participants.</title>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.g007" position="float" xlink:type="simple"/></fig>
<p>Upon finishing the task, participants were asked to name two important themes in the stories they recommended and solve a simple arithmetic question. Only those who correctly answered the arithmetic question were considered to have completed the task and paid. There were nine participants with corrupted session data, which were not included in the analysis. Of the 4,007 workers who accepted the task, only 2,643 completed it. Further, to ensure data quality (see below), we ignored recommendations made by participants who recommended more than 20 stories. The recommendations made by the remaining 1518 people (i.e., users) were saved in a database and are summarized in <xref ref-type="table" rid="pone-0098914-t001">Table 1</xref>. Only these recommendations were used in analysis. Recommendations data are available from the authors upon request.</p>
<table-wrap id="pone-0098914-t001" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.t001</object-id><label>Table 1</label><caption>
<title>Summary of experiments.</title>
</caption><alternatives><graphic id="pone-0098914-t001-1" position="float" mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.t001" xlink:type="simple"/>
<table><colgroup span="1"><col align="left" span="1"/><col align="center" span="1"/><col align="center" span="1"/><col align="center" span="1"/><col align="center" span="1"/></colgroup>
<thead>
<tr>
<td align="left" rowspan="1" colspan="1">policy</td>
<td align="left" rowspan="1" colspan="1">users</td>
<td align="left" rowspan="1" colspan="1">votes</td>
<td align="left" rowspan="1" colspan="1">avg.</td>
<td align="left" rowspan="1" colspan="1">std. dev.</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="1" colspan="1">random</td>
<td align="left" rowspan="1" colspan="1">199</td>
<td align="left" rowspan="1" colspan="1">1873</td>
<td align="left" rowspan="1" colspan="1">9.4</td>
<td align="left" rowspan="1" colspan="1">4.7</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1">fixed</td>
<td align="left" rowspan="1" colspan="1">217</td>
<td align="left" rowspan="1" colspan="1">1978</td>
<td align="left" rowspan="1" colspan="1">9.1</td>
<td align="left" rowspan="1" colspan="1">4.8</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1">reverse</td>
<td align="left" rowspan="1" colspan="1">221</td>
<td align="left" rowspan="1" colspan="1">1999</td>
<td align="left" rowspan="1" colspan="1">9.0</td>
<td align="left" rowspan="1" colspan="1">4.7</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1">activity</td>
<td align="left" rowspan="1" colspan="1">286 &amp; 193</td>
<td align="left" rowspan="1" colspan="1">2586 &amp; 1764</td>
<td align="left" rowspan="1" colspan="1">9.0 &amp; 9.1</td>
<td align="left" rowspan="1" colspan="1">4.5 &amp; 4.6</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1">popularity</td>
<td align="left" rowspan="1" colspan="1">174 &amp; 228</td>
<td align="left" rowspan="1" colspan="1">1570 &amp; 2162</td>
<td align="left" rowspan="1" colspan="1">9.0 &amp; 9.5</td>
<td align="left" rowspan="1" colspan="1">4.7 &amp; 4.5</td>
</tr>
<tr>
<td align="left" rowspan="1" colspan="1">total</td>
<td align="left" rowspan="1" colspan="1">1518</td>
<td align="left" rowspan="1" colspan="1">13932</td>
<td align="left" rowspan="1" colspan="1"/>
<td align="left" rowspan="1" colspan="1"/>
</tr>
</tbody>
</table>
</alternatives><table-wrap-foot><fn id="nt101"><label/><p>Number of participants and votes made under different visibility policies. The history-dependent orderings (activity and popularity) each have two independent experiments. The last two columns give the average and standard deviation of number of votes per user.</p></fn></table-wrap-foot></table-wrap><sec id="s4a">
<title>Visibility Policy</title>
<p>Our experiments allow controlling the presentation of stories and monitoring URL clicks and recommendations, but not tracking which stories are viewed. We studied the visibility policies described above. In each experiment, stories initially had no recommendations, and the popularity and activity policies used the same story order as the fixed policy. The fixed order was also used to break ties in the popularity and activity policies. The random policy was our control condition.</p>
<p>As our focus is on the effect of visibility, we eliminate any confounding effect of social influence <xref ref-type="bibr" rid="pone.0098914-Salganik1">[2]</xref> by not showing the number of recommendations the stories received or disclosing the method by which we ordered stories. We tested the reproducibility of results for the history-dependent activity and popularity policies by creating “parallel worlds” experiments <xref ref-type="bibr" rid="pone.0098914-Salganik1">[2]</xref>, in which we ran two instances of each policy starting from the same initial conditions.</p>
</sec><sec id="s4b">
<title>Data Quality Control</title>
<p>Amazon Mechanical Turk is an appealing platform for studies of human behavior. However, a major challenge for using Mturk is ensuring data quality <xref ref-type="bibr" rid="pone.0098914-Mason1">[14]</xref>, because some workers, i.e., spammers, fail to exert the effort necessary to evaluate stories. Instead they do the least work to get paid, e.g., click on the first story or on every story.</p>
<p>We used a multi-step strategy to reduce spam. First, we selected workers using qualifications provided by Mturk: they lived in the US, had completed at least 500 tasks on Mturk, and had a 90% or above approval rate. In addition, after workers finished recommending stories, we asked them to solve a simple arithmetic problem. A new problem was generated after an incorrect answer, preventing them from finding the solution by exhaustive search.</p>
<p>In spite of our selection process, we found large apparent variation in motivation. Some participants appeared not to make a serious effort in evaluating stories and simply recommended most or all of the stories. To exclude such people, our vetting procedure accepted only the recommendations from participants who recommended at most 20 stories. Such vetted participants were the users in our study. They generally spent more time evaluating each story. <xref ref-type="fig" rid="pone-0098914-g008">Fig. 8</xref> shows the distribution of session times (excluding the time required to read instructions and do the post-survey) and the average time taken by participants to recommend a story. While non-vetted participants spent a little more time on the task, it took a typical vetted participant (voting on at most 20 stories) 25 seconds to recommend a story, while a non-vetted participant required fewer than 10 seconds. These differences are statistically significant (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e038" xlink:type="simple"/></inline-formula>-values less than <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e039" xlink:type="simple"/></inline-formula> with Mann-Whitney tests). In addition, the rate at which participants clicked URLs, an action not required by the task but which suggested motivation, was higher for vetted (27%) than non-vetted participants (22%), with <inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e040" xlink:type="simple"/></inline-formula>-test indicating these proportions are different (<inline-formula><inline-graphic xlink:href="info:doi/10.1371/journal.pone.0098914.e041" xlink:type="simple"/></inline-formula>-value 0.01). Although the choice of the 20-recommendation threshold is somewhat arbitrary, timing results and URL clicks indicate that it appropriately weeded out unmotivated participants.</p>
<fig id="pone-0098914-g008" position="float"><object-id pub-id-type="doi">10.1371/journal.pone.0098914.g008</object-id><label>Figure 8</label><caption>
<title>Distribution of session time (left column) and average time per vote (right column) for vetted and non-vetted participants.</title>
<p>Participants are grouped according to their activity, i.e., number of votes, with each group (indicated by a colored bar) containing about 500 people. A few people with longer session times and times per vote are not included in the plots.</p>
</caption><graphic mimetype="image" xlink:href="info:doi/10.1371/journal.pone.0098914.g008" position="float" xlink:type="simple"/></fig></sec></sec></body>
<back>
<ack>
<p>We thank Kuai Yu and Suradej Intagorn for their help with setting up Mturk experiments, and Katrina Pariera and Jake de Grazia for discussions regarding experimental design.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="pone.0098914-Surowiecki1"><label>1</label>
<mixed-citation publication-type="other" xlink:type="simple">Surowiecki J (2005) The Wisdom of Crowds. Anchor.</mixed-citation>
</ref>
<ref id="pone.0098914-Salganik1"><label>2</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Salganik</surname><given-names>MJ</given-names></name>, <name name-style="western"><surname>Dodds</surname><given-names>PS</given-names></name>, <name name-style="western"><surname>Watts</surname><given-names>DJ</given-names></name> (<year>2006</year>) <article-title>Experimental study of inequality and unpredictability in an artificial cultural market</article-title>. <source>Science</source> <volume>311</volume>: <fpage>854</fpage>–<lpage>856</lpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Muchnik1"><label>3</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Muchnik</surname><given-names>L</given-names></name>, <name name-style="western"><surname>Aral</surname><given-names>S</given-names></name>, <name name-style="western"><surname>Taylor</surname><given-names>SJ</given-names></name> (<year>2013</year>) <article-title>Social influence bias: A randomized experiment</article-title>. <source>Science</source> <volume>341</volume>: <fpage>647</fpage>–<lpage>651</lpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Lampe1"><label>4</label>
<mixed-citation publication-type="other" xlink:type="simple">Lampe C, Resnick P (2004) Slash(dot) and burn: Distributed moderation in a large online conversation space. In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. New York, NY, USA: ACM, CHI’ 04, 543–550. URL <ext-link ext-link-type="uri" xlink:href="http://doi.acm.org/10.1145/985692.985761" xlink:type="simple">http://doi.acm.org/10.1145/985692.985761</ext-link>. doi:<ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1145/985692.985761" xlink:type="simple">10.1145/985692.985761</ext-link>.</mixed-citation>
</ref>
<ref id="pone.0098914-Lorenz1"><label>5</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Lorenz</surname><given-names>J</given-names></name>, <name name-style="western"><surname>Rauhut</surname><given-names>H</given-names></name>, <name name-style="western"><surname>Schweitzer</surname><given-names>F</given-names></name>, <name name-style="western"><surname>Helbing</surname><given-names>D</given-names></name> (<year>2011</year>) <article-title>How social influence can undermine the wisdom of crowd effect</article-title>. <source>Proceedings of the National Academy of Sciences</source> <volume>108</volume>: <fpage>9020</fpage>–<lpage>9025</lpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Payne1"><label>6</label>
<mixed-citation publication-type="other" xlink:type="simple">Payne SL (1951) The Art of Asking Questions. Princeton University Press.</mixed-citation>
</ref>
<ref id="pone.0098914-Buscher1"><label>7</label>
<mixed-citation publication-type="other" xlink:type="simple">Buscher G, Cutrell E, Morris MR (2009) What do you see when you’re surfing?: using eye tracking to predict salient regions of web pages. In: Proc. the 27th Int. Conf. on Human factors in computing systems. New York, NY, USA, 21–30.</mixed-citation>
</ref>
<ref id="pone.0098914-Joachims1"><label>8</label>
<mixed-citation publication-type="other" xlink:type="simple">Joachims T, Granka L, Pan B, Hembrooke H, Gay G (2005) Accurately interpreting clickthrough data as implicit feedback. In: Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. New York, NY, USA: ACM, SIGIR’ 05, 154–161. URL <ext-link ext-link-type="uri" xlink:href="http://doi.acm.org/10.1145/1076034.1076063" xlink:type="simple">http://doi.acm.org/10.1145/1076034.1076063</ext-link>. doi:<ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1145/1076034.1076063" xlink:type="simple">10.1145/1076034.1076063</ext-link>.</mixed-citation>
</ref>
<ref id="pone.0098914-Craswell1"><label>9</label>
<mixed-citation publication-type="other" xlink:type="simple">Craswell N, Zoeter O, Taylor M, Ramsey B (2008) An experimental comparison of click positionbias models. In: Proceedings of the international conference on Web search and web data mining. WSDM’ 08, 87–94.</mixed-citation>
</ref>
<ref id="pone.0098914-Yue1"><label>10</label>
<mixed-citation publication-type="other" xlink:type="simple">Yue Y, Patel R, Roehrig H (2010) Beyond position bias: Examining result attractiveness as a source of presentation bias in clickthrough data. In: Proceedings of the 19th International Conference on World Wide Web. New York, NY, USA: ACM, WWW’ 10, 1011–1018. URL <ext-link ext-link-type="uri" xlink:href="http://doi.acm.org/10.1145/1772690.1772793" xlink:type="simple">http://doi.acm.org/10.1145/1772690.1772793</ext-link>. doi:<ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1145/1772690.1772793" xlink:type="simple">10.1145/1772690.1772793</ext-link>.</mixed-citation>
</ref>
<ref id="pone.0098914-Blunch1"><label>11</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Blunch</surname><given-names>NJ</given-names></name> (<year>1984</year>) <article-title>Position bias in multiple-choice questions</article-title>. <source>Journal of Marketing Research</source> <volume>21</volume>: <fpage>216</fpage>–<lpage>220</lpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Kittur1"><label>12</label>
<mixed-citation publication-type="other" xlink:type="simple">Kittur A, Nickerson JV, Bernstein M, Gerber E, Shaw A, et al. (2013) The future of crowd work. In: Proceedings of the 2013 Conference on Computer Supported Cooperative Work. New York, NY, USA: ACM, CSCW’ 13, 1301–1318. URL <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1145/2441776.2441923" xlink:type="simple">http://dx.doi.org/10.1145/2441776.2441923</ext-link>. doi:<ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.1145/2441776.2441923" xlink:type="simple">10.1145/2441776.2441923</ext-link>.</mixed-citation>
</ref>
<ref id="pone.0098914-Bohannon1"><label>13</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Bohannon</surname><given-names>J</given-names></name> (<year>2011</year>) <article-title>Social science for pennies</article-title>. <source>Science</source> <volume>334</volume>: <fpage>307</fpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Mason1"><label>14</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Mason</surname><given-names>W</given-names></name>, <name name-style="western"><surname>Suri</surname><given-names>S</given-names></name> (<year>2012</year>) <article-title>Conducting behavioral research on Amazon’s Mechanical Turk</article-title>. <source>Behavior Research Methods</source> <volume>44</volume>: <fpage>1</fpage>–<lpage>23</lpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Crump1"><label>15</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Crump</surname><given-names>MJC</given-names></name>, <name name-style="western"><surname>McDonnell</surname><given-names>JV</given-names></name>, <name name-style="western"><surname>Gureckis</surname><given-names>TM</given-names></name> (<year>2013</year>) <article-title>Evaluating Amazon’s Mechanical Turk as a tool for experimental behavioral research</article-title>. <source>PLos ONE</source> <volume>8</volume>: <fpage>e57410</fpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Salganik2"><label>16</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Salganik</surname><given-names>MJ</given-names></name>, <name name-style="western"><surname>Watts</surname><given-names>DJ</given-names></name> (<year>2008</year>) <article-title>Leading the herd astray: An experimental study of self-fulfilling prophecies in an artificial cultural market</article-title>. <source>Social Psychology Quarterly</source> <volume>71</volume>: <fpage>338</fpage>–<lpage>355</lpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Hogg1"><label>17</label>
<mixed-citation publication-type="other" xlink:type="simple">Hogg T, Lerman K (2012) Social dynamics of digg. EPJ Data Science 1.</mixed-citation>
</ref>
<ref id="pone.0098914-Krumme1"><label>18</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Krumme</surname><given-names>C</given-names></name>, <name name-style="western"><surname>Cebrian</surname><given-names>M</given-names></name>, <name name-style="western"><surname>Pickard</surname><given-names>G</given-names></name>, <name name-style="western"><surname>Pentland</surname><given-names>S</given-names></name> (<year>2012</year>) <article-title>Quantifying social influence in an online cultural market</article-title>. <source>PLoS ONE</source> <volume>7</volume>: <fpage>e33785</fpage>.</mixed-citation>
</ref>
<ref id="pone.0098914-Brandtzaeg1"><label>19</label>
<mixed-citation publication-type="other" xlink:type="simple">Brandtzaeg PB, Heim J (2007) User loyalty and online communities: why members of online communities are not faithful. In: INTETAIN’ 08: Proceedings of the 2nd international conference on INtelligent TEchnologies for interactive enterTAINment. 1–10.</mixed-citation>
</ref>
<ref id="pone.0098914-Hodas1"><label>20</label>
<mixed-citation publication-type="other" xlink:type="simple">Hodas N, Lerman K (2012) How limited visibility and divided attention constrain social contagion. In: In ASE/IEEE International Conference on Social Computing.</mixed-citation>
</ref>
<ref id="pone.0098914-Salganik3"><label>21</label>
<mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Salganik</surname><given-names>MJ</given-names></name>, <name name-style="western"><surname>Watts</surname><given-names>DJ</given-names></name> (<year>2009</year>) <article-title>Web-Based experiments for the study of collective social dynamics in cultural markets</article-title>. <source>Topics in Cognitive Science</source> <volume>1</volume>: <fpage>439</fpage>–<lpage>468</lpage>.</mixed-citation>
</ref>
</ref-list></back>
</article>