{ From user Lonnie, Model Intro_to_rank_correl at 25-Mar-2010 11:21:16 AM } Softwareversion 4.3.0 { System Variables with non-default values: } Samplesize := 10K Usetable := 0 Displayoutputs Run: , Typechecking := 1 Checking := 1 Saveoptions := 2 Savevalues := 0 Distresol := 500 {!40000|Att_contlinestyle Graph_primary_valdim: 4} {!40000|Att_contlinestyle Graph_pdf_valdim: 1} Model Intro_to_rank_correl Title: Intro to Rank Correlation Description: This model, created during the Analytica User Group Webinar on 25 March 2010, was used to illustrate the uses and estimation of Rank Correlation. Author: Lonnie Chrisman~ Lumina Decision Systems Date: Wed, Mar 24, 2010 11:51 PM Saveauthor: Lonnie Savedate: Thu, Mar 25, 2010 11:21 AM Defaultsize: 48,24 Diagstate: 2,33,12,472,386,17 Windstate: 2,98,83,476,224 Fontstyle: Arial, 15 Fileinfo: 0,Model Intro_to_rank_correl,2,2,0,0,W:\Training\User Group Webinars\Rank-Correlation-Analysis.ana Module Soccer_data Title: Soccer Data Description: Data was collected from girls' varsity and JV soccer teams at a high school as part of a Masters project in sports Kinesiology, examing whether self-confidence and injury rates are statistically related. The athletes took a standard self-confidence survey before and after the season, and the coach tracked all injuries sustained during the season in detail which have been condensed here to an "injury score" (which scores both severity of injuries and number of injuries during the season).~ ~ This data is used here to demonstrate small-sample rank correlation analysis for the Analytica user group webinar. Author: Lenka Berenova and Lonnie Chrisman Date: Thu, Mar 25, 2010 9:16 AM Defaultsize: 48,24 Nodelocation: 88,56,1 Nodesize: 48,24 Diagstate: 2,79,17,587,475,17 Index Player Title: Player Definition: ['V1','V2','V3','V4','V6','V7','V8','V9','V10','V11','V12','V13','V14','V15','JV1','JV2','JV3','JV5','JV6','JV7','JV9','JV10','JV11','JV12','JV13','JV14','JV15','JV16','JV17'] Nodelocation: 100,48,1 Nodesize: 48,24 {!40000|Att_previndexvalue: ['V1','V2','V3','V4','V6','V7','V8','V9','V10','V11','V12','V13','V14','V15','JV1','JV2','JV3','JV5','JV6','JV7','JV9','JV10','JV11','JV12','JV13','JV14','JV15','JV16','JV17']} Constant Pre_season_self_conf Title: Pre-season Self Conf. Definition: Table(Player)(~ 84,72,82,90,87,86,96,93,91,78,99,86,90,71,94,92,91,92,83,93,104,62,65,93,95,70,87,96,80) Nodelocation: 100,104,1 Nodesize: 52,24 Valuestate: 2,260,2,416,568,0,MIDM Constant Post_season_self_con Title: Post-season Self Conf. Definition: Table(Player)(~ 90,88,91,82,«null»,90,100,«null»,98,67,98,76,97,76,100,93,«null»,98,78,100,«null»,71,86,«null»,89,79,93,88,83) Nodelocation: 100,160,1 Nodesize: 52,24 Valuestate: 2,36,43,313,451,0,MIDM Constant Injury_score Title: Injury Score Definition: Table(Player)(~ 1,2,0,0,0,3,0,0,3,0,0,6,3,0,0,1,0,0,0,0,0,1,0,0,0,0,1,0,0) Nodelocation: 100,216,1 Nodesize: 52,24 Valuestate: 2,515,11,360,558,0,MIDM Variable Rc_pre_conf_vs__inj Title: rc pre-conf vs. inj Definition: RankCorrel( Pre_season_self_conf,Injury_score, Player ) Nodelocation: 240,104,1 Nodesize: 48,24 Valuestate: 2,499,272,416,303,0,MIDM Function Rc_p_value(x,y : ContextSamp[I] ; I : Index=Run) Title: rc p value Definition: var n := Sum( x<>null and y<>null, I );~ var rc := RankCorrel(x,y,I);~ var z := 0.4856 * ln( (1+rc)/(1-rc) ) * Sqrt(n-3);~ CumNormal(z) Nodelocation: 104,320,1 Nodesize: 48,24 Windstate: 2,506,216,476,224 Paramnames: x,y,I Variable P_value_for_mono_rel Title: p-Value for mono relation exists Definition: Rc_p_value(Pre_season_self_conf,Injury_score,Player) Nodelocation: 240,217,1 Nodesize: 48,40 Valuestate: 2,557,238,416,303,0,MIDM Numberformat: 2,%,4,2,0,0,4,0,$,0,"ABBREV",0 Function Rc_dist(dataIndex : Index ; sampleRc : scalar) Title: Rc dist Description: Estimate the distribution for the underlying rank correlation given the measured sample rc and the number of samples (from the data index). Definition: If IsSampleEvalMode then (~ var x := BiNormal(0,1,xy,sampleRc, over:DataIndex);~ RankCorrel( x[@xy=1], x[@xy=2], dataIndex )~ ) else ~ sampleRc Nodelocation: 248,320,1 Nodesize: 48,24 Windstate: 2,521,215,476,309 Paramnames: dataIndex,sampleRc Index Xy Title: xy Definition: ['x','y'] Nodelocation: 248,384,1 Nodesize: 48,24 Objective Underlying_rc Title: Underlying RC Definition: Rc_dist(Player, Rc_pre_conf_vs__inj) Nodelocation: 368,104,1 Nodesize: 48,24 Valuestate: 2,356,182,575,339,1,PDFP Objective Prob_rc_0 Title: Prob rc>0 Definition: Probability(Underlying_rc>0) Nodelocation: 368,184,1 Nodesize: 48,24 Valuestate: 2,495,251,416,303,0,MIDM Numberformat: 2,%,4,2,0,0,4,0,$,0,"ABBREV",0 Objective Prob_independent Title: Prob independent Definition: Probability( abs(Underlying_rc) < 0.05 ) Nodelocation: 464,184,1 Nodesize: 48,24 Valuestate: 2,564,263,416,303,0,MIDM Numberformat: 2,%,4,2,0,0,4,0,$,0,"ABBREV",0 Objective Prob_rc__0_3 Title: Prob rc<-0.3 Definition: Probability(Underlying_rc < -0.3) Nodelocation: 368,248,1 Nodesize: 48,24 Valuestate: 2,551,226,416,303,0,MIDM Numberformat: 2,%,4,2,0,0,4,0,$,0,"ABBREV",0 Close Soccer_data Module Synthetic_dataset Title: Synthetic Dataset Author: Lonnie Date: Thu, Mar 25, 2010 9:16 AM Defaultsize: 48,24 Nodelocation: 200,56,1 Nodesize: 48,24 Diagstate: 2,235,36,550,276,17 Index Test_subject Title: Test Subject Definition: 1..25 Nodelocation: 96,48,1 Nodesize: 48,24 {!40000|Att_previndexvalue: [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25]} Variable Data_x Title: data x Definition: Table(Test_subject)(~ 6.6,5.3,7.8,5.9,3.6,3.7,8.4,9.699999999999999,3.6,3.6,10.2,7.3,5.9,4.2,3.4,9.1,8.1,2.3,7.5,4.2,8.9,3.4,6.6,3.9,5.4) Nodelocation: 96,104,1 Nodesize: 48,24 Variable Data_y Title: data y Definition: Table(Test_subject)(~ 5.9,8.300000000000001,8.800000000000001,2.5,0.8,5.6,11.6,8.5,0.1,9.4,11.3,10.7,10.4,3.1,1.3,0.1,6.1,1,3.1,8.1,5.6,9.1,0.7,3,0.6) Nodelocation: 96,160,1 Nodesize: 48,24 Defnstate: 2,268,246,416,303,0,MIDM Valuestate: 2,472,250,445,303,1,MIDM Graphsetup: {!40000|Att_contlinestyle Graph_primary_valdim:4} Numberformat: 2,D,4,2,0,0,4,0,$,0,"ABBREV",0 Objective Data_scatter Title: data scatter Definition: [data_y,data_x] Nodelocation: 96,224,1 Nodesize: 48,24 Valuestate: 2,105,191,465,349,1,MIDM Graphsetup: {!40000|Att_contlinestyle Graph_primary_valdim:4} Reformval: [Test_subject,Undefined] {!40000|Att_xrole: -2} {!40000|Att_yrole: -1} {!40000|Att_coordinateindex: Self} Variable Rc_xy Title: rc xy Definition: RankCorrel(data_x,data_y, Test_subject ) Nodelocation: 224,128,1 Nodesize: 48,24 Valuestate: 2,505,270,416,303,0,MIDM Close Synthetic_dataset Library Multivariate_distrib Title: Multivariate Distributions Description: A library of multivariate distributions.~ ~ In a multivariate distribution, each sample is a vector. This vector is identified by an index, identified by the I parameter of the functions in this library. A Mid value from a distribution function will therefore be indexed by I, whlie a Sample from a distribution function is indexed by both I and Run. These distribution functions can also be used from within the Random function to generate a single monte-carlo sample, which will be indexed by I.~ ~ This library also contains functions for generating correlated distributions. Correlate_with, for example, allows you to generate a univarite distribution with an arbitrary marginal distribution that has a specified rank correlation with an arbitrary reference distribution. Several functions may be used for generating serial correlations, where each distribution along an index is correlated with the previous point along that index. Author: Lonnie Chrisman, Ph.D.~ Lumina Decision Systems~ ~ With contributions by:~ John Bowers, US FDA.~ Max Henrion, Lumina Decision Systems Date: Fri, Aug 01, 2003 7:12 PM Saveauthor: Lonnie Savedate: Tue, Nov 11, 2008 1:59 PM Defaultsize: 48,24 Nodelocation: 328,56,1 Nodesize: 56,24 Nodeinfo: 1,1,1,1,1,1,0,0,0,0 Diagstate: 1,42,10,649,1009,17 Windstate: 2,401,199,483,316 Fontstyle: Arial, 15 Function Wishart( cv : Number[I,J,Run] ; n :positive ; I,J : Index ; ~ singleSampleMethod : optional hidden scalar) Title: Wishart(cv,n,I,J) Description: Suppose you sample N samples from a Gaussian(0,cv,I,J) distribution, X[I,R]. (R is the index that indexes each sample, R:=1..N). The Wishart distribution describes the distribution of sum( X * X[I=J], R ). This matrix is dimensioned by I and J and is called the scatter matrix. ~ ~ A sample drawn from the Wishart is therefore a sample scatter matrix. If you divide that sample by (N-1), you have a sampled covariance matrix. ~ ~ If you compute a sample covariance matrix from data, and then want to use this in your model, if you just use it directly, you'll be ignoring sampling error. That may be insignificant of N is large. Otherwise, you may want to use:~ Wishart( SampleCV, N, I, J) / (N-1)~ instead of just SampleCV in your model. The extended variance will account for the uncertainty from the finite sample size that was used to obtain your sample CV.~ ~ If you can express a prior probability on covariances in the form of an InvertedWishart distribution, then the posterior distribution, after having computed the sample covariance matrix (assumed to be drawn, by nature, from a Wishart), is also an InvertedWishart. Definition: var T := if i