sábado, 27 de febrero de 2016

Assignment 3: Pearson Correlation

1.- Research question:

I want to find ou whether we can predict the number of suicides in a given country, knowing the % of urban population it has.

For that matter I am going to use the following variables from the GapMinder dataset:

  • Urban Rate
  • Suicides per 100K people
2.- Code used (SAS)


LIBNAME mydata "/courses/d1406ae5ba27fe300 " access=readonly;
DATA new; set mydata.gapminder;

PROC SORT; BY COUNTRY;

PROC CORR; VAR urbanrate suicideper100TH;

RUN;

3.- Results:

2 Variables:urbanrate suicideper100th
Estadísticos simples
VariableNMediaDev stdSumaMínimoMáximoEtiqueta
urbanrate20356.7693623.844931152410.40000100.00000urbanrate
suicideper100th1919.640846.3001818410.2014535.75287SUICIDEPER100TH
Coeficientes de correlación Pearson
Prob > |r| suponiendo H0: Rho=0
Número de observaciones
 urbanratesuicideper100th
urbanrate
urbanrate
1.00000
 
203
-0.13038
0.0761
186
suicideper100th
SUICIDEPER100TH
-0.13038
0.0761
186
1.00000
 
191

4.- Interpretation:

There is almost no linear relation between the Urban Rate of a country and its number of suicides. We cannot make predictions.

Even if the r value was strong enough for us to make predictions, the p value of 0,0761 (higher than 0.05) does not allow us to refuse the null hypothesis which claims there is no relation vetween the variables.

domingo, 21 de febrero de 2016

Assingment 2: Chi Square Analysis

Coursera: Data analysis Tools
Week 2 assignment: Chi Square analysis

1.- Analysis  Question:

Does money really bring happiness? At least what I am willing to discover is whether individuals from a richer or poorer household are more or less likely to suffer from depression sympthoms.
For that matter, we are going to take into account two variable.
·         A numerical explanatory variable; “Household income per year”, which I have turned into a categorical one by making Average Household Income Groups (From now on, HHIG)
·         A categorical response variable, S4AQ1, which tells us if the individual has had a period of +2weeks of sadness, being:
o   1: The individual HAS suffered from these sympthoms
o   2: The individual HAS NOT suffered from these sypmthoms 




2.- Code Used:

LIBNAME mydata "/courses/d1406ae5ba27fe300 " access=readonly;
DATA new; set mydata.nesarc_pds;

/*Names for the variables*/
LABEL S1Q12B="Average household income group"
S4CQ1="Had a +2weeks low mood period";

/*Set appropriate missing data as needed*/
IF S4CQ1=9 THEN S4CQ1=.;

/*Creating the average household income groups*/
IF S1Q12B=1 OR S1Q12B=2 OR S1Q12B=3 OR S1Q12B=4 OR S1Q12B=5 OR S1Q12B=6 OR S1Q12B=7 OR S1Q12B=8 THEN HHIG=1;
ELSE IF S1Q12B=9 OR S1Q12B=10 OR S1Q12B=11 OR S1Q12B=12 OR S1Q12B=13 THEN HHIG=2;
ELSE IF S1Q12B=14 OR S1Q12B=15 OR S1Q12B=16 THEN HHIG=3;
ELSE IF S1Q12B=17 OR S1Q12B=18 OR S1Q12B=19 OR S1Q12B=20 OR S1Q12B=21 THEN HHIG=4;
/* HHIG is the average household income group
1.- 0 to 29.999$ - Low/ Mid-Low
2.- 30.000 to 69.999$ - Medium
3.- 70.000 to 99.999$ - High
4.- 100.000$ and more - Very High */

/*Starting the Chi Square anlysis*/
PROC FREQ; TABLES S4CQ1*HHIG/CHISQ;

RUN;

3.- Results and interpretation:

Procedimiento FREQ
Frecuencia
Porcentaje
Pct fila
Pct col
Tabla de S4CQ1 por HHIG
S4CQ1(Had a +2weeks low mood period)
HHIG
1
2
3
4
Total
1
1237
2.96
52.89
7.00
780
1.86
33.35
4.90
186
0.44
7.95
4.20
136
0.32
5.81
3.54
2339
5.59


2
16433
39.27
41.59
93.00
15132
36.16
38.30
95.10
4239
10.13
10.73
95.80
3706
8.86
9.38
96.46
39510
94.41


Total
17670
42.22
15912
38.02
4425
10.57
3842
9.18
41849
100.00
Frecuencia de ausentes = 1244
Estadísticos para la tabla de S4CQ1 por HHIG
Estadístico
DF
Valor
Prob
Chi-cuadrado
3
127.6301
<.0001
Chi-cuadrado de ratio de verosimilitud
3
129.3039
<.0001
Chi-cuadrado Mantel-Haenszel
1
113.1129
<.0001
Coeficiente Phi

0.0552

Coeficiente de contingencia

0.0551

V de Cramer

0.0552

Tamaño efectivo de la muestra = 41849
Frecuencia de valores ausentes = 124

Interpretation: The P value of less than 0.001 tells us that the differencces are statistically relevant, thus the conclusion being that, the higher the income of a household, the less likely to suffer from depression its dwellers are.

3.- Post HOC analyses:

We know that there are representative differences between groups, but we still need to learn which groups are really different.

As we have 6 possible impairments, the BONFERRONI ADJUSTMENT gives us an “Adjusted p value” of  0,00833 to use as a benchmark to compare against the p values for each individual comparison.

4.- Post HOC  analyses code:

DATA COMPARISON1; SET NEW;
IF HHIG=1 OR HHIG=2;
PROC SORT; BY IDNUM;
PROC FREQ; TABLES S4CQ1*HHIG/CHISQ;
RUN;

DATA COMPARISON1; SET NEW;
IF HHIG=1 OR HHIG=3;
PROC SORT; BY IDNUM;
PROC FREQ; TABLES S4CQ1*HHIG/CHISQ;
RUN;

DATA COMPARISON1; SET NEW;
IF HHIG=1 OR HHIG=4;
PROC SORT; BY IDNUM;
PROC FREQ; TABLES S4CQ1*HHIG/CHISQ;
RUN;

DATA COMPARISON1; SET NEW;
IF HHIG=2 OR HHIG=3;
PROC SORT; BY IDNUM;
PROC FREQ; TABLES S4CQ1*HHIG/CHISQ;
RUN;

DATA COMPARISON1; SET NEW;
IF HHIG=2 OR HHIG=4;
PROC SORT; BY IDNUM;
PROC FREQ; TABLES S4CQ1*HHIG/CHISQ;
RUN;

DATA COMPARISON1; SET NEW;
IF HHIG=3 OR HHIG=4;
PROC SORT; BY IDNUM;
PROC FREQ; TABLES S4CQ1*HHIG/CHISQ;
RUN;

5.- Post HOC analyses results and interpretation

Procedimiento FREQ
Frecuencia
Porcentaje
Pct fila
Pct col
Tabla de S4CQ1 por HHIG
S4CQ1(Had a +2weeks low mood period)
HHIG
1
2
Total
1
1237
3.68
61.33
7.00
780
2.32
38.67
4.90
2017
6.01


2
16433
48.93
52.06
93.00
15132
45.06
47.94
95.10
31565
93.99


Total
17670
52.62
15912
47.38
33582
100.00
Frecuencia de ausentes = 1052
Estadísticos para la tabla de S4CQ1 por HHIG
Estadístico
DF
Valor
Prob
Chi-cuadrado
1
65.3157
<.0001
Chi-cuadrado de ratio de verosimilitud
1
66.0145
<.0001
Chi-cuadrado adj. de continuidad
1
64.9445
<.0001
Chi-cuadrado Mantel-Haenszel
1
65.3138
<.0001
Coeficiente Phi

0.0441

Coeficiente de contingencia

0.0441

V de Cramer

0.0441


Test exacto de Fisher
Celda (1,1) Frecuencia (F)
1237
Alineado a la izquierda Pr <= F
1.0000
Alineado a la derecha Pr >= F
<.0001


Tabla de probabilidad (P)
<.0001
De dos caras Pr <= P
<.0001
Tamaño efectivo de la muestra = 33582
Frecuencia de valores ausentes = 1052

Procedimiento FREQ
Frecuencia
Porcentaje
Pct fila
Pct col
Tabla de S4CQ1 por HHIG
S4CQ1(Had a +2weeks low mood period)
HHIG
1
3
Total
1
1237
5.60
86.93
7.00
186
0.84
13.07
4.20
1423
6.44


2
16433
74.37
79.49
93.00
4239
19.19
20.51
95.80
20672
93.56


Total
17670
79.97
4425
20.03
22095
100.00
Frecuencia de ausentes = 721
Estadísticos para la tabla de S4CQ1 por HHIG
Estadístico
DF
Valor
Prob
Chi-cuadrado
1
45.9511
<.0001
Chi-cuadrado de ratio de verosimilitud
1
50.5550
<.0001
Chi-cuadrado adj. de continuidad
1
45.4881
<.0001
Chi-cuadrado Mantel-Haenszel
1
45.9490
<.0001
Coeficiente Phi

0.0456

Coeficiente de contingencia

0.0456

V de Cramer

0.0456


Test exacto de Fisher
Celda (1,1) Frecuencia (F)
1237
Alineado a la izquierda Pr <= F
1.0000
Alineado a la derecha Pr >= F
<.0001


Tabla de probabilidad (P)
<.0001
De dos caras Pr <= P
<.0001
Tamaño efectivo de la muestra = 22095
Frecuencia de valores ausentes = 721

Procedimiento FREQ
Frecuencia
Porcentaje
Pct fila
Pct col
Tabla de S4CQ1 por HHIG
S4CQ1(Had a +2weeks low mood period)
HHIG
1
4
Total
1
1237
5.75
90.09
7.00
136
0.63
9.91
3.54
1373
6.38


2
16433
76.39
81.60
93.00
3706
17.23
18.40
96.46
20139
93.62


Total
17670
82.14
3842
17.86
21512
100.00
Frecuencia de ausentes = 713
Estadísticos para la tabla de S4CQ1 por HHIG
Estadístico
DF
Valor
Prob
Chi-cuadrado
1
63.2565
<.0001
Chi-cuadrado de ratio de verosimilitud
1
72.3478
<.0001
Chi-cuadrado adj. de continuidad
1
62.6786
<.0001
Chi-cuadrado Mantel-Haenszel
1
63.2535
<.0001
Coeficiente Phi

0.0542

Coeficiente de contingencia

0.0541

V de Cramer

0.0542


Test exacto de Fisher
Celda (1,1) Frecuencia (F)
1237
Alineado a la izquierda Pr <= F
1.0000
Alineado a la derecha Pr >= F
<.0001


Tabla de probabilidad (P)
<.0001
De dos caras Pr <= P
<.0001
Tamaño efectivo de la muestra = 21512
Frecuencia de valores ausentes = 713

Procedimiento FREQ
Frecuencia
Porcentaje
Pct fila
Pct col
Tabla de S4CQ1 por HHIG
S4CQ1(Had a +2weeks low mood period)
HHIG
2
3
Total
1
780
3.84
80.75
4.90
186
0.91
19.25
4.20
966
4.75


2
15132
74.41
78.12
95.10
4239
20.84
21.88
95.80
19371
95.25


Total
15912
78.24
4425
21.76
20337
100.00
Frecuencia de ausentes = 531
Estadísticos para la tabla de S4CQ1 por HHIG
Estadístico
DF
Valor
Prob
Chi-cuadrado
1
3.7344
0.0533
Chi-cuadrado de ratio de verosimilitud
1
3.8393
0.0501
Chi-cuadrado adj. de continuidad
1
3.5816
0.0584
Chi-cuadrado Mantel-Haenszel
1
3.7342
0.0533
Coeficiente Phi

0.0136

Coeficiente de contingencia

0.0135

V de Cramer

0.0136


Test exacto de Fisher
Celda (1,1) Frecuencia (F)
780
Alineado a la izquierda Pr <= F
0.9769
Alineado a la derecha Pr >= F
0.0280


Tabla de probabilidad (P)
0.0049
De dos caras Pr <= P
0.0551
Tamaño efectivo de la muestra = 20337
Frecuencia de valores ausentes = 531

Procedimiento FREQ
Frecuencia
Porcentaje
Pct fila
Pct col
Tabla de S4CQ1 por HHIG
S4CQ1(Had a +2weeks low mood period)
HHIG
2
4
Total
1
780
3.95
85.15
4.90
136
0.69
14.85
3.54
916
4.64


2
15132
76.60
80.33
95.10
3706
18.76
19.67
96.46
18838
95.36


Total
15912
80.55
3842
19.45
19754
100.00
Frecuencia de ausentes = 523
Estadísticos para la tabla de S4CQ1 por HHIG
Estadístico
DF
Valor
Prob
Chi-cuadrado
1
12.9852
0.0003
Chi-cuadrado de ratio de verosimilitud
1
13.8344
0.0002
Chi-cuadrado adj. de continuidad
1
12.6790
0.0004
Chi-cuadrado Mantel-Haenszel
1
12.9846
0.0003
Coeficiente Phi

0.0256

Coeficiente de contingencia

0.0256

V de Cramer

0.0256


Test exacto de Fisher
Celda (1,1) Frecuencia (F)
780
Alineado a la izquierda Pr <= F
0.9999
Alineado a la derecha Pr >= F
0.0001


Tabla de probabilidad (P)
<.0001
De dos caras Pr <= P
0.0002
Tamaño efectivo de la muestra = 19754
Frecuencia de valores ausentes = 523

Procedimiento FREQ
Frecuencia
Porcentaje
Pct fila
Pct col
Tabla de S4CQ1 por HHIG
S4CQ1(Had a +2weeks low mood period)
HHIG
3
4
Total
1
186
2.25
57.76
4.20
136
1.65
42.24
3.54
322
3.90


2
4239
51.28
53.35
95.80
3706
44.83
46.65
96.46
7945
96.10


Total
4425
53.53
3842
46.47
8267
100.00
Frecuencia de ausentes = 192
Estadísticos para la tabla de S4CQ1 por HHIG
Estadístico
DF
Valor
Prob
Chi-cuadrado
1
2.4190
0.1199
Chi-cuadrado de ratio de verosimilitud
1
2.4312
0.1189
Chi-cuadrado adj. de continuidad
1
2.2450
0.1340
Chi-cuadrado Mantel-Haenszel
1
2.4187
0.1199
Coeficiente Phi

0.0171

Coeficiente de contingencia

0.0171

V de Cramer

0.0171


Test exacto de Fisher
Celda (1,1) Frecuencia (F)
186
Alineado a la izquierda Pr <= F
0.9468
Alineado a la derecha Pr >= F
0.0668


Tabla de probabilidad (P)
0.0136
De dos caras Pr <= P
0.1241
Tamaño efectivo de la muestra = 8267
Frecuencia de valores ausentes = 192

Interpretation: The groups that are actually different are the following. Meaning people from households with an income of less of 30.000$ are certainly more likely than the rest to suffer from depression.

Pair
Meaning
Statistically Different
1 vs 2
"Low/Mid-low" vs "Medium"
Yes
1 vs 3
"Low/Mid-low" vs "High"
Yes
1 vs 4
"Low/Mid-low" vs "Very High"
Yes
2 vs 3
"Medium" vs " Very High"
No
2 vs 4
"Medium" vs "High"
Yes
3 vs 4
"High" vs " Very High"
No