Monday, October 7, 2013

Regression, Part 2


Last time we ran a regression analysis, this time we will look at the output and interpretation. Your results should look like the following:

Descriptive Statistics
 
 
Mean
Std. Deviation
N
 
perceived stress
69.60
17.063
10
 
Time to complete exam
43.20
12.925
10
 
exam grade
82.10
12.879
10
 
 
Model Summaryb
Model
R
R Square
Adjusted R Square
Std. Error of the Estimate
1
.884a
.781
.718
9.056
a. Predictors: (Constant), exam grade, Time to complete exam
b. Dependent Variable: perceived stress
 
Correlations
 
perceived stress
Time to complete exam
exam grade
Pearson Correlation
perceived stress
1.000
-.868
-.874
Time to complete exam
-.868
1.000
.943
exam grade
-.874
.943
1.000
Sig. (1-tailed)
perceived stress
.
.001
.000
Time to complete exam
.001
.
.000
exam grade
.000
.000
.
N
perceived stress
10
10
10
Time to complete exam
10
10
10
exam grade
10
10
10
 
ANOVAa
Model
Sum of Squares
df
Mean Square
F
Sig.
1
Regression
2046.264
2
1023.132
12.474
.005b
Residual
574.136
7
82.019
 
 
Total
2620.400
9
 
 
 
a. Dependent Variable: perceived stress
b. Predictors: (Constant), exam grade, Time to complete exam
 
Coefficientsa
Model
Unstandardized Coefficients
Standardized Coefficients
t
Sig.
95.0% Confidence Interval for B
Correlations
Collinearity Statistics
B
Std. Error
Beta
Lower Bound
Upper Bound
Zero-order
Partial
Part
Tolerance
VIF
1
(Constant)
146.784
31.053
 
4.727
.002
73.354
220.214
 
 
 
 
 
Time to complete exam
-.518
.702
-.393
-.739
.484
-2.177
1.141
-.868
-.269
-.131
.111
9.025
exam grade
-.667
.704
-.504
-.948
.375
-2.332
.998
-.874
-.337
-.168
.111
9.025
a. Dependent Variable: perceived stress
There are many aspects that can be checked, based on the analyses we have run; however, I do not have the space to review them all. Please see Pallant (2013) for an in-depth discussion of them. 

Let's evaluate our model. Look in the Model Summary box and check the value under the heading R Square. This tells you how much of the variance in the DV (stress) is explained by the model (which includes the IVs exam time and grades). In this case, the value is .781 (see yellow highlight), so we can say that the model explains 78.1% of the variance in perceived stress. We had a very small sample, however, so it is best to use the adjusted R square .718 or 71.8%, which is a better estimate. To assess the statistical significance of the result, we need to look at the table labeled ANOVA. This tests the null hypothesis that multiple R in the population equals 0. In our example, the model reaches statistical significance of .005 (see blue text).  

Next, take a look at the table of Coefficients and the column labeled Beta. Ignoring any negative signs we can see that exam grade made the largest contribution (.504) to explaining the DV, when the variance explained by all other variables are controlled. The Beta value for exam time was slightly lower (.393) indicating it made less of a contribution (see red text) 

The results of the analyses allow us to the answer the two questions we posed at the beginning. The model, which includes the time to complete the exam and grade, explains 71.8% of the variance in perceived stress. Of these two variables, exam grade makes the largest contribution (beta = -.504), although exam time also made a statistically significant contribution (beta = -.393).

Next time we will look at the formation of research questions. Do you have an issue or a question that you would like me to discuss in a future post? Would you like to be a guest writer? Send me your ideas! Send me an email with your ideas. leann.stadtlander@waldenu.edu 

Pallant, J. (2013). The SPSS Survival Manual, 5th edition. Open University Press.

Friday, October 4, 2013

Regression, Part 1


Jamie asks: When you did blogs on several of the statistical analyses, you did not do one on regression. Would that be one you would consider blogging about? I've been getting stuck on that.  

Of course, Jamie! I am going to assume that you are interested in multiple regression. This is a little more complicated to try to cover in a blog post, but let's give it a try! Multiple regression is a more sophisticated extension of correlation and is used when you want to explore the predicative ability of a set of independent variables (IV) on one continuous dependent measure (DV). There are different types of multiple regression that allow you to compare the predictive ability of particular independent variables ad find the best set of variables to predict a dependent variable. I will be looking at a standard multiple regression. 

An example might be the time to complete an exam (IV1), the person's grade on the exam (IV2), and perceived stress (DV). Our research questions are – 1) How well does time on the exam and grade on exam predict perceived stress? How much variance in perceived stress scores can be explained by scores on these two IVs? 2) Which is the best predictor of perceived stress: time or grade? 

Let's do an example of a multiple regression together. So open SPSS, first go to Edit on the menu, select Options and make sure there is a check in the box No scientific notation for small numbers in tables. Enter the following data for your sample: 

Under Variable view (see tab at bottom of page), It should look like: 

Name
Type
Width
Decimals
Label
Values
Ignore the rest
Examtime
numeric
8
0
Time to complete exam
None
Ignore the rest
Grade
numeric
8
0
Exam grade
None
Ignore the rest
Stress
numeric
8
0
Perceived stress
None
Ignore the rest

 
Go back to Data View and enter the following: 

Examtime
Grade
Stress
20
63
85
45
89
65
36
75
82
59
92
45
56
96
50
27
66
90
39
70
77
52
89
70
43
82
85
55
99
47

 
Go to Analyze/ Regression/Linear. Move your continuous DV (stress) into the Dependent box. 
Move your IVs (exam time and grade) into the independent box 
For Method, make sure Enter is selected.
Click on the Statistics button
             Select the following: Estimates, Confidence Intervals, Model fit, Part and partial correlations, and Collinearity diagnostics
             In the Residuals section, select Casewise diagnostics and Outliers outside 3 standard deviations. Click on Continue.
Click on the Plots button
             Click on *XRESID and move to the Y box
             Click on *ZPRED and move to the X box
In the section labeled Standardized Residual Plots, tick the normal probability plot option. Click on continue
Click on the Save button
             In the section labeled Distances, select Mahalanobis box and Cook's
Click on Continue and then OK 

Next time we will look at the output and interpretation. Do you have an issue or a question that you would like me to discuss in a future post? Would you like to be a guest writer? Send me your ideas! Send me an email with your ideas. leann.stadtlander@waldenu.edu