Friday, April 14, 2017

Making Data Make Sense: Extreme Scores

What are extreme scores? They are scores far outside the norm for a variable or population, leading to the conclusion they are not part of your true population and probably do not belong in your analyses. A common operational definition for extreme scores is +/-3 standard deviations (SDs) from the mean.

Recall the standard normal distribution of a population has 68.26% of the population between +1 and -1 SD of the mean (see diagram: [34.13% between 0 to +1 SD] + [34.13% between 0 and -1 SD] = 68.26%).



So 95.44% of the population should fall between ±2 SD from the mean (34.13% + 34.13% + 13.59% +13.59% = 95.44%), and 99.74% of the population should fall within ±3 SD of the mean. In other words, the probability of randomly sampling an individual more than ±3 SD from the mean in a normally distributed population is 0.26% (.0026), which gives good justification for considering scores outside ±3 SD as suspect. The concern is these extreme scores are not part of the population of interest in your study. 

Next time we will consider the effects of extreme scores. Do you have an issue or a question that you would like me to discuss in a future post? Would you like to be a guest writer? Send me your ideas! leann.stadtlander@waldenu.edu

Wednesday, April 12, 2017

Missing Data as a Variable/ Best Practices

You may wish to examine missing data as an outcome itself, as there may be information in the missing-ness. The act of failing to respond vs. responding might be of interest. This can be examined through a "dummy variable," a new variable you create representing whether a person has missing data or not on a particular variable. You can then do some analyses to see if there are any relationships that develop.

Osborne (2013) provides some best practices in dealing with missing data, which are great to remember.
•            First, do no harm; be careful in your methodology to minimize missing data.

•            Be transparent. Report any incidence of missing data (rates by variable, and reason for missing data if known). This can be important information for readers.

•            Explicitly discuss whether data are missing at random (i.e., if there are differences between individuals with complete and incomplete data).

•            Discuss how you, as the researcher, dealt with the issue of incomplete data. 

Next time we will consider extreme scores. Do you have an issue or a question that you would like me to discuss in a future post? Would you like to be a guest writer? Send me your ideas! leann.stadtlander@waldenu.edu

Monday, April 10, 2017

Dealing with Missing Data

How can you deal with missing data? SPSS offers pairwise deletion, which means only those cases with complete data are included in the analysis. If you have few missing data and they are a result of randomness, then such a plan may be acceptable. However, if they are not randomly missing, you could introduce biases.

A second commonly used method is substituting the overall sample's mean for the missing data. The logic of this is that in absence of any other information, the sample's mean is the best representation of an individual's score. If only a few scores are missing, then this may be an acceptable alternative. However, keep in mind the more scores that are replaced, the more you are biasing the sample to the mean.

A third alternative is given by Osborne (2000, 2013) in which a prediction equation is developed through multiple regression. If you have quite a few missing scores, you may want to explore this alternative. 

Next time we will consider best practices and missing data. Do you have an issue or a question that you would like me to discuss in a future post? Would you like to be a guest writer? Send me your ideas! leann.stadtlander@waldenu.edu

Friday, April 7, 2017

Categories of Missing Data

There are two categories of missing data: data missing at random and data missing not at random. If data are missing randomly, we can assume that they will not bias not the results. However, data missing not at random may be a strong biasing influence.

Let us use an example from Osborn (2013), of an employee satisfaction survey given to schoolteachers. The teachers are surveyed twice, once in September and once in June. Missing random data would mean data missing in June had no relationship to any variable from the September survey (such as satisfaction in Sept., age, and years of teaching). An example, might be if we randomly selected 50% of the people who responded in September to again complete the survey in June, we would legitimately be missing half of the data in June (the 50% of people we did not ask). The missing data would be random and not related to a specific variable such as satisfaction, age, years teaching.

On the other hand, suppose only teachers who were satisfied responded to the survey in June (i.e., people who were dissatisfied were less likely to respond to the survey). Then the missing data are considered missing not at random and may substantially bias the results. Thus, the June survey would show a higher than expected satisfaction score (because unsatisfied people did not participate). 

Next time we will consider how to deal with your missing data. Do you have an issue or a question that you would like me to discuss in a future post? Would you like to be a guest writer? Send me your ideas! leann.stadtlander@waldenu.edu

Thursday, April 6, 2017

Dissertation Process Webinar

I conducted a one hour webinar on the dissertation process for health psych students. You are welcome to view it!


Wednesday, April 5, 2017

Happy Anniversary!

Today is the 5th anniversary of the dissertation blog! In the past 5 years we have discussed many topics and a book came out of the blog Finding Your Way to a Ph.D.: Advice from the Dissertation Mentor! My thanks to all who have supported the blog.

A special thanks to my dogs, Mandy & Murphy who patiently put up with me following them around with my camera!

Monday, April 3, 2017

Making Data Make Sense: Missing Data

In almost any research study, there will be missing or incomplete data. Missing data can happen for a number of reasons: participants fail to respond to questions, subjects withdraw or quit studies before they are completed, and data entry errors.

The problem with missing data is nearly all statistical techniques assume or require complete data. There can be legitimately missing data; an example might be a survey in which a person is asked if he or she is married, and if so how long. If you are not married, then you would be correct in leaving the "how long" portion of the question blank.

It is also important to realize legitimately missing data can be meaningful. The missing data allows a validity check and may inform the status of an individual. Osborn (2013) provides a great example. In cleaning the data from an adolescent health risk survey, he noticed some individuals indicated on one question they had never used illegal drugs, but later in the survey when asked how many times they used marijuana, indicated an answer greater than 0. Therefore, an answer they should have skipped (or be missing), showed an unexpected number. The author suggests several possible explanations, such as the subject was not paying attention and answered in error. However, a more intriguing possibility is some subjects did not view marijuana as an illegal drug, which is an interesting possibility that could be examined in future research.

One way of dealing with legitimately missing data is making the missing and present data two separate groups. Using the marriage survey example, we could eliminate non-married individuals from a specific analysis when looking at issues related to being married vs. not married. So instead of asking the silly research question, "How long, on average, do all people, even unmarried people, stay married;" we can ask two more refined questions: "What are the predictors of whether someone is currently married?" and "Of those who are currently married, how long on average have they been married?" 

Next time we will consider categories of missing data. Do you have an issue or a question that you would like me to discuss in a future post? Would you like to be a guest writer? Send me your ideas! leann.stadtlander@waldenu.edu