Learn Before
Sensitivity of the Mean to Outliers
A key characteristic of the mean is its high sensitivity to extreme scores, or outliers. In highly skewed distributions, the mean is pulled away from the median in the direction of the skew (the longer tail). Because extreme scores heavily influence its value, the mean may cease to be an accurate representation of the typical score in such datasets, leading researchers to often prefer the median for skewed distributions.
0
1
Tags
Ch.2 Psychological Research - Psychology @ OpenStax
Psychology @ OpenStax
OpenStax
OpenStax Psychology (2nd ed.) Textbook
Psychology
Social Science
Empirical Science
Science
KPU
Research Methods in Psychology - 4th American Edition @ KPU
Related
Positively Skewed Distribution
Sensitivity of the Mean to Outliers
Negatively Skewed Distribution
Which of the following best describes the shape of a skewed distribution?
A psychology researcher measures the reaction times of participants on a memory task. They find that while most participants respond very quickly (between 200 and 400 milliseconds), a small number of participants take significantly longer (over 1,500 milliseconds). Because the scores cluster at the lower end of the scale and the tail trails off toward the higher values, this distribution is __________ skewed.
A researcher conducting a study on digital literacy finds that most participants are highly proficient and score near the top of a 100-point scale. However, a small number of participants have almost no experience with technology and score very low. This creates a negatively skewed distribution. Analyze the relationship between the measures of central tendency in this scenario and arrange them in order from the lowest numerical value to the highest numerical value.
A researcher finds that a distribution of participant reaction times is heavily positively skewed, with a mean of $1,200 ext{ ms}and a median of $600 ext{ ms}. The researcher justifies reporting the mean as the 'typical' reaction time by arguing that it is the most comprehensive measure because it captures the specific performance of every slow-responding participant in the long tail. This justification represents a scientifically valid evaluation of how to represent the central tendency
In a skewed distribution, the 'tail' refers to the prominent peak where the majority of scores cluster most heavily.
In a skewed distribution, how does the 'tail' relate to the overall shape and clustering of the data?
Match each distribution type to the description that correctly identifies where scores cluster and where the tail extends.
A psychology researcher is analyzing the shapes of different data distributions. Match each research data distribution characteristic to its corresponding pattern of score clustering and tail direction.
When analyzing the shape of a skewed data distribution, a researcher finds a prominent peak where scores cluster heavily toward one end. To identify the direction of the skew, the researcher must locate the relative position of the _____, which trails off toward the opposite end.
Order the steps a researcher must take to evaluate the shape of a dataset's distribution and determine if it is skewed.
Excluding Outliers
Handling Valid Extreme Outliers
Example of an Outlier
Sensitivity of the Mean to Outliers
Defining Outliers using z Scores
Reaction Time Outlier Example
Impact of Outliers on the Range
Identifying Outliers Using z Scores
Handling Outliers
What term is used to describe an extreme score that falls significantly above or below the rest of the scores within a distribution?
In psychological research, an outlier in a dataset always indicates that an error occurred during data collection, such as an equipment malfunction or participant misunderstanding.
In psychological research, outliers can arise from various sources. Match each research scenario with the most likely reason for the extreme score described.
A researcher identifies a score in a dataset that falls significantly below the rest of the distribution. Arrange the following steps in the logical analytical sequence used to investigate and address this extreme value.
You are constructing a research protocol for a study on the cognitive effects of extreme stress. To ensure that your final dataset can systematically distinguish between genuine cases of 'stress resilience' (valid outliers) and 'task confusion' (measurement errors), which design element should you integrate into your data-collection phase?
Sexual Partners Survey Outlier Example
Outlier in Beck Depression Inventory Scores
A researcher is evaluating a dataset where a single participant's score is , while all other scores are between and . After confirming the participant understood the task perfectly and no errors occurred, the researcher decides to retain the score. This decision reflects a judgment that the extreme value is a(n) _____ which, despite its potential to skew the mean, represents a genuine and valid case within the study.
A researcher measures anxiety in 50 participants and finds that 49 scores fall between 20 and 45, but one participant scores 95. This extreme value, which may reflect a genuinely anxious individual or a data collection error, is called a(n) _____.
A researcher administering a depression inventory notices one participant scored far higher than the rest of the sample. If the researcher confirms this participant is indeed clinically depressed and no data entry or equipment errors occurred, this extreme score is no longer considered an outlier.
Match each description of a research scenario that produces an extreme score with its corresponding source of outlier.
When conducting data cleaning on a newly collected psychological dataset, order the stages a researcher should follow to systematically identify and categorize an extreme score.
Define what an outlier is in the context of a dataset's distribution, and describe the two broad categories of sources that can produce these extreme scores in psychological research.
Based on the provided context, diagnose the nature of this student's extreme score. Is this outlier a result of an error, or does it represent a genuine case? Explain your reasoning.
A researcher asks participants to report their age in years. Most responses range from to . One participant enters . Apply the definition of an outlier to explain what this value represents and identify the most likely reason for this extreme score based on common sources of outliers.
Which of the following best describes an outlier?
Because outliers are extreme scores that fall significantly outside the rest of a distribution, a researcher should always assume they are the result of equipment malfunctions, participant misunderstandings, or data entry errors.
An outlier is an extreme score that falls significantly above or below the rest of the distribution. Match each research scenario below to the most likely source of the outlier it describes.
While analyzing a distribution of questionnaire scores, a researcher spots an extreme score that falls significantly above the rest. Upon examining the raw data, the researcher notices the participant selected 'Strongly Agree' for every item, including directly opposing statements such as 'I am always anxious' and 'I am always calm.' This contradictory evidence indicates that the extreme score is not a genuine case, but rather an outlier resulting from a participant ____.
A researcher needs to judge whether a suspicious data point should be included in their analysis. Arrange the following steps in the most logical sequence to systematically evaluate the source of this outlier, from initial observation to final judgment.
While an outlier can result from mistakes, it can also represent a 'genuine case' on the variable being measured. Which of the following is an example of a genuine case?
When a researcher identifies an extreme score in a distribution, why is it important to carefully consider its source rather than immediately discarding it?
During a study on early language acquisition, a researcher finds that 49 out of 50 toddlers have a vocabulary of 20 to 50 words, while one toddler has a verified vocabulary of over 200 words. Because this unusually high score represents a genuine case rather than an equipment malfunction or researcher error, it should not be classified as an outlier.
Researchers must carefully break down the context of extreme scores to identify their source. Match each piece of analytical evidence to the specific source of the outlier it most strongly suggests.
A cognitive psychologist notices an extreme score in a reaction time distribution that falls significantly below the rest of the scores, indicating an incredibly fast response. Without checking the laboratory logs, the researcher immediately deletes the score, stating, 'Extreme outliers are always the result of equipment malfunctions or participant misunderstandings, so this data point is invalid.' How should this researcher's decision be evaluated?
Example of Finding the Median
Sensitivity of the Mean to Outliers
What does the median represent in a distribution of scores?
Learn After
A real estate agent is marketing a neighborhood of 10 homes. Nine of the homes are valued at approximately $200,000 each, while one newly built mansion is valued at $3,200,000. The agent advertises the 'average home value' in the neighborhood as $500,000. Which statement best evaluates the agent's use of this figure?
Example of the Mean's Sensitivity to Outliers
In a highly skewed distribution, how does the presence of extreme scores typically affect the mean?
In a distribution of reaction times, the mean will typically be more significantly shifted by the presence of a few extreme outliers than the median will.
Match each research scenario with the expected relationship between the mean and the median, based on how the presence of extreme scores (outliers) influences the mean.
A researcher is analyzing a dataset of participants' reaction times that is positively skewed due to a few participants who took an exceptionally long time to respond (outliers). Based on the sensitivity of the mean to these extreme scores, arrange the following values in order from the lowest numerical value to the highest numerical value within this specific distribution.
A researcher is developing a reporting protocol for a psychology study on reaction times. One participant's response is times slower than everyone else's. Arrange the following steps in the correct order to create a data-summary protocol that ensures the final report accurately represents the 'typical' participant by accounting for the mean's sensitivity to extreme scores.
In a skewed distribution, the mean is typically pulled away from the median in the direction of the shorter tail of the distribution.
A researcher is examining the weekly study hours of a group of psychology students. Most students study between and hours per week, but two students report studying hours per week. Which statement best explains why the mean of this dataset is not a good representation of the typical student's study hours?
A researcher is evaluating the central tendency of a dataset containing several extreme outliers that have significantly shifted the average value. After reviewing the data, the researcher concludes that the mean is a(n) _____ representation of the typical score because its high sensitivity to those outliers causes it to no longer accurately reflect the center of the distribution.
Analyze the impact of different distribution characteristics on central tendency measures by matching each concept with the statement that best describes its behavior or role in a dataset.
A researcher is evaluating a reaction time dataset and finds that a few extremely long response times have created a highly skewed distribution. The researcher must determine which measure of central tendency provides the most accurate representation of the typical score. Because the mean is pulled away from the center by these extreme scores, the researcher decides that the mean is not an accurate representation and evaluates that they should report the _____ instead.
Why do researchers often prefer the median over the mean when describing the typical score of a highly skewed distribution?
Because the mean utilizes every data point in a psychological dataset, it is always the most accurate representation of the typical score, even when the distribution contains extreme outliers.
A psychology researcher is analyzing different datasets from a recent study. Match each dataset scenario with the most appropriate statistical decision regarding its measure of central tendency.
A cognitive psychology researcher is analyzing three datasets of reaction times to understand how measures of central tendency behave under different conditions. Arrange the following datasets in order from the one where the mean is LEAST pulled away from the median to the one where the mean is MOST drastically pulled away from the median.
A peer reviewer is evaluating a research manuscript which claims that participants typically required fifteen trials to learn a maze, based entirely on the mean score. However, the data shows a severe positive skew: most participants learned it in four to six trials, but two participants required over eighty trials. To ensure the statistical summary is logically defensible and not distorted by extreme outliers, the reviewer should demand that the authors report the _____ instead.
In a highly skewed distribution, the mean is pulled away from the median in the direction of the _____.
A psychology researcher calculates both the mean and the median for a dataset of reaction times. If the dataset contains a few unusually slow reaction times (extreme high values), how will these outliers affect the relationship between the two measures?
A developmental psychology researcher observes the number of words spoken by toddlers in a 10-minute play session. Most toddlers speak between 15 and 20 words, but one highly verbal toddler speaks 110 words. If the researcher wants to describe the typical vocabulary output of the group without the summary being distorted by the extreme score, they should use the mean as their primary measure of central tendency.
A psychology researcher is analyzing how different dataset characteristics affect measures of central tendency. Match each dataset scenario with the resulting structural relationship between its mean and median.
A research committee is critiquing a draft report that incorrectly uses the mean to describe a psychological dataset containing extreme outliers. Arrange the logical sequence of arguments the committee should use to justify revising the report to use the median instead.