--- title: "Vol 4 3" book: "EDUCC 113Methods and Techniques of Educational Research" category: "EDUCC" publisher: "Ratan Prakashan Mandir Pvt. Ltd." type: "Educational Material" --- According to Latest Syllabus Read For Sure Success In University Examination RATAN TEXT BOOK METHODS AND TECHNIQUES OF EDUCATIONAL RESEARCH Vol-4 M.A.Education (Sem-IV) Dr. Mohini Agrawal Published by Ratan Prakashan Mandir Pvt. Ltd. 2nd Floor, Centre Plaza, Parinay Kunj, Lajpat Kunj Marg, Agra-282002 Copyright Authors & Publishers Published by Ratan Prakashan Mandir Pvt. Ltd. 2nd Floor, Centre Plaza, Parinay Kunj, Lajpat Kunj Marg, Agra-282002 ISBN :978-93-0970-215-0 Price 180.00 only Printed at : KIDS INTERNATIONAL PVT. LTD. C-60, 61, 62, 63, EPIP, Shastripuram, Agra - 282007 Ph. : +91 9719004921 UNIT: 16 QUANTITATIVE DATA, NATURE OF QUANTITAIVE DATA AND SCALES OF MEASUREMENT 16.1    Introduction 16.2    Learning Objectives 16.3    Quantitative Data, Nature of Quantitative data and scales of measurement a.    Nominal scale Self- Check Exercise -1 b.    Ordinal Scale Self-Check Exercise-2 c.    Interval Scale Self- Check Exercise-3 d.    Ratio Scale Self- Check Exercise-4 16.4    Summary 16.5    Glossary 16.6    Answers to Self - Check Exercise 16.7    References/ Suggestive Readings 16.8    Terminal Questions 16.1    Introduction In this unit, we will explore the fundamentals of quantitative data, including its characteristics, nature, and different scales of measurement. Quantitative data refers to information that can be measured and represented numerically, allowing researchers to apply statistical analysis and draw meaningful conclusions based on the data. 16.2    LEARNING OBJECTIVES: After going through this unit, the students will be able to: 1.    Understand the meaning of quantitative data and nature of quantitative data. 2.    Define the concept of different scales of measurement. 3.    Differentiate between different scales of measurement. 16.3 Quantitative Data, Nature of Quantitative data and scales of measurement: Quantification is the process of using numerical methods to describe observations of materials or characteristics. By establishing a specific portion of the material or characteristic as a measurement standard, researchers can ensure a valid and precise approach to data description. NATURE OF QUANTITATIVE DATA: Measurement is the most accepted method used in quantitative approach. Quantitative data includes numerical values that can be counted as well as measured. Quantitative data deals with amounts and quantities which involves the use of mathematical calculations and the use of statistical analysis. Examples of quantitative data are: Temperature, Height, Weight, Income etc. Quantitative data is mostly used in the field of physics, economics and psychology. There are three properties of Quantitative data which are listed below: a.    The property of order b.    The property of identity c.    The property of additivity A Nominal Scale: The nominal scale is the most basic form of measurement and the least precise method of quantification. It is used to categorize objects or entities based on qualitative differences rather than numerical values. In other words, a nominal scale groups items into distinct categories without any inherent ranking or order. For example, academic titles such as professor, associate professor, assistant professor, instructor, and lecturer, or classifications like male and female, are all examples of nominal data. In statistics and research, nominal scales are used to classify data without implying any hierarchy among categories. Unlike other measurement scales, nominal data do not follow a specific order. Instead, they serve the purpose of differentiation. Nominal data consist of counted data, meaning each individual belongs to only one category within a set, and all members of that set share the same defining characteristic. Examples include nationality, gender, socio-economic status, race, occupation, and religious affiliation. Although nominal scales lack an inherent order, in some cases, simple classification and counting provide a sufficient basis for statistical analysis. SELF CHECK EXERCISE-1 1 .Nominal data are ___________ data. 2 .A nominal scale is a type of scale where data is classified into: a.    Numbers b.    Ranks c.    Categories d.    All of the above An Ordinal Scale: An ordinal scale is a type of measurement scale where data is ranked based on magnitude or intensity, but the gaps between ranks may not be equal. Unlike nominal scales, which only classify data into categories, ordinal scales establish a hierarchy or order among the categories. Ordinal scales follow a serial arrangement, meaning they not only indicate differences between items but also suggest that these differences exist in varying degrees. This allows researchers to rank individuals or items from highest to lowest based on a particular criterion. The ranking is expressed in terms of relative position within a group, such as 1st, 2nd, 3rd, 4th, 5th, and so on. However, ordinal measurements do not have absolute values, and the actual differences between adjacent ranks are not necessarily equal. Although rankings appear evenly spaced, the true gaps between them may vary. This limitation is best understood through specific examples that illustrate how ordinal scales function in real-world scenarios.The following example illustrates this limitation: | Subject | Height In inches | Difference in inches | Rank | |---|---|---|---| | Rahul | 76 | | 1st | | Surinder | 68 | 8 | 2nd | | Bhanu | 66 | 2 | 3rd | | Pyare Lal | 59 | 7 | 4th | | Arun | 58 | 1 | 5th | Some other examples of ordinal scale are: Rankings in competition (1st, 2nd, 3rd) Likert scales in surveys (e.g. Strongly Agree, Agree, Neutral, Strongly Disagree, Disagree) Rating Scales (Poor, Fair, Good, Excellent) SELF CHECK EXERCISE-2 1 .Ordinal scale is a _____________ scale. 2 .Ordinal scale establishes a ____________ among the categories. An Interval Scale: An interval scale is a quantitative measurement scale that maintains both order and equal intervals between values. It uses a standardized unit of measurement to indicate how much of a characteristic is present. For example, the difference between scores of 90 and 91 is the same as the difference between 60 and 61, ensuring consistency in measurement. Many psychological tests and inventories are based on interval scales. With an interval scale, addition and subtraction are valid operations, but multiplication and division are not applicable because the scale lacks a true zero point. Unlike ratio scales, interval scales do not measure the complete absence of a trait. This means that a score of 90 does not indicate twice as much of a characteristic as a score of 45. Despite this limitation, the interval scale offers a significant advantage over nominal and ordinal scales because it allows researchers to quantify the relative amount of a characteristic rather than simply ranking or categorizing data. SELF EXERCISE-3 1.    Interval scale consists of: a.    Unequal intervals between two variables b.    Equal intervals between two variables c.    Both a and b d.    None of these RATIO SCALE: Ratio scales represent the highest level of measurement, offering the most precise way to quantify variables. Unlike other measurement scales, ratio scales allow all four mathematical operations: addition, subtraction, multiplication, and division. A ratio scale shares the equal interval properties of an interval scale but has two additional features: 1.    A true zero point – This means a variable can be completely absent. For example, zero length or zero height indicates a total lack of that characteristic. 2.    Real-number properties – Numbers on a ratio scale can be added, subtracted, multiplied, divided, and expressed in ratio relationships. For instance: o 5 grams is half of 10 grams o 15 grams is three times 5 grams o On a weighing scale, two 1-gram weights balance a 2-gram weight One key advantage in physical sciences is that many variables, such as weight, height, and temperature (in Kelvin), can be measured using ratio scales. In contrast, behavioral sciences often rely on interval scales, which are less precise because many psychological and social traits cannot be measured with a true zero. From Nominal to Ratio Scales: Increasing Precision Measurement scales progress in precision from nominal (least precise) to ratio (most precise). Researchers should aim to use the most accurate and detailed scale available for their variables. Challenges in Behavioral Research In behavioral sciences, many important qualities—such as intelligence, motivation, or personality—are abstract concepts that cannot be directly observed. Instead, researchers must define them based on observable behaviors. For example, intelligence might be measured using scores on an IQ test, but this is only an operational definition—not intelligence itself. Operational definitions come with limitations: •    They can be subjective, leading to disagreements among experts about their validity. •    Quantification does not guarantee accuracy—even when data is numerical, it can still contain ambiguities and inconsistencies. Some critics argue that an overemphasis on quantification in behavioral research can lead to superficial studies. The desire to imitate the precision of physical sciences may push researchers to measure small, easily quantifiable aspects of behavior, sometimes missing the bigger picture of real human experiences. The Role of Quantitative Methods in Scientific Progress Despite these challenges, quantitative research remains essential. Over time, researchers continue to improve operational definitions and observation techniques, making measurement more valid and reliable. Quantification is indispensable in most fields of research, playing a crucial role in science’s evolution from philosophical speculation to empirical, verifiable knowledge. SELF CHECK EXERCISE-4 1 .Which of the following scales of measurement possesses a true zero point? a.    Interval scale b.    Ordinal scale c.    Nominal scale d.    Ratio scale 16.4    SUMMARY: In this lesson we have learnt about quantitative data, nature of quantitative data and scales of measurement. Scales of measurement play an important role in quantitative data and these scales provide a framework which in turn is useful for analyzing data and classification of data. 16.5    GLOSSARY Arrangement – Arrangement refers to the organization or placement of items, ideas or events in a particular order or pattern. Properties - Properties refer to the characteristics or attributes that define an object, substance or concept. Statistical Analysis- Statistical analysis is the process of collecting, processing, analyzing and interpreting nume4rical data to uncover patterns, trends, relationships or insights within a dataset. 16.6    ANSWERS TO SELF CHECK EXERCISES SELF CHECK EXERCISE-1 Answer 1. Counted Answer 2. Categories. SELF CHECK EXERCISE-2 Answer 1. Measurement Answer 2. Hierarchy/order SELF CHECK EXERCISE-3 Answer 1. B SELF CHECK EXERCISE-4 Answer 1. D 16.7 REFERENCES/ SUGGESTIVE READINGS: Anastasi, A (1970) Psychological Testing London McMillan Best, John W. and Kahn, James V. (2004). Research in Education (7 Ed) New Delhi Prentice-Hall of India 16.8 TERMINAL QUESTIONS: 1 .Describe the meaning of quantitative data. 2 .Explain different types of scale of measurements with examples. 3 . Write in detail about the nature of quantitative data? UNIT :17 ORGANIZATION AND GRAPHICAL REPRESENTATION OF DATA 17.1    Introduction 17.2    Learning Objectives 17.3    Organization and Graphical representation of data Self- Check Exercise -1 17.4    Steps in the Construction of a Frequency Distribution Table Self- Check Exercise -2 17.5    Summary 17.6    Glossary 17.7    Answers to Self - Check Exercise 17.8    References/ Suggestive Readings 17.9    Terminal Questions 17.1    INTRODUCTION: In today’s data driven world, the ability to organiz4e and effectively present data is paramount. Whether in business, social trends or scientific findings, the way data is structured can profoundly impact the decision - making process and understanding. 17.2    LEARNING OBJECTIVES: After going through this unit, the students will be able to: 1.    Study about organization of data. 2.    Understand the graphical representation of data. 17 .3 Organization and Graphical representation of data In educational research, various tests are administered from time to time with a view to establishing individual differences, ascertaining relative status of students in a group and assessing significant group trends pertinent to the specific objectives of the project, based on surveys or experimental observations (which are made use of) for building up the scientific body of knowledge or for solving the immediate problems. In either case, investigator or classroom teacher comes across a mass of data comprising individual scores. This mass of unordered data is known as raw data and alt statistical investigations start with this type of fundamental information Classification and description of this data is needed for its interpretation and meaningful presentation. The classification leads the investigator to understand basic features of the data. A frequency distribution of data represents its arrangement in an orderly manner so as to indicate the frequency of occurrence of the different values of variables falling within arbitrarily defined intervals of the variable Its meaning will be made dear in the present lesson along with its graphical representation. Scores by themselves have no meaning, 55 out of 100 may be a good or a bad score depending on a variety of factors such as the difficulty of the test, the ability of the class, and the standard of marking by the examiner. Therefore, 55 out of 100 may mean average or poor or high score. We can not interpret the score of a student unless we know about the scores of other students who appeared in the same examination. For example, 53% may be a good score in one case and poor in another. Set-A: 30, 40, 45, 50, 53, 55                  Mean = 45.5 Set B: 50, 60, 53, 70, 75, 80                  Mean = 64.67 In set -A, score 53 is above the average score (45.5) and can be considered as a good score but in set-. B, it is below the average score (64.67) and therefore, here it is- a poor score. Hence to judge where a score is good or poor, reference must be made to the average score of the group. Even in two sets of scores with the same average, some scores may have different meanings. Consider the following two sets of scores. Set C: 15, 25, 35, 45, 55, 65, 75, 85, 95.              Mean = 55. Set-D: 35, 40, 45, 50, 60, 65, 70, 75. Mean = 55. In both the sets of scores the average of scores is 55, here the score 75 has a different meaning in each. In set-C, 75 is the third best score, where as in set-D, it is the best score. Thus, we observe that one must also know something about a particular score. Hence, the scores of two sets should not be compared unless the averages of the scores of the two tests and their distributions are similar. Example -1: Arrange the given scores in a frequency distribution 8,    6, 2, 5, 10, 9, 3, 2, 1, 5, 4, 5, 7, 6, 5, 3, 4, 7, 0, 7 2,    5, 6, 7, 4, 3, 4, 6, 8, 5, 7, 4, 3, 5, 3, 4, 9, 6, 1, 7 It is difficult to see from the above list how the scores are distributed. Inspection of these scores, however, shows that many scores occurred more than once. We observed that there are two 9s, two 8s, six 7s, and so on. This suggests that we may arrange the data in columns, as shown in Table-1. In first column we may arrange the possible scores in descending or ascending order, and in the other we record by tallies (opposite to the respective marks) the number of students scored these scores. When we have to record five tallies, we mark four tallies, thus III and the fifth is made across the four thus III. The tallies are then totaled in the next column. This column thus, lists the number of times each score occurred. Such an arrangement of data is known as frequency distribution and the number of times a particular score value occurs is known as frequency of that score. The value of score is generally represented by the symbol X and its frequency by the symbol 'f. The total number of frequencies is denoted by 'N' or Σf | Scores | Tallies | Frequencies | |---|---|---| | 10 | I | 1 | | 9 | II | 2 | | 8 | II | 2 | | 7 | IIII I | 6 | | 6 | IIIIX | 5 | | 5 | IIII III | 8 | | 4 | IIII I | 6 | | 3 | IIIIX | 5 | | 2 | III | 3 | | 1 | II | 2 | | 0 | I | 1 | | Total | | N = 40 | Arrangement in Table 1 is known as frequency array. The data have been classified in as many classes as there are score values within the total range of the variables. In the above example, the maximum score was 10 and the maximum possible score was 0. In this way we had only 11 classes. But when the range (the interval between the highest and lowest score) of scores is large, we shall have large number of classes if we classify in as many classes as there are score values. If the classes are more than twenty in number, it becomes difficult to handle them Therefore, in such cases we reduce the number of classes by arranging the data in arbitrarily defined sub-groups. Consider the following scores of data in arbitrarily defined classes, A range of values which incorporates a set of items is called a class. For example, 5-10, 10-15 or 5-9, 10-14 are the classes in a frequency distribution. We will understand this with another example Example: 2: Arrange the given scores in a frequency distribution 9,    7, 6, 3, 4, 9, 3, 1, 2, 7, 4, 2, 3, 9, 10, 10,    2, 8, 5, 5, 3, 1, 0, 10, 2, 9, 7, 8, 5, 0 | SCORES | TALLIES | FREQUENCIES | |---|---|---| | 10 | III | 3 | | 9 | IIII | 4 | | 8 | II | 2 | | 7 | III | 3 | | 6 | I | 1 | | 5 | III | 3 | | 4 | II | 2 | | 3 | IIII | 4 | | 2 | IIII | 4 | | 1 | II | 2 | | 0 | II | 2 | | TOTAL | | N = 30 | In first column we may arrange the possible scores in descending or ascending order, and in the other we record by tallies (opposite to the respective marks) the number of students scored these scores. The tallies are then totaled in the next column. This column thus, lists the number of times each score occurred. Such an arrangement of data is known as frequency distribution and the number of times a particular score value occurs is known as frequency of that score. The value of score is generally represented by the symbol X and its frequency by the symbol 'f’. The total number of frequencies is denoted by 'N'. Example: Consider the following scores of 50 students and arrange them in a frequency distribution. | 85 | 66 | 76 | 45 | 66 | 91 | 77 | 64 | 71 | 74 | |---|---|---|---|---|---|---|---|---|---| | 47 | 78 | 76 | 42 | 70 | 58 | 71 | 67 | 80 | 78 | | 73 | 48 | 68 | 87 | 71 | 72 | 65 | 69 | 73 | 84 | | 75 | 56 | 58 | 87 | 56 | 72 | 62 | 93 | 73 | 83 | | 97 | 81 | 51 | 61 | 53 | 72 | 62 | 79 | 88 | 79 | It is shown that the highest score is 97 and lowest score is 42. so the range is 55 (i.e. 97-42) Therefore, the distribution of scores can be conveniently arranged by dividing the range of 55 into ten or more classes if the class is taken to be of 5 points each and we take starting point 40, then the scores within the range 40 to 44 that is, all the given scores with the values 40, 41, 42, 43 and 44 will be grouped together to form the lowest class. All scores from 45 to 49 i.e., 45, 46, 47, 48 and 49 will form the next class. Similarly we shall group all scores within the classes 50 to 54, 55 to 59, 60 to 64 and so on. The highest class will be 95-99. In Table-2, these classes have been arranged serially from the smallest at the bottom to the largest at the top. Each class covers 5 scores. For each score we have marked tallies against the corresponding class. The first score of 85 is represented be a tally placed opposite- the class 85-89. The second score of 47 by tally placed opposite to class 45-49 and the third score 73 by a tally placed opposite to 85-89 The remaining scores have been listed the total number of tallies in each class i.e., the frequency being written in the next column 'f'. The total of 'fs' gives the number of scores (here 50) and are denoted by N. You should note that the beginning score of the lowest class was taken as 40 and not 42, which was the actual lowest score. Theoretically, classes of 42 to 46, 47 to 51, 52 to 56 etc., are as good as classes of 40 to 44, 45 to 49, 50 to 54 etc., but the second set is easier to handle from the point of view of tabulation as well as computations which will come later. Such a distribution is called grouped series. There are five kinds of continuous series: (i) Exclusive series (ii) Inclusive series (iii) Open end series (iv) Cumulative frequency series and (v) Mid-value series. Table - 2 | Class Intervals | Tallies | Frequencies | |---|---|---| | 95-99 | I | 1 | | 90-94 | I I | 2 | | 85-89 | IIII | 4 | | 80-84 | IIII | 5 | | 75-79 | IIII III | 8 | | 70-74 | IIII IIII | 10 | | 65-69 | IIII I | 6 | | 60-64 | IIII | 4 | | 55-59 | IIII | 4 | | 50-54 | II | 2 | | 45-49 | III | 3 | | 40-44 | I | 1 | | | | N = 50 | Conventions regarding the Formation of Class Intervals: In the frequency distribution as shown in Table-1, the original observations have been retained and we can reconstruct the data from the frequency distribution without loss. When the interval is of more than 1, as in the above example, it is of 5, some loss of information regarding individual observations is incurred. In such cases the original observations cannot be reproduced exactly from the frequency distribution. If the class interval is large in relation to the total range or the set of observations, this loss of information may be appreciable if the class interval is small, we shall have a large number of classes and there shall be very little gain in convenience over the utilization of the original data. Therefore, we choose such a class interval which is neither very small nor very large.The following conventions are generally used in the selection of class intervals: (i)The class interval should be well defined and should not overlap. Class intervals of 04, 5-9, etc. should be used in preference to 0-5, 5-10, etc. (ii)The class interval should be of such a size that with such class intervals the total range of observations is covered by 10 to 20 intervals. (iii)Start the class interval with a value which is a multiple of the size of the interval or a zero. For example, with a class interval of 5, the interval should start with the values of 0, 5, 10, 15, etc., while with a class interval of 2, the intervals should start with the values 0, 2, 4, 6, etc. (iv)The class intervals should be of uniform size and commonly of 3, 5 or 10, rarely of 7 and 20. SELF CHECK EXERCISE-1 1 .What do we call the mass of unordered data? a.Qualitative data b.Raw data c.Quantitative data d.All of the above 2 . Frequency distribution indicates: a.    Frequency of occurrence b.    Statistical analysis of data c.    Both a and b d.    None of these 17.4 Steps in Construction of a Frequency Distribution Table: 1 .Determine range of the scores. Range = Highest score - Lowest score 2 .Decide on how many categories of step intervals are required. Normally, the number of intervals used is not less than 10 and donot exceeds than 20. 3 .Divide the range by total number of intervals giving the actual size of each step interval. If the size of the interval used does not work out to be a whole number. Usually, the class intervals are taken to be equal to 3, 5 or 10. 4 .Construct the class interval column starting with the highest score or a convenient score. Subtract the class interval size from this score to get the lower score of the class. Repeat this procedure until the class interval column includes an interval into which the lowest score can be placed. 5 .Mark a tally for each individual score against the class interval in which it falls. 6 .Total up the tallies within each class interval and place them in the frequency column. Exact Limits of the Class Intervals: Where the variable under consideration is continuous and not discrete, we select a unit of measurement and record our observation on as discrete. For example, like in statistical measurement we measure the achievement of a child and award full scores of 26 or 80 instead of fractions series is discrete, though achievement is a continuous variable. When we record an observation in discrete form and the variable is a continuous one, (like the one we have just quoted), we 69.5 (70 — 7^)/" Class Interval" 74.5 Have just quoted), we imply that the value recorded-represents a value falling within certain limits, usually one-half unit above and below the value reported. Hence, a score of 44 on a continuous variable will represent a class interval 43.5 to 44.5 (e.g., all values greater than or equal to 43.5 and less than 44.5). These are called the exact limits of a class interval or class boundaries or end value. Similarly, the class interval "70-74" in the previous example will include all values greater than or equal to 69.5 (the lower Limit of 70) and below 74.5 (the upper limit of 74). i.e., it covers all values between 69.5 and 74.5. Cumulative Frequency Distribution: Sometimes we are not with the frequencies within the class intervals themselves, but rather with the number of percentage of values 'greater than or less than a specified value. This can be obtained by adding successively the individual frequencies. The new frequencies obtained by adding individual frequencies of class intervals are called cumulative class frequencies. If we denote individual class frequencies of the successive class intervals by f1, f2, f3, ...... fk, the cumulative frequencies will f1; f1 + f2, fi + f2 + f3 and so on. Table 3 shows how from the frequency distribution given in Table 2. we can get cumulative frequencies and cumulative percentages for the distribution. Table 4 | Exact Limits of Class Intervals | Class Intervals | Midpoints | Frequency (f) | Cumulative frequency (cf) | Cumulative Percentage Frequency (c%f) | |---|---|---|---|---|---| | 94.5-99.5 | 95-99 | 97 | 1 | 49 + 1 = 50 | 100.0 | | 89.5-94.5 | 90-94 | 92 | 2 | 47 + 2 = 49 | 98.0 | | 84.5-89.5 | 85-89 | 87 | 4 | 43 + 4 = 47 | 94.0 | | 79.5-84.5 | 80-84 | 82 | 5 | 38 + 5 = 43 | 86.0 | | 74.5-79.5 | 75-79 | 77 | 8 | 30 + 8 = 38 | 76.0 | | 69.5-74.5 | 70-74 | 72 | 10 | 20 + 10 = 30 | 60.0 | |---|---|---|---|---|---| | 64.5-69.5 | 65-69 | 67 | 6 | 14 + 6 = 20 | 40.0 | | 59.5-64.5 | 60-64 | 62 | 4 | 10 + 4 = 14 | 28.0 | | 54.5-59.5 | 55-59 | 57 | 4 | 6 + 4 = 10 | 20.0 | | 49.5-54.5 | 50-54 | 52 | 2 | 4 + 2 = 6 | 12.0 | | 44.5-49.5 | 45-49 | 47 | 3 | 1 + 3 = 4 | 8.0 | | 39.5-44.5 | 40-44 | 42 | 1 | 1 + 0 = 1 | 2.0 | | | | | N = 50 | | | SELF CHECK EXERCISE-2 1 .Arrange the following steps of construction of a frequency distribution table in sequence: a.    Construct the class interval column. b.    Total up the tallies. c.    Determine the range of the scores d.    Mark a tally for each individual score against the class interval. e.    Decide how many categories of intervals are required f.    Divide the range by the number of intervals. Choose the correct option from the following options: 1 .a, b, c, d, e, f 2 .d, b, a, c, f, e 3 .e, f, b, a, d, c 4 .c, e, f, a, d, b 2 .Which of the following is not continuous data? a.    A person’s weight b.    Volume of water in a bottle c.    Bikes manufactured in a factory in a day d.    None of these 17.5    SUMMARY: In this lesson, we have learnt about organization and graphical representation of data. We have learnt that mass of unordered data is known as raw data and all statistical investigations start with this type of fundamental information. 17.6    GLOSSARY: Interval - An interval is a set of real numbers with the property that any number that lies is also included in the set. Tallies – Tallies are a simple and visual way to represent numerical data and are commonly used for counting and recording of data. 17.7    ANSWERS TO SELF CHECK EXERCISES SELF CHECK EXERCISE:1 Answer 1. B Answer 2. A SELF CHECK EXERCISE-2 Answer 1. D Answer 2. C 17.8    REFERENCES/SUGGESTIVE RERADINGS: Ebel, Robert L (1966). Measuring Educational Achievement. New Delhi: Prentice Hall of India Pvt. Ltd pp. 481. Garrett, H. E. and Woodsworth, R. S. (1966). Statistics in Psychology and Education Bombay Vakils, Feffer and Simons Pvt. Ltd. 17.9    TERMINAL QUESTIONS 1.    What do you understand by frequency distribution? 2 .Write down the steps in construction of Frequency distribution table. 3 .Take scores of 100 students in any subject of your interest. Convert raw scores into frequency distribution. UNIT:18GRAPHIC REPRESENTATION OF FREQUENCY DISTRIBUTION 18.1    Introduction 18.2    Learning Objectives 18.3    Graphical representation of Frequency Distribution Self- Check Exercise -1 18.4    Drawing of a Frequency Polygon Self- Check Exercise -2 18.5    Drawing of a Cumulative Frequency Curve Self – Check Exercise-3 18.6    Summary 18.7    Glossary 18.8    Answers to Self - Check Exercise 18.9    References/ Suggestive Readings 18.10    Terminal Questions 18.1    INTRODUCTION: In this lesson we will learn about graphical representation of data, drawing of a graph, drawing of histogram, drawing of frequency polygon, drawing of a smoothed frequency polygon, drawing of cumulated frequency curve and also drawing up of a cumulative percentage curve which is also known as Ogive. 18.2 LEARNING OBJECTIVES: After going through this unit, the students will be able to: 1.    Study about histogram. 2.    Understand the concept of Frequency Polygon. 3.    Study about Smoothed frequency polygon. 4.    Learn about Cumulated frequency curve. 5.    Learn about Cumulated percentage curve. 18.3 Graphic Representation of Frequency Distribution: Graphic representation leads to the understanding of a data set. If the graph is well drawn, it is pretty easier to read and interpret from the table. We shall consider here only those graphs which are useful in visualizing the important properties of frequency distributions. These will be of great help in enabling us to comprehend the important features of frequency distributions and in comparing one frequency distribution with another. Four methods of graphic representation of frequency distribution are in general use. (i)    Histograms (ii)    Frequency Polygons (iii)    Cumulative Frequency Curve (iv)    Cumulative Percentage Curve or Ogive. Before proceeding to discuss the methods of representation of tabular data graphically, I would like to discuss some of the general conventions for the construction of graph and certain basic ideas in drawing of graph. Drawing of a Graph Before considering the procedure of constructing a frequency polygon or histogram, I would like to discuss with you the simple algebraic principles applicable to any graphic representation. Graphing or presentation of any distribution on a graph paper is done with reference to two lines i.e., co-ordinate axes. One, the vertical line or Y-axis and the other, the horizontal or X-axis Usually, these two lines are taken perpendicular to each other and their point of intersection is called origin 'O' This is the point of reference for both the axes. Some Conventions for the Construction of Graph: 1 .In a frequency distribution graph, it is customary to let the horizontal axis represent scores and the vertical axis the frequencies. 2 . The arrangement of the graph should proceed from left to right. The low or small numbers on the scale should be on the left, and the low numbers on the vertical scale should be at the bottom. 3 .The distance along either axis selected to serve as a unit is arbitrary and affects the appearance of the graph. Some writers suggest that the units should be so selected that the ratio of height to length should be roughly 3:5. This procedure has some aesthetic advantages. 4 .The vertical scale should be so selected that the zero point must fall at the point of intersection of the axes. When data start with a zero, it is customary to designate the point of intersection at the zero point and make a small break-off in the vertical line A similar procedure is to be followed for the horizontal line. 5 .Both the horizontal and vertical axes should be appropriately labeled. 6.    Every graph should be assigned a descriptive title which states precisely what it is about 7.    The scale of measurement (or units selected) should always be mentioned on the graph. Now let us go to the procedure of drawing a graph. Drawing of a Histogram: A histogram is a graph, where the frequencies are represented in the form of rectangular bars. The height of such rectangular bars is the frequency and the width of such bars is the length of respective class intervals (taking into consideration, the exact limits). Example: Draw a histogram of the following distribution in Table 4. Table - 4 | Class Interval | Exact Limits | Frequency (f) | |---|---|---| | 88-90 | 87.5-90.5 | 1 | | 85-87 | 84.5-87.5 | 1 | | 82-84 | 81.5-84.5 | 2 | | 79-81 | 78.5-81.5 | 2 | | 76-78 | 75.5-78.5 | 5 | | 73-75 | 72.5-75.5 | 2 | | 70-72 | 69.5-72.5 | 5 | | 67-69 | 66.5-69.5 | 6 | | 64-66 | 63.5-66.5 | 8 | | 61-63 | 60.5-63.5 | 4 | | 58-60 | 57.5-60.5 | 4 | | 55-57 | 54.5-57.5 | 3 | | 52-54 | 51.0-54.5 | 4 | | 49-51 | 48.5-51.5 | 2 | | 46-48 | 45.5-48.5 | 1 | | | | N = 50 | Solution: Step 1-Draw two straight lines perpendicular to each other, the vertical line near the left side of the graph paper and the horizontal line near the bottom of it. Step 2 - Label the vertical line (the Y-axis) OY, and the horizontal line (the X-axis) OX Put the 0 (origin) where the two lines interest. Step 3- Lay off the scores (or C) of the frequency distribution at regular distances along X-axis. The exact lower limits of the scores (or C1) are to be taken. Select X unit such that it will allow all the scores (or C.) to be represented easily on the graph. Step 4- Mark on the Y-axis successive units to represent the frequencies of the different intervals Choose a Y scale which will make the largest frequency (the height) of the histogram approximately 60% to 80% (or roughly 75%) of the width of the figure. Step 5-Draw rectangles Note: Units are mentioned on the upper right corner of the graph Scores Exercise: Draw a histogram of the following distribution: Table-5 | C.I. | Exact Limits | Frequencies (f) | |---|---|---| | 53-57 | 52.5-57.5 | 1 | | 48-52 | 47.5-52.5 | 1 | | 43-47 | 42.5-47.5 | 6 | | 38-42 | 37.5-42.5 | 4 | | 33-37 | 32.5-37.5 | 3 | | 28-32 | 27.5-32.5 | 14 | | 23-27 | 22.5-27.5 | 19 | | 18-22 | 17.5-22.5 | 22 | | 13-17 | 12.5-17.5 | 2 | Now follow the steps 1 to 5 of the previous example and get the histogram. Hint: Length on X-axis is 9 x 5= 45 divisions. Maximum height should be about 45 x 3/4 = '33.7, division which when divided by the maximum number (frequency 22) gives approximately 1.5 div. This is a suitable, but not rigid unit on Y-axis. SELF CHECK EXERCISE -1 1.I n the construction of graph, the arrangement of graph should proceed from: a.    Left to Right b.    Right to Left c.    Horizontal to Vertical d.    Vertical to Horizontal 2.I n a frequency distribution graphs, the horizontal axis represents ______ and the vertical axis represents the________. 3.    In histogram, the frequencies are represented in the form of ________bars. 18.4    Drawing of a Frequency Polygon: A frequency polygon is a closed figure where the frequencies are represented by the height of the perpendicular from the mid-values of each class interval. In the construction of a frequency polygon, after the usual steps in the drawing of any graph, points plotted are above the midpoints of the intervals on the X-axis. Frequencies are represented in each case by a dot. When all the points are located, they are joined by a sense of short lines to form a frequency curve. In order to complete the polygon, additional intervals at the high and low ends of the distribution are included on the X-axis. The frequencies on each of these additional intervals is naturally Zero. The polygon is closed by joining the mid-points of these two extreme intervals on the X- axis. Example: Represent the frequency distribution in Table 4 as a frequency polygon. Solution: Follow the steps 1 to 4 of histogram, except for step 3, instead of taking the exact limits of scores, take the mid-points of the class intervals. These mid points are to be marked off on the X-axis 5.From the midpoint of each interval on the X-axis, go up in the Y direction at a distance equal to the number of frequencies of the intervals. Place points at all these locations 6.Join all the points with straight lines to each other to get the frequency polygon. Notice that the end points of the polygon have been positioned on the X-axis at the midpoints of the intervals on either side of the two extreme intervals. Exercise: Represent the frequency distribution in Table-5 in a frequency polygon. Drawing of a Smoothed Frequency Polygon: Sometimes the researcher wishes to know what sort of frequency distribution he would have obtained the data had been collected from a larger sample, He can get this idea to some extent by smoothing the original polygon. Smoothing a frequency polygon requires the calculation of averaged frequencies for all intervals. The average frequency for any interval is calculated by adding the frequency of that interval to the frequencies of two adjacent intervals, and dividing the total by 3 For example, in Table-6 the averaged frequency for the interval 49-51 is the frequency of the interval (2) added to the frequency of the interval 46-48(1) and the frequency of the interval 52-54(4) divided by 3. (2 + 1 + "4 " )/3= 2.3 This procedure is repeated for all of the class intervals, when an interval has only an adjacent interval, then missing interval is given a frequency of 0. For example, the averaged frequency for the interval 46-48 is (0 + 1 + "2 " )/3 = 1.0 Table-6 shows the frequency distribution with averaged frequency for the of the class intervals. The smoothed frequencies are plotted for the polygon. If this is super imposed over the original frequency polygon, then comparison can be made as to the shape of the curves. Table 6 Frequency distribution (Smoothed Frequencies) | Class Interval | Midpoints (X) | Frequencies | Averaged or Smoothed Frequencies | |---|---|---|---| | 88-90 | 89 | 1 | 0.6 | | 85-87 | 86 | 1 | 1.3 | | 82-84 | 83 | 2 | 1.6 | | 79-81 | 80 | 2 | 3.0 | | 76-78 | 77 | 5 | 3.0 | | 73-75 | 74 | 2 | 4.0 | | 70-72 | 71 | 5 | 4.3 | | 67-69 | 68 | 6 | 6.3 | | 64-66 | 65 | 8 | 6.0 | |---|---|---|---| | 61-63 | 82 | 4 | 5.3 | | 58-60 | 59 | 4 | 3.6 | | 55-57 | 56 | 3 | 3.6 | | 52-54 | 53 | 4 | 3.0 | | 49-51 | 50 | 2 | 2.3 | | 46-48 | 47 | 1 | 1.0 | SELF CHECK EXERCISE-2 1 .A frequency polygon is a __________ figure. a.    Open b.    Closed c.    Unlocked d.    None of these 18.5    Drawing of a Cumulative Frequency Curve: A curve representing accumulated frequency is usually termed as Accumulative frequency curve or graph A cumulative frequency is defined as frequencies accumulated progressively from the bottom of the distribution to upwards. To illustrate, we can take the help of Table-3, then. We can plot the graph with the help of these added frequencies. The cumulative frequencies (cf) are plotted vertically at the upper exact limit of each class interval (or scores) because cumulative frequencies are cumulated through each interval rather than their mid values. Once the points are plotted, they are connected by straight lines (free hand) beginning with zero at the lower limit of the lowest interval and ending with upper exact limit of the interval at the upper end of the X-axis. Example: Plot a cumulative frequency graph of the distribution given in Table-7. Table-7 Cumulative Frequency Distribution | Class Interval (Scores) | Class Interval (Exact Limits) | Frequency (f) | Cumulative Frequency (f) | |---|---|---|---| | 88-90 | 87.5-90.5 | 1 | 50 | | 85-87 | 84.5-87.5 | 1 | 49 | | 82-84 | 81.5-84.5 | 2 | 48 | | 79-81 | 78.5-81.5 | 2 | 46 | | 76-78 | 75.5-78.5 | 5 | 44 | | 73-75 | 72.5-75.5 | 2 | 39 | | 70-72 | 69.5-72.5 | 5 | 37 | | 67-69 | 66.5-69.5 | 6 | 32 | | 64-66 | 63.5-66.5 | 8 | 26 | | 61-63 | 60.5-63.5 | 4 | 18 | | 58-60 | 57.5-60.5 | 4 | 14 | | 55-57 | 54.5-57.5 | 3 | 10 | | 52-54 | 51.5-54.5 | 4 | 7 | | 49-51 | 48.5-51.5 | 2 | 3 | | 46-48 | 45.5-48.5 | 1 | 1 | The procedure for plotting a cumulative frequency curve is the same as for frequency polygons except that of the cumulative frequencies are plotted on the upper exact limits of each of the class interval instead of midpoints, as explained earlier. Summary of the steps is given below: 1 .Draw the axes. 2 .Mark off upper exact limits of each of the class interval on X-axis and cumulative frequencies on the Y-axis, by selecting appropriate units. 3 .Plot cumulative frequencies for each class interval on the upper exact limit of that interval. 4 .Join the points to get the cumulative frequency curve. Cumulative frequency curve can be smoothed in the same manner in which frequency polygon can be smoothed Drawing of a Cumulative Percentage Curve or Ogive: The cumulative percentage curve or Ogive differs from cumulative frequency curve. In it, frequencies are expressed as cumulative percents of frequency on the X- axis In Table-8 cumulative frequencies have been converted into cumulative percents. The procedure of plotting an Ogive is exactly the same as that of plotting cumulative frequency curves except that instead of cumulative frequencies, the percentage of cumulative frequencies are plotted on the upper exact limits of each class interval Example: Plot an Ogive of the distribution given in Table-8. Table-8 Cumulative Percentages | Class Intervals | Upper exact limit | Individual frequencies (f) | Cumulative frequency (cf) | Cumulative Percentage Frequency (c%f) | |---|---|---|---|---| | 41-43 | 43.5 | 1 | 86 | 100.0 | | 38-40 | 40.5 | 4 | 85 | 98.8 | | 35-37 | 37.5 | 5 | 81 | 94.2 | | 32-34 | 34.5 | 8 | 76 | 8.4 | | 29-31 | 31.5 | 14 | 68 | 79.1 | | 26-28 | 28.5 | 17 | 54 | 72.8 | | 23-25 | 25.5 | 9 | 37 | 43.0 | | 20-22 | 22.2 | 13 | 28 | 32.6 | | 17-19 | 19.5 | 8 | 1 | 17.4 | | 14-16 | 16.5 | 3 | 7 | 8.1 | | 11-13 | 13.5 | 4 | 4 | 4.7 | | 8-10 | 10.5 | 0 | 0 | 0.0 | Step 1- Find out the exact upper limit of each class interval Step 2-Accumulate the successive frequencies. Step 3- Convert these cumulated frequencies into percentages to obtain cumulative percentage frequencies Step 4- Plot these cumulative percentages on the upper exact limits of the class intervals in the same manner as we have done in the case plotting of a cumulative frequency curve. The only difference is that on the Y-axis, instead of cumulative frequencies, cumulative percentage frequencies are marked off from 0 to 100. Step 5-Join the points smoothly to get Ogive ù ' Cumniilative% age frequency Curve (Ogive) 40 j »0. K_ Scale on X axis : 1 Unit = 3 scores on Y axis : 1 Unit= 10 frequencies Exercise: Plot Ogive for the given frequency distribution | | II | III | IV | V | |---|---|---|---|---| | C.I. | Exact limits | F | Cf | Cf% | | 53-57 | 53-57 | 1 | | | | 48-52 | 48-52 | 1 | | | | 43-47 | 43-47 | 6 | | | | 38-42 | 38-42 | 4 | | | | 33-32 | 33-32 | 8 | | | | 28-32 | 28-32 | 14 | | | |---|---|---|---|---| | 23-22 | 23-22 | 10 | | | | 18-22 | 18-22 | 4 | | | | 13-17 | 13-17 | 2 | | | When to use the Frequency Polygon and Histogram: The frequency polygon is less precise than the histogram. Because it does not represent accurately in terms of area the frequency in each interval. In comparing two or more graphs plotted on the same axis, the frequency polygon is likely to be more useful. SELF CHECK EXERCISE-3 1 .Accumulated frequency curve represents: a.    Undistributed frequency b.    Distributed frequency c.    Accumulated frequency d.    All of these 2 .A cumulative frequency is defined as: a.    Frequencies accumulated from upwards to bottom b.    Frequencies accumulated from bottom of distribution to upwards c.    Both a and b d.    None of these 18.6    Summary: After going through this lesson, we have learnt about different histograms, polygons, Cumulative frequency curve and Ogive curve. These were understood by us in terms of their representation. 18.7    GLOSSARY: Horizontal axis: A horizontal axis runs horizontally from left to right on a graph or chart. Vertical axis: A vertical axis runs vertically from bottom to top on a graph or chart. Percentage: Percentage is a way of expressing a proportion or a fraction of a whole in terms of hundredths. 18.8    ANSWERS TO SELF CHECK EXERCISES: SELF CHECK EXERCISE- 1 Answer 1. A Answer 2. Scores, Frequencies Answer 3. Rectangular SELF CHECK EXERCISE-2 Answer1. B SELF CHECK EXERCISE-3 Answer 1. C Answer 2. B 18.9    REFERENCES/SUGGESTIVE READINGS: Sanswal, N.D. (2020). Research Methodology and Applied Statistics. (1st ed.). Shipra Publications. Kaul, L. (2009). Methodology of Educational Research. (4th ed.). Vikas Publishing House Private Limited. Best, John W. and Kahn, James V. (2004). Research in Education (7 Ed.). New Delhi: Prentice-Hall of India. 18.10 TERMINAL QUESTIONS: 1 .Take scores of 100 students in any subject of your interest. Convert raw scores into frequency distribution and draw a Histogram. UNIT: 19 MEASURES OF CENTRAL TENDENCY (MEAN, MEDIAN, MODE) 19.1    Introduction 19.2    Learning Objectives 19.3    Measures of Central Tendency 1.    The Arithmetic Mean Self- Check Exercise -1 2.    The Median Self- Check Exercise -2 3.    The Mode Self-Check Exercise-3 19.4    Summary 19.5    Glossary 19.6    Answers to Self - Check Exercise 19.7    References/Suggestive Readings 19.8    Terminal Questions 19.1    INTRODUCTION: In this unit, we will learn about measures of central tendency which includes mean, median and mode. In the field of statistics, central tendency can be referred to as a central value for any probability distribution. Measures of central tendency are the statistical measures which are used to describe the center of a data set. These measures provide a single value that represents the central value around which the data tend to cluster. They help in understanding the central or average value of a data distribution, aiding in data analysis and interpretation. 19.2    LEARNING OBJECTIVES: After going through this unit, the students will be able to: 1.    Define the concepts of mean, median and mode. 2.    Identify the conditions when different measures of central tendency should be used. 3.    Compute mean, median and mode for given data 19.3    Measures of Central Tendency: Statistical methods provide a variety of measures to describe a distribution. The most often used descriptive measure is one which determines the 'central tendency' of the distribution or a set of data. A more popular term for a measure of central tendency is an average. People generally talk about the average price, average height, average score etc. What do we mean by an average? Usually we mean some measure which, in some sense, summarizes a set of data. For example, if we are told that the average IQ of a group of students is 118, this single figure leads us to think of this group as generally bright. We know that not all the students in the group have an I.Q. of 118. Some have a lower 1,Q. for the group which permits us to gain a general impress of the intelligence level of the group of students as a whole. Types of Statistical Averages: For the sake of simplicity of our study, various averages may be divided as follows: Mathematical Average: (i) Arithmetical Mean, (ii) Geometric Mean, (iii) Harmonic Mean Positional Average: (i) Median, (ii) Mode The following three measures of central tendency are commonly used in Education: (a) The Arithmetic Mean or Simple Mean, (b) The Median, (c) The Mode. 1 .The Arithmetic Mean The 'Arithmetic Mean' is perhaps the most familiar and most frequently used measure of central tendency. Moreover, it is relatively easy to calculate and is widely used in statistical research. Generally, when the term mean is used without a modifier (adjective) the arithmetic mean is implied. The arithmetic mean represents the typical value of the dataset by distributing the total value equally across all observations. The mean is sensitive to extreme values, often reflecting the central tendency when the data is normally distributed. The arithmetic mean or simply the mean of a set of observations is obtained by adding all the values and dividing the total by the number of the values. If 'X' denotes measurements X in question, then the mean is always denoted by ' ' (read xbar) or 'M', and 'N' stands for the n of the values and symbol 'Σ' (Sigma) denotes the sum of all the terms following it, we have the formula: M or (∓) =(0 + 1+ "1" )/3= 0.67 i.e. Mean = "sum of the scores or values" /"number of the scores or values" Example 1: The following are the ungrouped scores of 12 students 17, 25, 44, 16, 20, 15, 10, 30, 19, 22, 31, 11 Calculate mean score for the given set of scores. Sol. By definition of the mean Mean = XX/N X = “17+ 25+ 44+ 16+ 20+ 15+ 10+ 30+ 19+ 22+ 31+11” /12 = 260/12 = 21.6 Example 2: The following are the ungrouped scores of 10 students 47, 52, 51, 50, 52, 53, 51, 51, 49, 55. Calculate mean score for the given set of scores. Sol. By definition of the mean Mean = XX/N X ("47 " + " 52 " + " 51 " + " 5" 0 + " 52 " + " 53 " + " 51" + " 51" + " 49 " + " 55" )/10 =         = 51.1 Deviation: Deviation of any X value from its mean is its difference from the mean. In example 2, the deviation of 53 from its mean 51.1 is 1.9 i.e., 53-51.1 and similarly the deviation of 47 from its mean is -4.1 i.e. (47 – 51.1 = -41). Here, we can say that Deviation (Score-Mean) or x = (X - M). Where: x denotes deviation of score X from its true mean. 1.Short-Cut Method of Calculating the Mean In this method we first take an arbitrary (assumed) working mean and calculate the mean of the deviations of the observations from this working mean. If 'x' denotes deviation of any 'X' value from assumed mean (A.M.) then: Mean or (x ) A.M. + XX/N                    (1) Now we can solve example-1, by taking the assumed mean (A.M.) equal to 50 | Serial No.                      X | X' = X - A.M. | |---|---| | 1                                47 | -3 | | 2                            52 | 2 | | 3                              51 | 1 | | 4                          50 | 0 | | 5 | 52 | 2 | |---|---|---| | 6 | 53 | 3 | | 7 | 51 | 1 | | 8 | 51 | 1 | | 9 | 49 | -1 | | 10 | 55 | -5 | | | | Σx =11 | Mean or (x) =A.M. + XX/N = 50 + 11/10 = 50+ 1.1 + 51.1 As regards the final result, any assumed mean will do, but in order to keep the computation to the minimum, care should be taken to select an assumed mean as close as possible to the actual mean, as far as one can judge. (II)    The Arithmetic Mean of Grouped Data: One simple method to calculate the mean of grouped data is given below. Mean or (x) = XfX/N Where XfX is the sum of product of frequency (f) and the midpoints (X) of class intervals Step-1. Find out the mid-point of each class interval by finding out the average of upper and lower scores of the class. Step-2. Multiply each mid-point value (X) by the frequency (f) in that class to get the value of fx. Step-3. Add all values of fx to obtain XfX Step-4. Apply the following formula to get the value of mean. Mean or (X) = XfX/N Example-2: From the given frequency distribution, calculate the value of Mean | Score | F | |---|---| | 30-34 | 2 | | 25-29 | 2 | | 20-24 | 5 | | 15-19 | 6 | | 10-14 | 4 | | 5-9 | 1 | Solution: | Score (C.I.) | Midpoints intervals (X) | F | fx | |---|---|---|---| | 30-34 | 32 | 2 | 64 | | 25-29 | 27 | 2 | 54 | | 20-24 | 22 | 5 | 110 | | 15-19 | 17 | 6 | 102 | | 10-14 | 12 | 4 | 48 | | 5-9 | 7 | 1 | 7 | | | | N = 20 | XfX = 385 | Mean or (X) = XfX/N =       = 19.25 Example-3: From the given frequency distribution, calculate the value of Mean | Score | F | |---|---| | 80- 84 | 5 | | 75 – 79 | 1 | | 70 – 74 | 2 | | 65 – 69 | 4 | | 60 – 64 | 3 | | 55 – 59 | 7 | Solution: | Score (C.I.) | Midpoints intervals (X) | F | fx | |---|---|---|---| | 80-84 | 82 | 5 | 410 | | 75-79 | 77 | 1 | 77 | | 70-74 | 72 | 2 | 144 | | 65-69 | 67 | 4 | 268 | | 60-64 | 62 | 3 | 186 | | 55-59 | 57 | 7 | 399 | | | | N = 22 | XfX= 1484 | Mean = EfX/N = 1484/22 = 67.45 (III)    The Arithmetic Mean from Group Data: (short method). The steps in the process of computing a mean from a frequency distribution according to the short-cut method are as follows: | Step 1. | Arrange the data in a frequency distribution to get the total number of items, i.e. N- Ʃ f, and the width of class intervals, i.e., (i) | |---|---| | Step 2. | By mere inspection estimate the interval in the distribution which is most likely to contain the mean. The midpoint of this interval will be designated as AM (Assumed Mean). | | Step-3. | In the fourth column, Le, the column next to that of frequencies, mark the deviation 'x' of the midpoint of the interval from the assumed mean in steps, i.e. in terms of interval. Those that are greater than the assumed mean are designed plus (+) and the small are designed minus (-) | | Step-4. | The next operation is to multiply each frequency (f) by its corresponding deviation (x'). The products (fx') are written in the fifth column. Care should be taken to observe signs. | | Step-5. | Find the algebraic sum of fx' i.e. (Ʃfx') from fifth column. | | Step-6. | Divide Ʃfx' by N. Sometimes Ʃfx' will be positive and sometimes negative. This sign depends on the position of the assumed mean. | | Step-7. | Multiply XfX'/Nby the site of the class interval (i) of the class interval is also called the width of the class interval and is equal to the difference between lower or upper limits of two consecutive class intervals. | | Step-8. | The final step is to calculate the mean by applying the following formula. Mean or (X ) = A.M. + XfX'/N x i | Remember that the above formula is only when all classes are of equal width. Example -3: Calculate mean of the following distribution. | Class Interval (C.I.) | Frequency (f) | |---|---| | 53-57 | 1 | | 48-52 | 1 | | 43-47 | 2 | | 38-42 | 3 | | 33-32 | 5 | | 28-32 | 4 | | 23-22 | 2 | | 18-22 | 1 | | 13-17 | 1 | Solution: Step-1      Adding up all the frequencies, we get N (Total)or Ʃf=20 Step-2      As we need mid-points of the class intervals so we re-write the frequency distribution, taking the mid-point of each of the C.I in column II Step-3      Let us take mid-point 30 of class interval 28-32 as our assumed mean (A.M. = 30). Step-4      Since A.M. 30 (deviation of midpoint in C.1. units) corresponding to 28-32 is zero, and that for class interval 23-27 is-l' for, class interval 18-22 is - 2 and soon. Similarly for Class Interval 33-37 value for 'x 'is + 1' and that for 38-42 is + 2' and so on. In the forth column we note the values of x'. Step-5       Calculate the products 'fx', i.e. product of frequency and corresponding deviation x' and write the values in the 5th column fx. In the CI. 53-57, f= I and x +5, hence fx + 5. In C. 48-52, f= 1, x = +4 and so on. The particular sign of '+' or should not be ignored. | (I) C.I | (II) Mid-point | (III) (f) | (IV) (x') | (V) f x' | |---|---|---|---|---| | 53-57 | 55 | 1 | +5 | +5 | | 48-52 | 50 | 1 | +4 | +4 | | 43-47 | 45 | 2 | +3 | +6 | | 38-42 | 40 | 3 | +2 | +6 | | 33-37 | 35 | 5 | +1 | +5 | | 28-32 | 30 AM | 4 | 0 | +0 | | 23-27 | 25 | 2 | -1 | -2 | | 18-22 | 20 | 1 | -2 | -2 | | 13-17 | 15 | 1 | -3 | -3 | N = 20     |-l(&+26@&-7)]S/V= 19 Step-6       Find the algebraic sum of the 5th column, (first add all the '+' values and then add the '-' values, the algebraic sum will be the value of Ʃfx'. Step-7      Calculate mean by applying the following formula. x =A.M. + ^JX^/N xi Step-8      Put the respective values of A.M. = 30 (assumed mean) ƩfX' = 19 (calculated in 5th column) 1 = 5 (size of C.I.) N=20 (Total number of frequencies) Step-9      Mean = 30 + 19/20 x 5 = 30 + 95/20= 30 + 4.75 + 34.75 The Mean from Combined Groups. If there are 'n' groups whose means and number of cases in each group (N) are known, we can compute their weighted mean using the following formula Mcomb = (W_l M_1 + N_2 M_2..........+N_n M_n)/(N_1 + N_2+............N_n) Example-4: Compute the Mcomb for the table given below: | Group | Mean (M) | N | |---|---|---| | I | 20 | 10 | | II | 15 | 20 | | III | 25 | 30 | Solution: Some Properties of Arithmetic Mean - Following are the important properties of Arithmetic mean are- 1 .The sum of the deviations of all the measurements in a set from their arithmetic mean is zero. If x or M is the mean then (X-) =0.x 2 .    The sum of squares of deviations about the arithmetic mean is less than the sum of squares of deviations about any other value. The deviations of scores 2, 8, 16, 4 , 6 and 0 form their mean 6 are -4, 2, 10-2, 0 and -6. The squares of these deviations are 16, 4, 100, 4, 0, 36. The sum of these squared deviations is 160. If we take deviations from some other value, say 8, the deviations are -6, 0, 8, -4, -2, -8 squaring these we have 36, 0, 64, 16, 4, 64 the sum of these is 184, which is greater than the sum of squared deviations from the mean. Selection of any other value will demonstrate the same result. SELF CHECK EXERCISE-1 1.Sum of scores divided by number of scores is the formula of: a.    Mean b.    Mode c.    Median d.    None of these 2.    Median and Mode comes under the category of: a.    Mathematical average b.    Positional average c.    Both a and b d.    None of the above 2. The Median For most purposes arithmetic mean is used to summarize the general character of collection of data. But if collection of data contains extreme values, which markedly affect the arithmetic mean, the simple mean is inappropriate or even misleading. For example, if a sample of five values is 45, 57, 65, 90, 98, its arithmetic mean is 71.0. This value is not a true representative of the sample and misleads us. In such cases another value called the 'Median' is used. Median is the middle value in a dataset which is in ascending or descending order. If there is an even number of values, the median is the average of two middle values. Median of a set of data is that point below and above which 50% of the values lie. In other words, it is the half way mark when the data are ranked (arranged in order of magnitude). In the above example, values are in ascending order and the middle value is 65. Therefore, the median is 65. Median does not take into consideration the absolute difference in successive pairs or values. The median is often called positional average. In the arithmetic mean each value contributes to the final result according to its magnitude. In median only the middle value represents the whole data. In a series consisting of an odd number of values, the median is the middle most value. If a series is arranged in numerical order i.e., it is the arithmetic mean of the two middle values. For example median of 31, 32, 35, 36, 38 and 42 is ((35 + 36))/2 = 35.5 (I)    Calculation of Median from Grouped Data In case of grouped data, ie, frequency distribution, we first prepare the cumulative frequency distribution and then the formula for calculation of the median is: Mdn = I + ((^/2 - F»/f X i Where: Mdn is median, (I) is the exact lower limit of the interval which contains the median, the point below which half the number of cases lie (i) is the width of the class interval in which the median falls (f) is the frequency of the same interval and (F) is the cumulative frequency of the preceding classes or intervals. Example 1 - Calculate the median from the data given below. | Class Interval | Frequency | |---|---| | 40-44 | 2 | |---|---| | 35-39 | 2 | | 30-24 | 4 | | 25-29 | 6 | | 20-24 | 10 | | 15-19 | 15 | | 10-14 | 6 | | 5-9 | 5 | | | N = 50 | Solution: Step-1 We first add the frequencies and get N or Σf = 50 Step-2 We have the following working formula for the median Mdn = I + W/2 - F»/f X i We need the exact lower limit of the median class and F (the cumulative frequency, i.e., sum of the frequencies of the preceding classes). Step-3 We take the exact limits and prepare a cumulative frequency distribution. Step-4 Here N= 50. Therefore: N/2 = 50/2 = 25 The cumulative 'fs' are required to decide about the median interval. Hence the median class (class in which the median lies) is 15-19. Because in this interval we are getting the required number (25) of cumulative frequencies Step-5Calculate median by applying the formula Mdn = 1 +W/2-FT)/f*i | C.I. | Limit | f | cf | |---|---|---|---| | 40-44 | 39.5-44.5 | 2 | 50 | | 35-39 | 34.5-39.5 | 2 | 48 | | 30-34 | 29.5-34.4 | 4 | 46 | | 25-29 | 24.5-29.5 | 6 | 42 | | 20-24 | 19.5-24.5 | 10 | 36 | | 15-19 | 14.5-19.5 | 15 | 26 | | 10-14 | 9.5-14.5 | 6 | 11 | | 5-9 | 4.5-9.5 | 5 | 5 | Here: 1 14 5, (the exact lower limit of the C 1. 15-19 is 14.5) F= 11 (cumulative frequencies preceding the median class interval 15-19) f=15 (frequency in the Mdn class), and i= 5 (size of the class interval) Putting the values in the above formula Mdn = 14.5 +(50/2 - 11)/15 X 5 = 14.5 + (25 - 11)/15 X 5 = 14.5 + 14/(15_3) X 5 = 14.5 + 14/3 = 14.5 + 4.67 = 19.17 Example 2 - Calculate the median from the data given below. | Class Interval | Frequency | |---|---| | 185-205 | 4 | | 165-184 | 8 | |---|---| | 145-164 | 14 | | 125-144 | 20 | | 105-124 | 13 | | 85-104 | 5 | | 65-84 | 4 | | | N = 68 | Solution: Step-1      We first add the frequencies and get N or Σf = 68 Step-2      We have the following working formula for the median Mdn = I + ((^/2 - FV/f X i We need the exact lower limit of the median class and F (the cumulative frequency, i.e., sum of the frequencies of the preceding classes). Step-3      We take the exact limits and prepare a cumulative frequency distribution. Step-4      Here N= 68. Therefore: N/2 = 68/2 = 34 The cumulative 'fs' are required to decide about the median interval. Hence the median class (class in which the median lies) is 125-145. Because in this interval we are getting the required number (34) of cumulative frequencies Step-5 Calculate median by applying the formula Mdn = 1 + (( N/2-F))/ f x i | C.I. | Limit | F | cf | |---|---|---|---| | 185- 204 | 184.5-204.5 | 4 | 68 | | 165-184 | 164.5-184.5 | 8 | 64 | |---|---|---|---| | 145-164 | 144.5-164.5 | 14 | 56 | | 125 – 144 | 124.5-144.5 | 20 | 42 | | 105-124 | 104.5-124.5 | 13 | 22 | | 85-104 | 84.5-104.5 | 5 | 9 | | 65-84 | 64.5-84.5 | 4 | 4 | Here: 124.5, is the exact lower limit of the class interval 125- 145 F= 20 (cumulative frequencies preceding the median class interval 125-145) f=15 (frequency in the Mdn class), and i= 20 (size of the class interval) Putting the values in the above formula Mdn = 124.5 + (34-22) / 20 x 20 = 124.5 + 22/20 x 20 = 124.5 + 34-22 = 124.5 + 12 = 136.5 (II)    Special Case: When the frequency of Median interval is zero (0). Calculation of Median when: (a) the frequency distribution contains gaps; and (b) the first or last interval has indeterminate limits. (a)Difficulty arises when there are gaps or zero frequency which Mdn falls. We can take the following example. | C.I. | F | Cf | |---|---|---| | 20-21 | 2 | 16 | |---|---|---| | 18-19 | 2 | 14 | | 16-17 | 4 | 12 | | 14-15 | 0 | 8 | | 12-13 | 0 | 8 | | 10-11 | 4 | 8 | | 8-9 | 4 | 4 | As N=16, N/2 = 8; which is the of in both the Ci (12-13) and (14-15) giving two values of median, if the ordinary method is applied median may either be 13.5 or 15.5. Similarly we can get cases where more than two such values are obtained (if there are many zero frequencies. There are many ways of adjustment in such cases. One of the easiest way is to take the mean of the exact lower limit of the first and the upper limit of next class interval in which cumulative frequency is N/2. In this case cf N/2, in 12-13 and 14-15, the mean of II.5 and 15.5 is 13.5. Hence here median is 13.5. (b) The mean of grouped data cannot be calculated if the distribution is open ended (e.g. 80 and above, at the higher end or 20 and below, on the lowest end, but the median is readily computed since each score is simply counted as one frequency whether accurately classified or not, provided that the class in which median falls is defined. The calculation of median will be discussed in detail during P.C.P. SELF CHECK EXERCISE-2 1 .Median is the _______ value in a dataset. a.    First b.    Second c.    Middle d.    Last 2 .Median of a set of data is that point below and above which _____ of the values lie. a.    50% b.    25% c.    75% d.    40% 1.The Mode Mode is the number that occurs with highest frequency. If no number repeats, the dataset is said to have no mode or it can have multiple modes if two or more numbers occur with the same highest frequency. The mode is the value of the variate that occurs most commonly i.e., that value of variate for which frequency is maximum. If there is only one value which occurs the maximum number of times, then the distribution is said to have one mode or to be unimodal. On the other hand, if more than one value is repeated most and the same number of time the distribution is said to be multimodal. In a sense, mode is the most representative of observations because it tells us which value is most common. But certain difficulties arise in the calculation and use of this value in case value of the variate has same frequency (rectangular or uniform distribution) the mode has no meaning. It is greatly, moreover, as mentioned earlier a distribution may have more than one mode. If the distribution is unimodal there is hardly any difficulty in its calculation. We simply have to see from the data which of the observations occurs most. Example - Find the mode of the given observations: 1,2,3,4,5,6,5,6,5 Solution: We observe that the distribution is unimodal and 5 occurs maximum number of times i.e. 3 times. Therefore, the mode is 5. (I)    Calculation of Mode from Grouped Data - In case of grouped data, the crude mode is the mid- point of the class interval with maximum (f) while true mode is calculated by the following formula Mode = 3 Median—2 Mean This formula is applicable only when the values of mean and median are equal or almost equal I.e., difference should not be more than 0.5. If the difference is more than 0.5 then the value of mode will get affected accordingly. Example: Compute the mode for the following frequency distribution. | CI | F | |---|---| | 50-53 | 1 | | 46-49 | 1 | | 42-45 | 3 | | 38-41 | 2 | | 34-37 | 2 | | 30-33 | 1 | Solution:   For computing the mode, we have to compute mean and median first. (i)    Calculation of Mean : Assume 43.5 as assumed mean (the mid-value of the interval 42-45). | CI | F | X | fx' | |---|---|---|---| | 50-53 | 1 | +2 | +2 | | 46-49 | 1 | +1 | +1 | | 42-45 | 3 | 0 | 0 | | 38-41 | 2 | -1 | -1 | | 34-37 | 2 | -2 | -4 | | 30-33 | 1 | -3 | -3 | | | N=10 | | Ʃfx' = -5 | Here: A.M. = 43.5, Ʃfx' = -5, N=10 and I = 4. Put these values in the following formula Mean A.M. + (xm/Nx i EJx' = 43.5 + (-5)/W x 4 = 43.5 (-20)/10= 43.5 - 2.00 = 41.50 (ii)    Calculation of Median | Class intervals | F | Cf | |---|---|---| | 50-53 | 1 | 10 | | 46-49 | 1 | 9 | | 42-45 | 3 | 8 | | 38-41 | 2 | 5 | |---|---|---| | 34-37 | 2 | 3 | | 30-33 | 1 | 1 | Here: N/2 = 10/2 = 5 The interval in which we get cumulative frequencies is the median interval Thus: 38-41 is the median interval with lower limit 37.5 I = 37.5, F = 3, f = 2 and i = 4 (Now put these values in the formula of median) Mdn =I+^N/2~F^/fxi = 37.5 + ((10/2 - 3))/2x 4 = 37.5 + 1 x 4 = 37.5 + 4 + 41.5 Ans. Mode = 3 median - 2 mean: (This formula is applicable when the values of mean and median are equal almost equal) i.e. the difference between mean and median value is less than 5. =3(41.50) - 2, (41.50) = 124.50 - 83.00 = 41 Ans II.    Calculation of Mode Directly: Mode can also be calculated directly by applying the following formula Z or M0 = I+ (f-m - fS)/(2f_m - f_l - f_2 ) xi In the formula: 1    = exact lower limit of the modal class fm = Frequency of the modal class fi = Frequency of the class preceding the modal class f2 = Frequency of the class succeeding the modal class i = Size of the class interval Example: Calculate mode from the following frequency distribution | Class interval | F | |---|---| | 35-39 | 1 | | 30-34 | 2 | | 25-29 | 4 | | 20-24 | 6 | | 15-19 | 1 | | 10-14 | 1 | Solution: The first step is to identify the modal class, i.e., the class having the highest frequency. Through inspection, mode lies in the class interval 20- 24. The second step is to apply the formula. Here: 1 = 19.5, exact lower limit of the class interval 20-24 fm = 6, frequency of the class interval 20-24. f1 = frequency of the proceeding interval 1.e., 15-19. f2 = 4, frequency of the succeeding interval i.e., 25-29. i = 5, the, size of the class interval 20-25 Now put these values in the following formula M0 = I+ (f-m - /_l)/(2/_m - fy - f-2 ) xi Mode = 19.5 +(6 - l)/(2 X6-1-4)X5= 19.5 + 5/(12 -5) X 5 19.5 + 25/7= 19.5 + 3.57 = 23.0 The Characteristics of Arithmetic Mean: 1.    The mean is best known and well understood measure of central tendency. 2.    The value of the arithmetic mean is determined by every score in the distribution. 3.     It is greatly affected by extreme values. 4.    The sum of deviations about the arithmetic mean is zero. ie., Σ(X—M) = 0 5.    The sum of squares of the deviations from the arithmetic mean is less than those computed about any other point. 6.     In every case it has a determinate value. 7.     The standard error is minimum (Will be discussed later). The Advantages of Arithmetic Mean: 1.    The mean is the most commonly known and used measure of central tendency. 2.      Its calculation is not difficult. 3.     Its standard error is lowest. 4.    Mean can be manipulated algebraically. The Limitations of Arithmetic Mean: 1.    The value of arithmetic mean, maybe greatly distorted by extreme values and therefore, in such cases it may not be a typical value 2.     It cannot be calculated from open ended distributions. The Characteristics of Median: 1.     The Median is a central position. 2.     It is affected by number of values, not by the size of extreme values. 3.    The sum of the deviations about the median, signs ignored (all considered positive), will be less than the total about any other point. 4.     It is most typical when the central values of the data are closely grouped. The Advantages of Median: 1.     It is easy to calculate median. 2.     It is not distorted by extreme values. 3.     It is more typical measure of central tendency, because it does not depend upon quantity of the values. 4.     It may be calculated even when the distribution is open ended. 5.    Median is resistant to outliers. This makes it more reliable when dealing with skewed or asymmetric databases. The Limitations of Median: 1.    The median is not so widely used as the arithmetic mean. 2.    The items must be arranged according to the values of the variate (in ascending or descending order) before the median can be computed. 3.    The median is less reliable than the mean because it has large standard error than that of arithmetic mean (will be considered later). 4.    The median cannot be manipulated algebraically i.e. we cannot find the median of a combined group if medians of sub-groups are known. 5.    Median does not take into account actual values of the dataset beyond the middle point. The Characteristics of Mode: 1.    The mode has not most usual value. 2.    Mode is entirely independent of extreme values. 3.     It represents the value that occurs most frequently in a dataset. 4.      Unlike the mean, mode is not influenced by extreme values or outliers in the data. It simply reflects the value that appears most often. The Advantages of Mode: 1.     It is most typical and therefore the most descriptive value. 2.     It is simple to find mode by observation when there are a small number of items. 3.     It is not necessary to arrange the values if they are a few in number. 4.    Since extreme values are few in number, the mode is not affected by extreme values. The Limitations of Mode: 1.    There may be multimodal distributions. 2.     Its significance is limited when a large number of values is not available. 3.     In a small number of items the mode may tint exist, for none of the values may be repeated. Selection of An Appropriate Measure of Central Tendency: We observe that no one central value can be said to be good for all types of enquiries and under all conditions. In selection of a central value consideration must be given to the characteristics and limitations of various central tendencies. Each has its own field of importance and usefulness. In actual practice two or three values of a series maybe required for a proper understanding of the given data. In most of the cases, arithmetic mean would be found to be an ideal value, but if a very large number of items ma series have small values and only on or two items a big value, the arithmetic mean would give a fallacious conclusion, In such cases other values would give better results than the arithmetic mean If the purpose of enquiry is to study phenomena which are incapable of direct advantages in measurement, like intelligence etc, median has a distinct advantage If the enquiry relates, to say, "typical size or form' or 'average size of readymade clothes' or 'average outputs' then mode is the best value. SELF CHECK EXERCIS- 3 1 .Which of the following is included in measures of central tendency? a.    Average deviation b.    Mode c.    Standard deviation d.    None of these 2 .Mode is the number that occurs with: a.    Highest frequency b.    Lowest frequency c.    Medium Frequency d.    All of the above 3 . 3Median – 2 Mean is the formula of: a.    Average b.    Median c.    Mean d.    Mode 19.4    SUMMARY – In this chapter, we have learned about various measures of central tendencies (mean, median, mode) with their characteristics, advantages and limitations. We have also learned the process of computing the measures of central tendencies. 19.5    GLOSSARY- Grouped data: Grouped data refers to a set of data that has been organized into groups or intervals rather than listing individual values. Grouped data is commonly used in statistical analysis to simplify large datasets and identify patterns or trends. Ungrouped data: Ungrouped data refers to set of raw, individual data points that are not organized into groups or intervals. Ungrouped data is typically used when the individual values are important for analysis or when the dataset is relatively small and manageable without the need for grouping. Frequency: Frequency refers to the number of times a particular value occurs within a dataset or a specific category. 19.6    ANSWERS TO SELF CHECK EXERCISES SELF CHECK EXERCISE -1 Answer 1. A Answer 2. B SELF CHECK EXERCISE- 2 Answer 1. C Answer 2 – A SELF CHECK EXERCISE-3 Answer 1 – B Answer 2 – A Answer 3 – D 19.7 REFERENCES/ SUGGESTIVE READINGS: Best, J.W and J.V.Kahn, Research in Education (7th Ed.) New Delhi: Prentice Hall of India Pvt Ltd.1998 Sanswal, N.D. (2020). Research Methodology and Applied Statistics. (1st ed.). Shipra Publications 19.8 TERMINAL QUESTIONS: 1 .What are the various modes of central tendency? Explain briefly. 2 .What are the limitations of mode? 3 .What are the characteristics of mean? 4 .What are the advantages of median? 5 .Compute mean, median and mode for the following set of data. | Class Intervals | Frequencies (f) | |---|---| | 41-43 | 1 | | 38-40 | 4 | | 35-27 | 5 | | 32-34 | 8 | | 29-31 | 14 | | 26-28 | 17 | | 23-25 | 9 | | 20-22 | 13 | | 17-19 | 8 | | 14-16 | 3 | | 11-13 | 4 | | 8-10 | 0 | UNIT: 20 MEASURES OF VARIABILITY (RANGE, QUARTILE DEVIATION, STANDARD DEVIATION AND VARIANCE) 20.1    Introduction 20.2    Learning Objectives 20.3    Measures of Variability or Dispersion 1.    Range Self- Check Exercise -1 2.    Quartile Range Self- Check Exercise -2 3.    Average Deviation and Standard Deviation Self-Check Exercise-3 20.4    Summary 20.5    Glossary 20.6    Answers to Self - Check Exercise 20.7    References/ Suggestive Readings 20.8    Terminal Questions 20.1    INTRODUCTION: This chapter will delve into the fundamental concepts that quantify the dispersion or spread of data within a dataset. Understanding variability is crucial in statistical analysis as it provides insight into the distribution and reliability of datapoints. This chapter explores various measures that capture different aspects of variability, including range, variance, standard deviation, inter quartile range, mean, absolute deviation and coefficient of variation. Through clear explanations, this chapter aims to equip readers with the tools necessary to effectively interpret and analyze variability within datasets, enabling informed decision making and robust statistical inference. 20.2    LEARNING OBJECTIVES: After going through this unit, the students will be able to: 1.    Identify the conditions when different measures of variability should be used. 2.    Compute standard deviation for given data. 20 .3 Measures of Variability or Dispersion Central tendencies are central values around which the individual observations lie. They give us an idea of location of the distribution, but tell us nothing as to how the individual scores are scattered. Thus, each of the following six series, has 5 as the mean, though the patterns of individual measurements are different in all the series. (i)      5.5, 5, 5, 5, 5        (iv) 0, 9, 2, 2, 2, 3, 4, 12, 6. 10. (ii)    3, 4, 5, 6, 7           (v)     1.9, 10, 3, 7.0 (iii)    2, 3, 4, 5, 6, 7, 8 (vi) 6, 5, 5, 5, 5, 4 Thus, we observe that measurement of central tendency alone is not enough, we should also know how the individual scores are clustered around or scattered away from the central value. This characteristic of distribution is called variability. Dispersion or the spread is the degree to which the numerical data tends to spread about the average value of that data. Measures of variability are statistical measures that describe the spread or dispersion of a set of datapoints. Measures of variability refer to statistical tools used to quantify the extent of spread or dispersion within a dataset. They provide insight into how the individual datapoints are distributed around a central tendency such as the mean or median. These measures help to characterize the diversity or consistency of data points, which is essential for drawing meaningful conclusions and make accurate predictions in statistical analysis. The following are the main measures of variability or dispersion of the individual observations in the population. Range - The range of a distribution is the difference between the maximum and the minimum values of the variate. It is easily calculated. The ranges of the above six series are 0, 4, 6, 12, 9 and 2 respectively. It is readily understood and gives us some idea of the amount of dispersion present, but it is a crude and unstable measure of variability since it depends on two extreme values in the series and indicates virtually nothing about the general form of the series. Therefore, quite frequently, range misleads regarding dispersion of the distribution. Range= Maximum value – Minimum value EXAMPLE 1 Let’s calculate range for the given set of data: 74,    80, 92, 64, 70, 99, 82 Here, Maximum value = 99 Minimum value = 64 Range = Maximum value – Minimum value = 99 – 64 = 35 MERITS OF RANGE: It is simple to understand. It is easy to calculate. It is widely used in statistical quality control. DEMERITS OF RANGE: It cannot be calculated in case of open- ended series. It is not based on all observations. It is affected by extreme values in series. SELF CHECK EXERCISE-1 1 .Measure of variability helps to characterize: a.    Diversity of data points. b.    Similarity of data points c.    Both diversity and similarity of data points d.    None of the above 2 . The formula Maximum value – Minimum value signifies: a.    Mean b.    Median c.    Average d.    Range Quartile Range - Another way of describing the dispersion of a distribution is in terms of Quartile Range In determining this measure 25 percent of lowest values and 25 per cent of the upper values are disregarded. If Q1 is first quartile and Q3 is third quartile, then Quartile Range is Q3 - Q1. Half of this range is termed as semi-inter quartile range or quartile deviation (Q.D. or simple Q). Quartiles are the values that divide the data into 4 equal parts. Q1 is known as Lower Quartile Q2 is known as Middle Quartile or Median Q3 is known as Upper Quartile Let’s understand this with the help of an example: Example: Find the quartile deviation of the following set of data 62, 18, 22,11, 40, 41,70 Rearrange data into ascending or descending order 11, 18, 22, 40, 41, 62, 70 We can find Q1 by using the formula ¼ (n+1)th term Q1 = ¼ (7+1) = ¼ (8) = 8/4 = 2nd term i.e. 18 Hence, Q1 = 18 Similarly, we will find Q2 by using the following formula Q2 = ½ (n+1)th term =½ (7+1)th term =1/2 (8) = 8/2 = 4th term i.e. 40 So, Q2= 40 Now, we will find Q3 by using the formula Q3 = ¾ (n+1)th term = ¾ (7+1)th term = ¾ (8)th term = 6th term which is 62 So, we will compute quartile Range by using the formula Q3-Q1 = 62-18 = 44 The range and the quartile range both do not take into consideration the value of each individual item of the distribution and therefore, both lack in descriptive value. A good measure of dispersion should depend upon the amount by which the scores deviate from the measure of central tendency, say mean, the next measure depends upon deviation of each individual item from the mean. Merits of Quartile Deviation 1.I t is easy to calculate. 2.It is not very much affected by the extreme values of a series. Limitations of Quartile Deviation 1.It is a positional measure based on only 25th and 75th percentile. SELF CHECK EXERCISE -2 1.I n Quartile Range, data is divided into ____ equal parts a. Three b. Four c. Two d. Five 2.Q2 in Quartile Range is known as: a.    Lower Quartile b.    Upper Quartile c.    Middle Quartile or mean d.    None of the above 1 .AVERAGE DEVIATION AND STANDARD DEVIATION Average Deviation (A.D.) - Average deviation (A.D.) tells us how much, on average, each data point in a set differs from the overall average (mean). It’s like asking, "On a typical day, how far is each number from the center?" To find it, we: 1.    Work out the mean (average) of the numbers. 2.    Find how far each number is from that mean (ignoring negative signs). 3.    Take the average of all those distances. A The formula to compute A.D. is: A.D. = (ZMW Where: Ʃ|x| indicates magnitude of deviation by ignoring—ve sign. We calculate X-M = x for all the scores and get |x| by disregarding its proper sign (+or-). To calculate A.D. add all the |x| and divide the sum by N. Example - Calculate A.D. of the following values of X = 9,7,5,11,1,5,7,3 Solution | X | x = X – M | \|x\| | |---|---|---| | 9 | +3 | 3 | | 7 | +1 | 1 | | 5 | -1 | 1 | | 11 | +5 | 5 | | 1 | -5 | 5 | | 5 | -1 | 1 | | 7 | +1 | 1 | | 3 | -3 | 3 | | Ʃx = 48 | Ʃx = 0 | Ʃ\|x\| = 20 | Here : Ʃx = 48, N = 8, Mean = 6 and Ʃ|x| = 20 Therefore:  A.D. = (SklW =     = 2.5 A. D. from Grouped Data: The A.D. from grouped data can be calculated after calculating its mean. The deviations of the mid-points of the class intervals are calculated and multiplied by their respective frequencies. The sum (ignoring the signs) divided by N gives us the required A.D. value The formula is A.D = QyMwn The A.D. is rarely used in modern statistics, but is often found in the old experimental literature. Let’s understand this with an example Compute Average Deviation from the following set of Grouped Data | CLASS INTERVAL | FREQUENCY | |---|---| | 195-199 | 1 | | 190-194 | 2 | | 185-189 | 4 | | 180-184 | 5 | | 175-179 | 8 | | 170-174 | 10 | | 165-169 | 6 | | 160-164 | 4 | | 155-159 | 4 | |---|---| | 150-154 | 2 | | 145-149 | 1 | | 140-144 | 1 | | CLASS INTERVAL | MID POINT (X) | FREQUENCY (f) | fX | x | Fx | |---|---|---|---|---|---| | 195-199 | 197 | 1 | 197 | 26.2 | 26.2 | | 190-194 | 192 | 2 | 384 | 21.2 | 42.4 | | 185-189 | 187 | 4 | 743 | 16.2 | 64.8 | | 180-184 | 182 | 5 | 910 | 11.2 | 56 | | 175-179 | 177 | 8 | 1416 | 6.2 | 49.6 | | 170-174 | 172 | 10 | 1720 | 1.2 | 12 | | 165-169 | 167 | 6 | 1002 | -3.8 | -22.8 | | 160-164 | 162 | 4 | 648 | -8.8 | -35.2 | | 155-159 | 157 | 4 | 628 | -13.8 | -55.2 | | 150-154 | 152 | 2 | 304 | -18.8 | -37.6 | | 145-149 | 147 | 1 | 441 | -23.8 | -71.4 | | 140-144 | 142 | 1 | 142 | -28.8 | -28.8 | |---|---|---|---|---|---| | | | N=50 | Sum=8540 | | | Average Deviation= df^/^f) where x= X-M x is Deviation of scores from mean M= 8540/50 =170.8 X is Midpoint M is Mean A.D.= G/MWi) = 502/50 = 10.04 Merits of Average Deviation 1.    It is easy to calculate. 2.    Average Deviation reflects the variability or spread of the data. 3.    Calculating Average Deviation involves straightforward arithmetic operations. Limitations of Average Deviation Average deviation is sensitive to changes in sample size. Standard Deviation: (S.D. OR σ) The standard deviation is the most widely used measure of dispersion. It is positive square root of the mean of the squares of the deviations from the arithmetic mean and is denoted by a Greek letter a (sigma) and in short it is called S.D. SD or σ =<(Œa-M)A2)/^v) or (Zd^/N Where: σ: stands for standard deviation, X: is the value of variate M: is the arithmetic mean of the data, N: is the number of items d: is the deviation of mean from any score When the data are grouped in the form of frequency distribution, the above formula takes the following from. σ =^HX-M^/N) or 1/N J([(N£fd*2 ) - &fd*2 ^2]) I.     Calculation of Standard Deviation: The standard deviation is seldom calculated from the above formula because it is too time consuming. Unless the mean (M) is around number, the deviations (X-M) will be in fractional values, the squaring of which results in a prohibitive amount of calculation. The standard deviation, therefore, is generally calculated with the help of the following formula: σ = 1/N ^NXfd^2 ) - &fd*2 )A 2 ] ) Where: 'x deviation of the score (mid point) from the assumed mean (AM) in Gil units. f = frequency of different intervals or scores. N = number of cases Le sum (20) of all the frequencies or scores. i = size of the class intervals of the distribution. Example-1 Calculate the SD of the following values of X : 9,7,5,11,1,5,7,3 Solution: Taking score 5 as assumed mean, we obtain the following table: Here: Ʃx' = 8, | X | x1 = X-5 | x'2 | |---|---|---| | 9 | 4 | 16 | | 7 | 2 | 4 | | 5 | 0 | 0 | | 11 | 6 | 36 | | 1 | -4 | 16 | | 5 | 0 | 0 | | 7 | 2 | 4 | | 3 | -2 | 4 | | N=8 Scores | Ʃx'=+8 | Ʃx'2=80 | Ʃx² = 80 and N = 8 σ = 1/N V([(N£x'A2 ) - Q>y2 ] )     Put the different values in this formula. = 1                    =                   = 1      =     x 24 = 3 88 Note: Here size of the class interval is not to be used as we have only ungrouped observations. II.    Calculation of S.D. from Frequency Distribution or Grouped Data. When the data have been arranged in a frequency then the standard deviation is computed by the use of formula: σ =l/^V([(/VSA'A2)-E/xy2]) In the above formula: i    - represents size of the class interval x' - represents step deviations from the assumed mean in terms of class interval and N = Ʃf The steps required in the process of computing standard deviation from the distribution are outlined below: Step-1       Arrange the data in a frequency distribution table and add up all the frequencies to get N. i.e., Ʃf and find size of the class interval. Step-2       Complete the x' and fx' column to get Ʃfx' and Ʃfx'2. Step-3       Multiplying each f x' by the x' and write the result in the next column headed by fx'2. Add the products of this column to get to get Ʃ f x'2. Stop-4      Make use of the formula for σ Example -2. Calculate S.D. for the following distribution: | X | F | |---|---| | 48-52 | 2 | | 53-57 | 3 | | 58-62 | 5 | | 63-67 | 9 | | 68-72 | 10 | | 73-77 | 12 | | 78-82 | 7 | | 83-87 | 2 | |---|---| | 88-92 | 3 | | 93-97 | 1 | Solution: First, we can arrange the data as per our usual order. Assure 70 as assumed mean, which is the mid point of 68-72 class interval. | (1) C.I. | (2) f | (3) x' | (4) f x' | (5) fx'2 | |---|---|---|---|---| | 93-97 | 1 | +5 | +5 | 25 | | 88-92 | 3 | +4 | +12 | 48 | | 83-87 | 2 | +3 | = 6 | 18 | | 79-82 | 7 | +2 | +14 | 28 | | 73-77 | 12 | +1 | +12 | 12 | | 68-72 | 10 | 0 | 0 | 0 | | 63-67 | 9 | -1 | -9 | 9 | | 58-62 | 5 | -2 | -10 | 20 | | 53-57 | 3 | -3 | -9 | 27 | | 48-52 | 2 | -4 | -8 | 32 | | | Σf = 54 | | +48 13 -36 | Efx'2 = 219 | | | | | Σfx' = 13 | | |---|---|---|---|---| Here : Σfx1 = 13.    Σfx2 = 219   and Σf = 54. Step-1.      From column (1) i = 53-48 or 57-52 = 5 Step-2.     Taking mid-point of interval 68-72 as assumed mean and completing column (4), i.e. Σfx' = 13. Step-3.      Multiplying each number in the column (4) or 'fx' by the corresponding number in the column (3) or 'x' we get column (5) of fx2. Adding column (5) we get Σfx2 = + 219 Step - 4.     Put the required values in the formula. σ = 1/W y/([(N£fx'*2 ) - (EfxY2 ] ) σ = 5/54 V([(54 X 219) - (13)A2 ] ) = 9.99 Ans. We should remember that the measures of dispersion are expressed in the same units of measurement (such as inches, cm, grams, pounds etc. in which the observations themselves are measured. Exercise: Find out the SD of the following table Table-a | C.I. | F | |---|---| | 20-21 | 2 | | 18-19 | 2 | | 16-17 | 4 | | 14-15 | 0 | | 12-13 | 4 | Table-b | C.I. | F | |---|---| | 40-44 | 2 | | 35-39 | 3 | | 30-34 | 6 | | 25-29 | 10 | | 20-14 | 15 | | 10-11 | 0 | |---|---| | 8-9 | 4 | | | N=16 | | 15-19 | 20 | |---|---| | 10-14 | 16 | | 5-9 | 12 | | 0-4 | 8 | | | N = 92 | Effect on the Standard Deviation of Adding or Multiplying by a Constant: (a)    If a constant (positive or negative) is added to all the observations of a sample the same constant is added to the mean but the standard deviation remains uncharged. Let 'X' be the variable and constant 'C' is added to each value of 'X'. If 'M' denotes the mean of original values of X then the mean would be equal to (X+C) or (XC) which is clearly equal to X-M. This is the same as the deviations of the original variable from the mean. It proves that addition of the constant does not change the deviations The SD will remain the same Example- Let us consider small sets of 4 values, say 3.4.7, 10 and of 5, 6, 9, 12 The mean of the new set is '8'. The deviations from the mean in both the sets from their means are the same, which are -3, -2, -1 and 4 Therefore, the SD of both the sets is the same. (b)    If each score in a frequency distribution is multiplied or divided by a constant then the SD of the resultant distribution is also multiplied or divided to the same extent, as the individual scores are multiplied or divided by a constant. You should multiply the values in the preceding example and verify the above property. The standard deviation is the most commonly used measure of dispersion. The range and quartile range, though simple in calculation, are based on only two out of the whole bulk of observations Moreover, the range of a small sample will not agree with the range of the distribution from which it measures and is rigidly defined. In some applications of statistics, range is useful as in quality control. But in education, range and quartile range are seldom used. In fact, none of the measures of dispersion is as useful as the SD Most of the methods of statistical analysis have been evolved around the SD and its square. We use the standard deviation in the greatest dependability of the measures is required Variance: The term variance is used to denote the square of the standard deviation. Variance or σ² = (X(X - M)A2)/?V or i*2/N*2 [NƩ f x'2 – (ƩEX')2] in case of grouped data. 4.7.1  Uses of Measures of Variability(1)   Uses of Range: (i)    When the data are too scant or too scattered to justify the computation of a more precise measure of variability. (ii)    When knowledge of extreme scores or of total spread is all that is wanted. (iii)    Further calculations that depend upon it are likely to be needed. (iv)    Interpretations related to normal distribution curve are to be done. (2)    Uses of Quartile Deviation (i)    When the median is the measure of central tendency. (ii)    When there are extreme scores which have disproportionate influence on S.D. (iii)    When the concentration around the median is of primary interest (l.e. mid cases). (iv)    When the upper and lower class intervals are open (3)    Uses of A.D (i)    When it is desired to weigh all deviations from the mean according to their size. (ii)    When extreme deviations would influence the SD. unduly. (4)    Uses of S.D. (i)     When the statistic having the greatest stability is sought. (ii)    When the extreme deviations should exercise a proportionally great effect upon the variability. (iii)    When coefficients of correlation and other statistics are subsequently to be computed SELF CHECK EXERCISE- 3 1.Standard Deviation is denoted by a.    Alpha b.    Sigma c.    Rho d.    None of these 2.    Positive square root of the mean of the squares of deviations from the arithmetic mean is the formula of: a.    Average Deviation b.    Mode c.    Standard Deviation d.    Quartile Deviation 20.4    SUMMARY: In this lesson we have learnt about measures of variability and how to compute the measures of variability. Measures of variability refer to statistical tools used to quantify the extent of spread or dispersion within a dataset. They provide insight into how the individual datapoints are distributed around a central tendency such as the mean or median. 20.5    GLOSSARY: Deviation: Deviation refers to the act of departing or diverging from a standard, norm or expected course of action. Range: Range typically refers to the difference between the highest and lowest values in a set of data or the extent or variation between limits. 20.6    ANSWERS TO SELF CHECK EXERCISES: SELF CHECK EXERCISE-1 Answer 1. A Answer 2. D SELF CHECK EXERCISE-2 Answer1. B Answer 2. C SELF CHECK EXERCISE-3 Answer 1. B Answer 2. C 20.7 REFERENCES/SUGGESTIVE READINGS: Guilford. J. P. (1973). Fundamental Statistics in Psychology and Education (3 Ed). New York McGraw Hill Book Co Koul, Lokesh (1988) Methodology of Educational Research New Delhi: Vikas Publishing House Pvt. Ltd. 20.8    TERMINAL QUESTIONS: 1 .What are the various measures of variability? Explain briefly. 2 .Compute Average deviation for the given data sets: | Class      Intervals (Scores) | Frequencies (f) | Class   Intervals (Scores) | Frequencies (f) | |---|---|---|---| | 88-90 | 1 | 70-72 | 5 | |---|---|---|---| | 85-87 | 1 | 67-69 | 6 | | 82-84 | 2 | 64-66 | 8 | | 79-81 | 2 | 61-63 | 4 | | 76-81 | 5 | 58-60 | 4 | | 73-75 | 2 | | | ***** 308