--- title: "Complete Book 0" book: "Basic Stats with SLM" category: "General" publisher: "Ratan Prakashan Mandir Pvt. Ltd." type: "Educational Material" --- According to Latest Syllabus Read For Sure Success In University Examination RATAN TEXT BOOK BASIC STATISTICS M.A.Economics Sem-II Dr. Chandani Verma Published by Ratan Prakashan Mandir Pvt. Ltd. 2nd Floor, Centre Plaza, Parinay Kunj, Lajpat Kunj Marg, Agra-282002 Copyright Authors & Publishers Published by Ratan Prakashan Mandir Pvt. Ltd. 2nd Floor, Centre Plaza, Parinay Kunj, Lajpat Kunj Marg, Agra-282002 ISBN :978-81-69687-31-7 Price 610.00 only Printed at : KIDS INTERNATIONAL PVT. LTD. C-60, 61, 62, 63, EPIP, Shastripuram, Agra - 282007 Ph. : +91 9719004921 UNIT-1MEASURES OF CENTRAL TENDENCY STRUCTURE 1.1    Introduction 1.2    Learning Objectives 1.3    Measures of Central Tendency Self-Check Exercise-1 1.4    The Arithmetic Mean Self-Check Exercise-2 1.5    Weighted Arithmetic Mean Self-Check Exercise-3 1.6    The Median Self-Check Exercise-4 1.7    The Quartiles, Quintiles, Deciles and Percentiles Self-Check Exercise-5 1.8    The Mode Self-Check Exercise-6 1.9    Comparison of Mean, Median and Mode 1.9.1    Characteristics of Mean, Median and Mode Self-Check Exercise-7 1.10    The Geometric Mean Self-Check Exercise-8 1.11    The Harmonic Mean Self-Check Exercise-9 1.12    Summary 1.13    Glossary 1.14    Answers to Self-Check Exercise 1.15    References/Suggested Readings 1.16    Terminal Questions 1.1    Introduction Dear Students, Let me welcome you to the course on basic statistics which you will be studying through a course of a limited number of lessons. The basic attempt underlying these lessons is to make statistics as simple and clear as possible. Before I start explaining to you the various measures of central tendency and dispersion. I would like you to introduce to few basic concepts. “Statistics’ consists’ of two parts descriptive ‘statistics and statistical inference. Descriptive statistics deals wish the collection, organization and presentation of data, while Statistical inference deals with generalization from a part to the whole. Statistical /inference deals with the development of methods as well as their use. A population can be defined as the totality of all possible observations on measurements or outcomes, thus we may have human population, cattle population, population of students enrolled in Himachal Pradesh University etc. A population may be either finite or infinite. Related to the concept of a population is the concept of sample, which is a set of measurements or outcomes selected from the population. An important type of probability sample is the random sample. Both populations and samples can be described by, stating their characteristics. Numerical characteristic of a population are called parameters; the characteristics of a sample, given in the form of some summary measure; are called statistics. With respect to a phenomenon that can be measured, is known as a variable which means, a homogeneous quantity that can assume different values at different points of observation. If a phenomenon can only be counted but not measured, we speak of an attribute. The definition of a variable, stresses the possibility of variation at different point of observation. On the other hand, a quantity that cannot vary from one observation to another is called constant A continuous variable is a variable that can assume any value on the numerical axis or a part of it. In contrast to a continuous variable, a discrete variable is one that can assume only some specific values on the numerical axis. The final concept to be introduced at this stage is that of distribution. In the case of a sample we have & frequency distribution, while in the case of population we speak of a probability distribution. A distribution which confines to one ‘variable—is known as an univariate distribution, whereas when it deals with two or more variables, we have a bivariate and multivariate distribution. 1.2 Learning Objectives After going through this unit, you will be able to •    Understand and Calculate Measures of Central Tendency •    Analyze and Interpret Different Measures of Central Tendency •    Apply Measures to Diverse Data Sets 1.3 Measure of Central Tendency With this unit, we begin our formal discussion of the statistical ‘methods for summarizing and describing numerical methods for summarizing and describing numerical data. The objective here is to find one representative value which can be used to locate and summarize the entire set of varying values. This one value can be used to make many decisions concerning the entire set. We can define measures of central tendency (or location) to find some central value around which the data tend to cluster. Self-Check Exerfcise-1 Q1. What do you mean by measures of central tendency? Explain their significance in statistical analysis. Q2. Describe the different measures of central tendency and discuss situations where each measure is most appropriately used. Measures of central tendency, enable us to get an. idea of the entire data. For example, it is impossible to remember, individual earnings of crores of earning people in India. But if the average income is ‘obtained, a single value will represent the entire, population. These measures ‘also enable us to compare two or more sets of data to facilitate comparison. For example, the average production figures of one particular year may be compared with the production figures of previous years. A good measure of central tendency should possess, as far as possible, the followingproperties. (i)    It should be easy to understand (ii)    It should be simple to compute (iii)    It should be based on all observations (iv)    It should be uniquely’ defined (v)    It should be capable of further algebraic treatment. (vi)    It should not be unduly affected by extreme values. Following are some of the important measures of central, tendency which are commonly used. Arithmetic mean Weighted Arithmetic mean Median Mode Geometric Mean Harmonic Mean Arithmetic Mean The arithmetic mean (or mean or average) is the most commonly used and readily understood measure, of central tendency. In statistics the term average refers to any of the measures of central tendency. The arithmetic mean is defined as being equal to the sum of the numerical values of each and every observation divided by the total number of observations. Symbolically, it can be represented as : V - Sx X n where Ex indicates the sum of the values of all the observations, and N is the total number of observations. For example, Let us consider the monthly sale of ten firms in Rs. lakh 10, 18, 20, 28, 30, 25, 40, 30, 20, 10. If we compute the arithmetic then) X— 10 +18 + 20 + 28 + 30 + 25 + 40 + 30 + 20 +10 10 221 — i0- = Rs. 22.10 lakhs. Therefore the average monthly sale is Rs 22.10 Lakhs. 1.4    The Arithmetic Mean Example I : 100 crates of golden delicious ‘large’ variety of apples are sold at Rs. 40. big case (of 18 kg) another 250 crates of ‘medium’ variety at Rs. 30 per big case and 400 cases of ‘small’ variety at Rs. 25 per case, in the Simla market on a particular day. Then (100x 40) + (250 x 30) + (400x 25) X —---------------------- 100 + 250 + 400 21500 500 = Rs. 28.67 kg. per case. symbolically, if x1, x2, x3 ................. xn are the values of a variable, the mean is computed by the formula X1 + X2 + X3 +..........Xn X N n ZXi/N                                     (1) i=1 Where Z = the sum of X = the mean of values, x1 = the value of variables, N = Number of values. Self-Check Exercise-2 Q1. The pocket allowance for 10 boys is Rs. 15, 20, 30, 22, 25, 18, 40, 50, 55 and 65. Find out the average pocket allowance. Q2. Eight persons earn the following income: 30,    36, 34, 40, 42, 46, 54, 62 Find out the arithmetic mean. 1.5    Weighted Arithmetic Mean In computing the arithmetic mean we give equal importance to each item in the series. This equal importance may be misleading if some items in a distribution are more important than others. In such cases ‘various items’ are given proper weights. The weights assigned to each item being proportional to the importance of the item in the distribution. For example, if we want to have an idea of the change in the cost of living of a certain group of people, then the simple mean will not give the correct result, because the commodities to be selected will, not be of equal importance. It is, therefore, necessary to calculate weighted mean in such cases. So that weighted average or weighted arithmetic mean W1X1 + W2X2 +..........+ WnXn X W1 + W2 + W3 +.........+ wn wx x —--- ZW .................. (2) where W1, W2, W3 +.......... Wn stand for the weights of the items X1, X2, X3........Xn respectively. Example 2 : Table 1.1 shows the heights of the students in a class. Table 1.1 Heights of Students in Economic Class | Heights in Inches X | Number of Students | |---|---| | f | fx | | 59 (X1) | 2 (f,) | 118 | f1X1 | | 60 (X2) | 5 (f2) | 300 | f2X2 | | 61(X3) | 12 (f3) | 732 | f3X3 | | 62 (X4) | 20 (f4) | 1240 | f4X4 | | 63 (X5) | 35 (f5) | 2205 | f5XS | | 64(X6) | 48 (f6) | 3072 | f6X6 | | 65 (X7) | 30 (f7) | 1950 | f7X7 | | 65 (X8) | 18 (f8)- | 1188 | f8X8 | | 67 (X9) | 7 (f9) | 469 | f9X9 | | 68(X10) | 3(f10) | 204 | f10X10 | | | Ef=180 | EfX= | 11.478 | Arithmetic Mean of the series fiXi + f2X2 + f3X3..........+ f1Xn X fl + f2 + f3 +fn+ fn EfX V = N ............ (3) 11.478 180 = 63.77" app. Short-cut Method We can find the arithmetic mean by another method, where ‘we assume one of the heights to be the mean and then add a collection factors, with the help, of the following formula : X = Xa + ED N (4.1) (for data without any weights or frequencies) or X = Xa + EfD Ef (= N) (for data with frequencies) ........... (4.2) and correspondingly X = X + EwD Ew(= N) (for data with weight) ........... (4.3) Where Xa stands for the assumed mean D stand for the value of the deviation of the variate from the assumed mean = X - Xa , f stands for frequencies, w stands for the weights. Let us apply this method to the data of Table 4.1. We assume the mean to be 64". Normally, the assumed mean should be the value of the variate with the highest frequency or the one, which lies in middle, In case of & symmetrical distribution or mildly skewed series, both these would tend to coincide. Our assumed mean 64" has the highest frequency (48) and is almost in the middle of the series.’ Worksheet I based on Table 1.1 for computing Mean Height of Students in Economics Class - 42 142 | Height (X) in inches (1) | Frequency (f) (2) | Deviations from Xa - 64 D=X— Xa (3) | fxD (4) = (2), (3) | |---|---|---|---| | 59 | 2 | — 5 | — 10 | | 60 | 5 | — 4 | — 20 | | 61 | 12 | — 3 | — 36 | | 62 | 20 | — 2 | — 40 | | 63 | 35 | — 1 | — 35 | | 64 | 48 | 0 | 0 | | 65 | 30 | +1 | + 30 | | 66 | 18 | +2 | + 36 | | 67 | 7 | +3 | + 21 | | 68 | 3 | +4 | + 12 | | | v = N = 180 | | -141+99 = -42 | X = X + f = 64 + aN — 64 — 0.23 = 63.77". In this method, we avoid multiplication of big number. The Arithmetic Mean from Grouped Date : Long Method When dealing with a frequently distribution, we do not ordinarily have the original date from which the frequency distribution was made. | Sales Group (Rs.) | Number of Sales Zones f | Midpoint of the group (Rs.) X | fX X | |---|---|---|---| | 1,100—4,500 | 1 | 2,800 | 2,800 | | 4,600— 8,000 | 4 | 6,300 | 25,200 | | 8,100—11,500 | 10 | 9,800 | 98,000 | | 11,6007—15,000 | 11 | 13,300 | 1,46,300 | | 15,100—18,500 | 6 | 16,800 | 1,00,800 | | 18,600 —22,000 | 4 | 20,300 | 81,200 | | 22,100—30,000 | 2 | 26,050 | 52,100 | | 30,100—38,000 | 2 | 34,050 | 68,100 | | | 40 | | 5,74,500 | -   SfX    574500 X = WN = -40“ = Rs- 14’362'50 Thus the average sale of the 40 sale zones is Rs. 14,362.50. It is obvious that for finding out the mean of grouped data, we have first to find out the mid value with its corresponding frequency, sum up these products and divide it by the sum ‘of frequencies. Whenever the two values for X do not agree, it is due to inadequacy of the mid-value assumptions. “It is almost always true’ that none of the mid values is actually the true concentration point of its class. For groups to the left of the group of maximum frequency the mid value of the group to the right of the group of maximum, frequency, the mid-value of the a group frequently exceeds the mean of the group. Although all the mid value assumptions, are usually incorrect, there is a definite tendency for the error to offset each other provided the distribution is approximately symmetrical.” Shortcut method for computing mean Let us now compute the value of the mean with the help of an assumed mean. Mean i X = X a + f N Where Xa stands for the assumed mean D stands for the deviations of the mid values, from the assumed mean i.e. = X - Xa N = S f Let us assume the value of fee mean to be Rs. 13,000, the mid value with the maximum frequency, and calculate the mean, as shown in work sheet No. 3. Worksheet No.3 for Calculating the mean (by Short-Cut Method) | Sales group (Rs) | Mid-value of the group (X) | Frequency (f) | Deviation from Xa= 13,300 (D) | f X D | |---|---|---|---|---| | 1,100-4,500 | 2,800 | 1 | —10,500 | —10,500 | | 4,600—8,000 | 6,300 | 4 | —7,000 | —28,900 | | 8,100—11,500 | 9,800 | 10 | —3,500 | —55,000 | | 11,600—15,000 | 13,300 | 11 | 0 | 0 | | 15,100—18,500 | 16,800 | 6 | + 3,400 | + 21,000 | | 18,600—22,000 | 20,000 | 4 | + 7,000 | + 28,000 | | 22,100—30,000 | 26,050 | 2 | + 12,750 | + 25,500 | | 30,000— 38,00 | 38,050 | 2 | + 20,750 | + 41,500 | | | | 40 | | —73,500 | | | | | | + 116,000 | | | | | | + 42,500 | * FE. Croxton, D.J. Cowden and Sidney Klein : Applied General Statistics. Third Edition. 1971. p. 159. X = Xa + — aN 42,500 = 13,300 +   ’    = 1,300 + 1062.50 40 = Rs. 14,362.50 (the same as found out with the direct, method) Another short-cut method for computing the mean, goes for further simplification: by dividing the deviation (D’s) by the highest common factor (i); among the class interval sizes, i would equal the common class interval when, the size of all class intervals in a frequency distribution is uniform. However, when the size of class intervals varies, it will stand for the highest common factor, among, them. Let us show its working sheet No.4. Here, ^fDx X — X a + ---X i N ............... (4.4) D Where D' = — stands for deviation of mid values from the assumed mean in term of i. i Worksheet No. 4 for Calculating the Mean | Sales group (Rs) | Mid-values of the group (X) | Frequency (f) | Deviation D = (X = Xa ) | D D' = — i where i = 250 | fD' | |---|---|---|---|---|---| | 1,100—4,500 | 2,800 | 1 | — 10,500 | — 42 | — 42 | | 4,600—8,000 | 6,300 | 4 | —700 | — 28 | — 112 | | 8,100—11,500 | 9,800 | 10 | —35,000 | — 14 | — 130 | | 11,600—15,000 | 13,300 | 11 | 0 | 0 | 0 | | 15,100—18,500 | 16,800 | 6 | + 35,000 | + 14 | + 84 | | 18,600—22,000 | 20,300 | 4 | + 7,000 | + 28 | + 112 | | 22,100—30,000 | 26,050 | 2 | + 12,750 | + 51 | + 102 | | 30,100—38,000 | 34,050 | 2 | + 20,750 | + 83 | + 166 | | | | 40 | | | —294 | | | | | | | + 464 | | | | | | | + 170 | Now ^fDx X — X a + X i N 170 = 13,000 + — x 250 + 13,300 + 1062.50 40 = Rs. 14362.50 Since when classes very in width, the distribution is invariably skewed, and as skewness, increases our mid value assumptions offset each other less closely and therefore, the mean worked out from a frequency distribution with unequal class intervals may differ markedly from the mean computed from the unclassified date. 1.5.1. Prostrates of the Arithmetic Mean 1. 2. 3. An important property of the mean is that is algebraic some of the deviations of fee various values from the mean is equal zero, i.e. Yd = 0 Where d stand for deviations of the values from the mean = X - X The sum of the square of fee deviations of a set of number X, from any number a is a minimum if and only if a = X . If f1 numbers have, mean m1, f2 numbers have mean m2.........fk numbers have mean mk, then the mean of all the numbers combined is given by X = fm + f2m2 +.... + fkmk f + f +......f 12    k Self-Check Exercise-3 Q1. Define weighted arithmetic mean. Q2. A student obtained 60 marks in English, 75 in Hindi, 63 in Mathematics, 59 in Economics, and 55 in Statistics. Calculate the weighted mean of the marks if weights are respectively 2, 1, 5, 5 and 3. 1.6 The Median An important characteristic of the mean is that it is affected by all the values, especially by extreme values, Assume 5 groups with the following income: Rs. 20,000, Rs. 25,000, Rs. 27,000, Rs. 28,000, Rs. 1,50,000. The mean is X = 1/5 (20,0004-25,000 + 27,000 + 28,000 + 1,50,600) = Rs. 50,000. The mean of the first four incomes is Rs. 25,000, but with the inclusion of the extremely high income the mean jumps sharply to Rs. 50,000. It is obvious, that this mean of Rs. 50,000 does not adequately represent values in the frequency, distribution. In such cases where the frequency distribution is skewed and has extreme values a measure of central location or tendency called the median, is in many, cases, more appropriate. Media is affected by the position, but riot by the size of the Items. The median is positional measure of central tendency; it divides the series into two equal halves when presented as an array. The median is easily found out by arranging the data in the form of an array and locating the observation, that has just as many observation above it, as below it, if the n +1 number of observations in the series is odd, e.g. 37 or 75, then the Median Value of 2 th observation. If the number of observations in the series is an even number there, will be no one value that divides them into two equal parts. In such a case, the median = arithmetic mean of the two middle N ^ N 1^ values — th and I y + 1 I th. and the median value, may not coincide with an actual value in the series. Before going on to consider the computation of the median for grouped, data, let us calculate the value of the median for the, sales of the 40 Divisional Sales Managers HIMGU Products arrayed 40 in Table 1.1: We want to find the value which is so located that = 20 items will be on either side 2 of it. This is, of course, the value of the arithmetic mean of the 20th and 21st items and counting from either side reveals that the value of the median is 12,950. 1.6.1    The Median From Grouped Data For computing the value of the median for a frequency distribution, we count half, of the frequencies from either end of the distribution, in order for determine the value, on either side of which half of the frequencies fall. We locate the median, by interpolation, with the help of the following formula: i ( N +1 Median = L + r c        cfi (5.1) 1    f median \ 2 where l1 = lower class boundary of median class (i.e. the class containing the median). i = size of the median class interval. fmedian = frequency of median class. N = Number of items in the data ( =Sf ); Cf1 = cumulative frequency of the group preceding the median class. Geometrically, the median is the value of X. (abscissa), corresponding to that vertical line which divides a histogram of frequency curve into two parts having equal areas. This value of X is sometimes denoted by Xm. Let as now compute the value of the median for our sales in 40 zones. Worksheet 9, for computing the value of the median | Soles Group | Sales Group | Sale Zones | Cumulative | |---|---|---|---| | (Rs.) | class Boundaries | Frequency f | frequency | | (1) | (2) | (3) | (4) | | 1,100—4,500 | 7,050—4,550 | l | 1 | | 4,600—8,800 | 4,550—8,050 | 4 | 5 | | 8,100— 11,500 | 8,650—11,550 | 10 | 15 | | 11,500— 15,000 | 11,550—15,050 | 11 | 26 | | 15,100—18,500 | 15,050 —18450 | 6 | 32 | | 18,600—22,000 | 18,550—22,050 | 4 | 36 | | 22,100—30,000 | 22,050—30,050 | 2 | 38 | | 30,100—38,000 | 30,050—38,050 | 2 | 40 | | | | 40 | | 40 +1 Median Sales Zone =       20.5th zone. 2 If we take column 2, 3, and 4 together; we can describe the frequency table as under : In one zone, sales are less than Rs. 4,550 : in 5 zones, sales are below Rs. 8,050 ; in 15 zones are below 11,550; The fourth sales group Rs. 11,550 to 15,050 has 11 sales zones, starting from 16th to 264. Hence, our median sales zone i.e. 20.5th. zone lies in this group. Having located the median class, we can now very easily calculate ‘the value of the median, by inter-polation. i    (N +1 Median = L = r        I ~ 1   f median \ 2 Now l1 = 11,550, i =3,500, fmedian = 11 N +1    40 +1 2 2 = 20.5 and Cf1 = 15 3.500 Hence Median = 11,550 + 11 (205 - 15) =11,350 + 1750 = Rs. 13,300. We would get exactly the same result if we start from the other end of the median group. i     ( r N + 1^ Median = L — r        I cfm  ~   I(5-2) 2    f median \     2  J                                  v7 where l2 = upper real limit of the median group and Cfm = the cumulative frequency of the median group. 3.500 Hence Median = 15,050 —  11 (26 — 22.5) = 15,050 — 1750 = Rs. 13,000 (the same as found with the help of formula 5.1) The value of the median obtained from the frequency distribution Rs. 13,300 is in fairly close agreement with that of 12,950.00 found from the array. “Unless the data contains gaps or irregularities; we can expect rather close agreement when dealing with, a continuous variable, and likewise for a discrete variable if the data are not broken.” We have now computed the value of the mean and the median for the frequency distribution of zonal sales of HIMCU Products. The mean was Rs. 13,362.40 and the median Rs. 13,300. The mean exceeds the median because the distribution is skewed to the right. If a distribution is symmetrical, the mean and median are equal. The median is not affected by the presence of unequal or open-endclass intervals. The median of a frequency distribution can also be located graphically, with the help of the ogive in the following way: (a) Compute N +1 2 and locate, this point on the vertical scale. (b)    At this point draw a perpendicular to the Y = axis and extend it, so that it intersects the ogive. (c)    Drop a perpendicular to the X-axis from the point of intersection and read the value of the median on the X-axis’. The value of the median, located graphically, will be approximately equal to the value computed arithmetically. Self-Check Exercise-4 Q1. Find out the median in the following series. | Size (less than) | 5 | 10 | 15 | 20 | 25 | 30 | 35 | |---|---|---|---|---|---|---|---| | Frequency | 1 | 3 | 13 | 17 | 27 | 36 | 38 | Q2. Calculate the median from the following series | Class Interval | 10-20 | 20-40 | 40-70 | 70-120 | 120-140 | |---|---|---|---|---|---| | Frequency | 4 | 10 | 26 | 8 | 2 | 1.7. The Quartiles, Quintiles, Deciles and Percentiles The median, we have, seen, divides a set of data into two equal parts. An extension of this idea is dividing the set into four (quartzes), five (quintiles), ten (deciles) or hundred (percentiles) equal parts. These values of quartiles, denoted by Q4, Q2 and Q3 are called the first (or lower), the second, (or middle), and third (or upper) qualities respectively, the value Q2 being equal to the median. Similarly, the deciles are denote by D1, D2.......D9, and the percentiles are denoted by P1, P2.......... P99. D5 = P10 = Q2 = Median. Similarly, P25 and P75 correspond to the lower and upper quartiles respectively. Collectively quartiles, deciles, percentiles and other values obtained by equal subdivision of the data are called quintiles, Quintiles are computed in the same way as the median, 1.    We first divide the total number of observations by, the number of equal parts into which the series is to be divided, by 4, 5, 10 or 100 for quartiles, deciles and percentiles respectively.’ 2.    Multiply by the order of this quantile i.e. by 1 in case, of first quartile, decile etc. by 3 in case of third quartile, by 7 in case of 7th decile or percentile; and by 56 in case of P56. 3.    Determine with the help of cumulative fluency column the class interval in which our positional, measure lies, and 4.    Interpolate the value of tine quartile, in the same way as the median. Let us compute the value of some, of these positional measures, for the data on zonal sales, with the help of worksheet No 6 (given earlier). i ( N Q1 = l'+ f117 - ^f)                (6) where l1 = lower limit of the first quartile group i = the size of the first quartile class interval fq1 = frequently of the first quartile group Cf2 = cumulative frequency of the group proceeding the first quartile group Now, for the zonal sales data The 10th zonal sales lie, in the group, 8,100-11,500 (where 6th to l5th zonal sales lie) whose real ; class limits are 8,050 and 11,550 Hence l1, = 8,050, i = 3,500 fq1 = 10 and Cf1 = 5 3 500 Therefore, Q1 = 8,050 + 10 (10—5) = 8 0-50 + 1,750 = Rs 6,800 *Ibid p -866 Similarly; i (3N r’/-^ Q3 = l1 + fq3 [ 4   Cf1 J                                (7) 3N_ 3 + 40 Here we first locate -4- - —4— = 30 zonal sales and we find that it lies in the group, 15,100— 18,500 whose real class limits are 15,050 and 18550. Hence i1 = lower boundary of the third quartile group = Rs. 15,050 i = 3,500, fq2 = 6, and Cf1 = 26. 3,500 Therefore, Q = 15,050 +       (30-26) 36 = 15,050 +2,333.33 = Rs. 17,383.33 We may also, similarly, compute the value of 7th decile (D7) and 56 percentile (P56). i (7 N r/^ 7-1   fD7[ 10   ^ J                          (8) 7N The 7th decile item i.e. 10 th = 28th item lies in the group 15,100 — 18,500, whose class boundaries are 15,050 and 18,550 (the same as for Q3). Now l1 = 15,050, l = 3,500, f D7 = 6, Cf1 = 25. Hence 3,500 D = 15,050 +      (28—26) 76 = 15,050 + 1,157.14 = Rs. 16,207.14 n , i <56N P56 -11 + T1100- Cf1 J                           (9) P56 56 N 56 x 40 100 -   100 = 22.4th item lies in the group, 11,600 — 15,000, with class boundaries as 11,550 and 15,050 (the same as median group). Now l1 = 11,550, i = 3,500, f p54 = 11, Cf1 = 15. Hence. 3.500 -P54, = 81,550 + 11 (22.4—15) = 11,550 x 2,354.54 = Rs. 13,904.54. The quartiles can also be located on an ogive in the same way as the median. Self-Check Exercise-5 Q1. Define a)    Quartiles b)    Deciles c)    Percentiles Q2. From the following data, calculate Q1, Q3, D5 and P25 21,    15, 40, 30, 26, 45, 50, 54, 60, 65, 70. 1.8.    The Mode Another measure of location of a frequency distribution is the mode. The mode of a distribution is any value at which the frequency density is at a maximum. This implies that it is any value of the variable that occurs most frequently. In simple terms we may say that mode is the most popular value of the variable, i.e. a value which occurs maximum number of times, it further implies that if the frequency curve has one peak (i.e.) one maximum as in the figure 4.1 (a), there is only one mode, whereas if the frequency curve has, two (or more) peaks (i.e. two or more maxima), as in the figure 4.1 (b) the distribution has two (or more) modes. On the other hand, if we have a rectangular distribution* there is no mode, Example 4. Let us suppose a student took 5 tests in a semester with the results : 58,    70, 70, 75, 85. Then the mode is there but not so well defined 70 occurs twice while all other scores occur only once. If, however, the scores are: 59,    70, 73, 75, 85, there is one mode. Since the mode is the most typical value of a series of values, it is not at all affected by the presence of one or a few extreme values. Further, the mode cannot be readily located in an ungrouped data or even an array. For grouped data, the mode may be located more readily. While there are several ways of computing the mode, it is usually sufficient for practical purposes to use the midpoint to the modal group. In some cases, it may suffice to say that the mode lies between such and such limits. If, however, we have to determine the modal value of the point of maximum concentration, this can be done with the following formula: M0 = l1 + fm + fl (fm - fl) + (.fm — f2) x i (...10) where l1 = lower limit of the modal group. fm = modal or maximum frequency. f1, = frequency of the group preceding the modal group. f2 = frequency of the group following the modal group. i = size of the modal class interval. For our zonal sales data, the modal group will be 11,600—15,000, since it has the highest frequency (11). The class boundaries of this group are 11,550 and 15,050. Hence l1 = 11,550, i = 3,500 fm -11, f1, = 10 f2 = 6. 11 -10 Therefore, M0 = 11,550 + x 3,500 (11 -10) + (11 - 6) = 11,550 + 1 x 3,500, = 11,550 x 583.83 6 = Rs. 12,133.33. The mode is nearer to the lower class limit than to the upper one of the modal group because the frequency of the group following the modal group (6) is much less than that of the group, preceding the modal one. (10). The interpolation of mode shown above can also be located graphically, as in fig. 1.2. Fig. 1.2 The mode is easy to compute and may to applied to qualitative as well as quantitative data. One may be investigating, for example, consumer preferences for six brands of tea, A, B, C, D, E, and F. let the preferences be A = 15, B = 35, C = 25, D = 50, E = 40, F = 45 Here the modal preference is tea D. Suppose a garments store wishes to stock men’s shoes. An investigation shows that size 6 has the greatest demand. This is the modal value of the distribution of shoe sizes. A study of the local and suburban passengers using the local and suburban transport system may show a bimodal distribution, because there will be two peak hours, one in the morning, the other in the evening. Self-Check Exercise-6 Q1. Define Mode. Q2. Find the mode from the following data: | Class: | 155-157 | 157-159 | 159-161 | 161-163 | 163-165 | |---|---|---|---|---|---| | Frequency | 4 | 8 | 26 | 53 | 89 | 1.9. Comparison of the Mean, Median, and Mode The relationship between the mean, median and mode for a unimodal frequency distribution is shown in fig. 1.3. When a distribution is symmetrical, the mean, median and mode coincide. When a distribution is skewed to the right, t hen Fig. 1.3 (b). Mean > Mode. Mean (Rs. 16.362.50) > median (Rs. 13,140.91) > mode (Rs. 12,133.33). For example, income distribution is frequently skewed to the right, where the majority of families have incomes concentrated at lower levels and then the number of families tapers off as the income goes up. In this case, the mean is pulled up by the extreme high incomes and the relation among the three measures is as stated above i.e. Fig. 1.3 Mean > Median > Mode For unimodal frequency curves which are moderately skewed, the median travels 2 3 rd the distance between the mean and, the mode and we have the following empirical relation : Mean - Mode = 3 (Mean - Median)        ......(11) When a distribution is skewed to the left, then as in fig. 4.3. (c) mode > median > mean. The grades of a class where the majority have high grades, with only a few with low grades, is an example. Here the mean is pulled down below the median by the extremely low grades. 1.9.1.    Characteristics of Mean, Median and Mode* We may now briefly review the relative place of the three measures of central tendency— Mean, Median and Mode : (i)    The arithmetic mean is the most widely known and used, and whenever, common people talk of an average, they nave the arithmetic mean in mind, Median ‘is less known than the mean, but it is based on a simpler concept. Similarly, the mode is less well known; however, the concept of the mode is the most typical of a group of items. (ii)    The arithmetic mean possesses certain mathematical properties, which the two other measuresdo not. We have already referred” to the two important properties of the mean viz, Sdx = 0 and Sdx2 = minimum. These properties have become the basis of reference for measures of dispersion (to be discussed in the next lesson).” The use of the mean is extremely important for fitting the normal curve. Median is also sometimes used as the basis of some measures of dispersion, since the sum of the, absolute deviations from it is minimum. (iii)    Among the three measures, it is only fee arithmetic mean which permits algebraic, treatment. SX Since, X - ^- we can compute the value of any of the three parameters, X , SX and N, if the other two are known. Thus, X = —; EX = N.X N EX and N = — . X Further, if we know a series of arithmetic means, we can compute the arithmetic mean of the combined series, using appropriate weights. Thus NX = N1Xi + N2X2 + N3X3+  NnX -  N1Xi + N2X2 + N3X3 +.... + N n Xn or X =----------------------- N = (Ni + N2 + nn ) (iv)    The arithmetic mean may by calculated from ungrouped data, from arrayed data, from the frequency distribution or as .noted in (iii) above, merely, from our knowledge of SX and N. Further, the mean from the grouped data tends closely to approximate that from ungrouped data, the closeness being greater, the more symmetrical is the distribution. Median cannot be computed form ungrouped data, without first being put in the. form of an array. Its value, as computed from a frequency distribution, will agree approximately with that from a ray, if the items in the median class are distributed regularly. The mode is easily located in a frequency distribution, but with some difficulty in an array. Further, the interpolation of the mode in the modal class is at best only an approximation. (v)    When skswness is present in a distribution, the class intervals, generally, cannot be of uniform size. Thus, lack of uniformity in the size of class intervals tends to make the mean from the grouped data deviate from that computed from ungrouped data. Since most of the frequency distributions are skewed to the right, the value of X from these distributions, tends to exceed that from ungrouped data. The median generally remains unaffected by the variation in the width of the class interval. The upper quartile and the higher deciles and percentiles may, however, get affected by this phenomenon, making the values concerned, less reliable. The mode can be located fairly satisfactory, if the class intervals, following and preceding the modal class, have the same width as that of the modal group. Otherwise, modal’s value becomes less reliable: As stated earlier, the presence of skewness affects the mean, the most, and the median, less. The mode is hardly affected by it. Further the skewness pulls the mean and median in die-direction of the excess tail i.e., of skewness. (vi)    The existence of a few extreme values affects the mean greatly, the median is only (slightly affected and the mode is not at all influenced. Because of this characteristic, mean is not a suitable measure of central’ tendency for income distribution, especially if a few persons, at the top have incomes far exceeding those of most people. Similarly, wherever we have reason to suspect heterogeneity in data, the median shall be preferable to the mean. Again, whenever we have a series in which we know the number, but not the exact value of the extreme items, the median and the mode can be determined, but ‘not the mean.’ (vii)    The presence of open end classes makes the value of mean, largely conjectural because the mid value of these classes cannot be determined. Since these open-end classes occur at the beginning or at the end of a frequency distribution they hardly affect the median. Open-end classes do not affect the determination of the mode, unless the distribution is of J-type or J-reverse type in which case mode cease to be a measure of central tendency. (viii)    The arithmetic mean is hardly a suitable measure of central tendency for a frequency distribution if the data, are irregular or broken, because it would tend to deviate for from the mean from ungrouped data! The median, for its location required that data to be regularly distributed over the median class; hence, if the class in which median lies, has gaps or the data are irregularly distributed median would not be an appropriate measure of central tendency. If the data around the mode has gaps, the mode is not likely to be well defined. But irregularity of gaps elsewhere do not affect the mode. As described above, the choice of the measure of central tendency depends upon (a) the nature of the distribution of the data and (b) the concept of central tendency, desired for a particular purpose. With symmetrical or near symmetrical distributions any of the three measures may be used. With skewed distributions, the mean, not being a typical value, the median or the mode are to be preferred. Other things being, equal the mean is preferable to other measures. Besides the three measures of central tendency, the mean, the median, and the mode, we use two other measures occasionally in business and economics. These are the geometric mean and the harmonic mean. Of these, the geometric, mean is more, important and is used for averaging rates of change and constructing index numbers. Self-Check Exercise-7 Q1. Discuss the relationship between mean, median and mode. Q2. What are the characteristics of mean, median and mode? 1.10 The Geometric Mean The geometric, mean (G) may be defined as the nth root of the product of n items i.e. G = n VKX1X2X3......X,,)] = (X1X2X3........Xn)1/n                                 ......... (11) Logarithms, are used to calculate the nth root 1 Thus log G = n [log X1, + logX2 + .........+ log Xn] 1 or log G = , Slog X                             .......(12.1) or G = Antilog |1S logX|                             (12.2) For frequency distributions, each logarithm must be multiplied by the corresponding frequency so that T       filogXi + f2logX2 + fn logXn Log G =         N(=Sf) Sf logX N (13) Geometric mean is especially useful in the construction of index number it is an average most suitable when large weights have to be given to small values of observations and small weights to large, value of observations. Here, as in the case of the arithmetic mean, we first find out the mid values of the classes look up the logarithms of these mid values, multiply each logarithmic mid value fry the corresponding frequency, sum up these and divide by the number of items and then find but anti-logarithm of this quotient. Thus log of geometric mean is equal to the arithmetic mean of the log, values of the variate in the unit disrate (frequency distribution or weighted mean incase of grouped frequency distribution. Application of Geometric Mean: Average Rates of change and the Compound Interest Formula. Let us assume that the output of an industry increased 20 percent is die first year, 40 per in the second year, 30 percent in the third year and 10 percent in the fourth” year. Then the index of industrial output will be: | At the end of | zero year | 100 | |---|---|---| | | 1st year | 120   (20% increase) | | | 2nd year | :      168   (40% increase) | | | 3rd year | :      218.4 (30% increase) | | | 4th year | 240. 2 (10% increase) | What is the average rate of increase during, these four years ? We can clearly see that the output in 1st year was 1.20, times that of the zero year, that in the 2nd year 1-40 times of that of the 1st year, and so on. The geometric mean (G), would show the average rate of increase. Thus G = 4^ (1.20 x 1.40 x 1.30 x 1.10) or log G = 41 (log 1.20 + log 1.40 + log 1.30 + log 1.10) = 41 (0.0792 + 0.1461 + 0.1139 + 0.0414 1 = 4 (6.3806) = 0.09515. Hence G = Anti log (0.09515) = 245 This implies feat the average rate of change of output over 4 years was 24.5 percent (1.245 x 100—100). If r is average rate of change, then Pn = P0 (1+r) n                                      ...... (14.1) where P0 is initial and Pn, the figure after n years of change. This is the familiar compound interest formula. This may also be written as. r = 1/n p \ P 7 1 (14.2) Example 5. India’s net national product was Rs. 13,294 crores in 1960-61; it rose to Rs. 18,755 crores in 1970-71, at 1960-61 prices. What is the average rate of change over the decade? Using formula (14.1) and (14.2), we have 18,755 = 13,294 (1 + r)10 or r = 18755 1/10 13294 1 ( 18755 V/10 Le 1 + r = IÏ32947 Applying logarithms, we get 1 log (l + r) = 10 (log 18755—log 13294) 1 = 10 (4,2731- 4.1237) 1 = 10 (0.1494) = 0.01494 Hence, l + r = Anti log (0.01494) = 1.035 r = 0.135 — 1 = 0.035 or 3.5 percent Thus India’s N.N.R on the average, increased at a compound rate of 3.5 percent per annum over, the decade 1960-61 to 1970-71. Another use of the geometric mead, through compound interest formula is in discounting and capitalization. Example 6: Suppose, we have to choose between producing’ electricity through a hydle project or a thermal project. It is estimated that the cost of construction of the hydel project would, be Rs.100 crores and its life would be 80 years and that of the thermal project. Rs. 40 crores and its life would be 25 years. Both Projects are expected to yield an annual net income of Rs. 10 crores. Which of the projects should be chosen? For comparing the relative net returns versus costs of the two Projects, we shall have to compute the present value of returns of both projects and compare these to the costs. Present value can be found out by discounting the future returns at, the rate of ‘discount, r (let us say, 5 percent (and by using the compound interest formula: Po = ^ Pn (1 + r)n (14.3) 10      10       10 +++ 1 ohydro (1 + r) (1 + r)2    (1 + r)3 10 (1 + r)80 crores 11  1 --+ t----“r+ t----rr 1 + r (1 + r)2 (1 + r)3 1 (1 + r)80 crores Now applying the formula for the sum of a geometric series we have, the sum of the present value of net returns over 80 years* a (1 - r )n *Sn = —1-----where Sn is the sum of a geometric series for n terms, a is the first terms of the series, and r is the common ratio [1/1+r) in our present case]. = 10 P 1 ----x 1 + r 80 1 -I — I 1- V1 + r 1 1 + r = 10 P 1 ----x 1 + r ( 1 A 80 1 V1+r J 1 + r -1 1 + r = 10 P i \80' 1 -V rJ r Let Hence log P 0hydro 10 K r 80 1 -I — I V1 + r J 10 0.00 1 -I----1---- V1 + 0.05 80 hl       1 1 = 200 p - 7T30 [ = 200 - 200 (1.05) -80 . x = (1.05)-80 x = - 80 log 1.05 = -80(0.002166) = - 0.173280 = 1.82672 x = Anti-log 1.82672 = 0.671 = 200 - 200 (1.05)-80, therefore, becomes. = 203 - 200 (0.671) - 200 - 134.2 = Rs. 65.8 crores. Applying the above formula, to the thermal projects we get P 0thermal Now let x then log x = 200 [1-(1.05) -25] = (1.05) -25 = -25 log 1.05 = - 25 (0.002166) = - 0.054150 = 1.94585 . Therefore, x Hence P 0therma = Anti-log’ (1.94585) = 0.88277, = 200 (1-0.88277) = 200 (0.11723) = Rs. 23,446 crores. Now Returns/Cost ratios for the twoprojects, will be: (i) Returns/Cost ratio for Hydro Project = 65.8/100 = 0.658. (ii) Returns/Cost ratio for me Thermal Project = 23.44/40 = 0.586 app. Clearly, at a rate of discount of 5 percent, the hydro project would yield a higher net return/ cost ratio, and maybe preferable. The use of geometric mean; for averaging, index number will be studied, when we take up the study of Indian price index numbers. The geometric mean is preferred for this purpose, because its assigns equal weight to equal ratio of change. The geometric mean is, very useful in averageing ratios and percentages. It also helps in determining the rates of increase and disease: It is also capable of further, algebraic treatment so that a combined geometric; mean can easily be computed. Geometric mean cannot be computed if any observation has either a value zero or negative. Self-Check Exercise-8 Q1.   Define Geometric mean. Discuss the application of geometrical mean. Q2.   Calculate the geometrical mean of the following series: 180,    190, 240, 386, 492, 662 1.11 The Harmonic Mean The Harmonic mean H is the reciprocal of the arithmetic mean of the reciprocal of the values. The harmonic mean is a measure of central tendency for the data expressed, as rates such as kilometers per hour, tones per day, kilo meters per liter etc. Thus H = 1 11    1 --- + --- + .... + --- X1   X2 X n N 1 S — X N H = N S — X (15.1) For purpose of computation, (15.1) may be put as — = S — HX (15.2) N Having found the value of , we can get H; as its reciprocal. H Example 7. Assume that me 3 grades of, Golden delicious apples are being quoted in Shimla Market as follows: Large : 4 per rupee; Medium : 6 per rupee and small : 11 per rupee. - _ 4 + 6 +11 _ „ The arithmetic mean X - ----3----- 7 per rupee i.e 14.3 paise per apple. This is the price we must pay if we spend equal amounts of money for each grade. Paying 14.3 p. for each of the 21 apples we shall spend Rs. 3/- for the lot. The harmonic mean gives different results. 1N T H =  1 s—    s— XX N = 5.91 app. for Re 1/- or 1692 paise per apples This is the price we shall pay if equal number of apples are bought at each price. The harmonic mean for frequency distributions is hardly ever used. In this case. H N = S f S f X (15.3) or 1 H f X N (15.4) The harmonic mean hardly adds much to the information furnished by the arithmetic mean. It may, however, be useful when, data are generally quoted in terms of problems solved per minute, miles or kilometers covered per hour, units sold per rupee etc. If, in case of a highly skewed distribution, the plotting of the reciprocals of the data (or class mid values), results in an approximately normal curve, harmonic mean could be useful. But these instances are rather unusual. The harmonic mean is useful for computing the average rate of increase of profits, or average speed at which a journey has been performed, or the average price at which an article has been sold. Another measure of central tendency, which is more, of theoretical rather than practical interest is the root mean square or the quadratic mean. Root mean Square (R.M.S - 3 < SX2 ^ I N J (16) This type of average is used sometimes in physical applications. Self-Check Exercise-9 Q1. Define Harmonic mean. Q2. Calculate the Harmonic mean of the following: 2,    4, 7 12, 19 1.12 Summary 1.         Mean, medium, mode, geometric mean and harmonic mesa are parameters which attempt to characterize the central tendency of the distribution. 2.    (a) The arithmetic mean is the sum of all observations in the series divided by the number of observations. (b) Formula for ungrouped data Y- £X X — -- N (c) Formula for grouped data 3. (a) (b) (i) (ii) (iii) V SfX Long method: X —    . - -  SfX Short-cut method : X — Xa + ——. N - _ -   sfD' . Short-cut method (in terms of highest common factor) : X — Xa +—— x i . The median is a positional measure of central tendency dividing the series when arrayed, into two equal halves. i I N +1 For grouped data, Median = li + “ I “ f \ 2 Cfi I. 4. (c) (a) (b) Quantiles divide a series when arrayed or a frequency distribution into n equal parts. The mode is the most frequent or typical value or a series. It cannot be generally easily located for ungrouped data. (fm ~ fl)      + i M0 l + (fm - fl) + (fm - fl) (a)    For a symmetrical distribution, Mean = Median = Mode. (b)    (i) For positively skewed distributions Mean > Median > Mode. (ii)    When the distribution is slightly skewed Mean - Mode = 3 (Mean - Median). (c)    The mean possesses a number of advantages over the median and mode, but is less suitable when data are irregular or broken, highly skewed, with unequal or open end class intervals. (a) The geometric mean is the nth roof of the product of the value of n items. It is useful to measure the average rate of change, for discounting and capitalization, and for constructing price indices. (b) (S logX G = Anti - log I 6 I N (Sf logX A or Anti - log I n I for grouped data. 7. The harmonic mean is the reciprocal of the arithmetic means of the reciprocal, of the values H = 1 111 11 X1  X2 X3 N 1.13 Glossary •    Arithmetic Mean: The sum of all data points divided by the number of data points; commonly known as the average. •    Weighted Arithmetic Mean: An average where each data point is multiplied by a weight reflecting its importance before calculating the mean. •    Median: The middle value of a data set when arranged in order; if even, it is the average of the two middle values. •    Mode: The value that appears most frequently in a data set; useful for identifying the most common category. •    Geometric Mean: The nth root of the product of all data points, used for data sets with exponential distribution or growth rates. •    Harmonic Mean: The reciprocal of the arithmetic mean of the reciprocals of the data points, useful for rates or ratios. •    Central Tendency: A measure that identifies the center of a data set, providing a representative value. •    Skewed Data: A data set where values are not symmetrically distributed, often requiring the median for a more accurate measure. •    Outliers: Data points that are significantly different from the rest of the data; can affect the arithmetic mean. •    Bimodal/Multimodal: A data set with two or more modes, indicating multiple frequently occurring values. 1.14    Answers to Self-Check Exercise Self-Chek Exercise-1 Answer to Q1. Refer to Section 1.3. Answer to Q2. Refer to Section 1.3. Self-Chek Exercise-2 Answer to Q1. Rs. 34. Answer to Q2. Rs.43. Self-Chek Exercise-3 Answer to Q1. Refer to Section 1.5. Answer to Q2. 60.63. Self-Chek Exercise-4 Answer to Q1. M=21. Answer to Q2. M=52.69. Self-Chek Exercise-5 Answer to Q1. Refer to Section 1.8. Answer to Q2. Q1=26, Q3=60, D5=45, P25=26. Self-Chek Exercise-6 Answer to Q1. Refer to Section 1.8. Answer to Q2: 163.576 Self-Chek Exercise-7 Answer to Q1. Refer to Section 1.9. Answer to Q1. Refer to Section 1.9. Self-Chek Exercise-8 Answer to Q1. Refer to Section 1.10. Answer to Q2: 317.9 Self-Chek Exercise-9 Answer to Q1. Refer to Section 1.11. Answer to Q2: 4.86 1.15    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Kmenta, J. (1971). Elements of econometrics. Macmillan. 3.    Jain, T. R., &Aggrawal, S. C. (2020). Business statistics. V K Publications Pvt. Ltd. 4.    Yamane, T. (1973). Statistics: An introductory analysis. Harper & Row. 1.16    Terminal Questions Q1. Distinguish between mean, median and mode. Q2. The grades of 36 students in an Auditing test are given in the following table: | Grades (Less than) | 40 | 50 | 60 | 70 | 80 | 90 | 100 | |---|---|---|---|---|---|---|---| | No. of Students | 3 | 7 | 13 | 23 | 29 | 33 | 36 | Find Mean, Median and Mode. ***** UNIT-2MEASURES OF DISPERSION AND SKEWNESS STRUCTURE 2.1    Introduction 2.2    Learning Objectives 2.3    Dispersion 2.3.1    Measures of Absolute Dispersion 2.3.1.1    The Range 2.3.1.2    The 10-90 Percentile Range 2.3.1.3    The Quartile Deviation or the Semi-inter Quartile Range 2.3.1.4    The Mean of Average Deviation 2.3.1.5    The Standard Deviation 2.3.2    Measures of Relative Dispersion 2.3.2.1    Co-efficient of Variation Self-Check Exercise-1 2.4    Skewness Self-Check Exercise-2 2.5    Kurtosis Self-Check Exercise-3 2.6    Summary 2.7    Glossary 2.8    Answers to Self-Check Exercise 2.9    References/Suggested Readings 2.10    Terminal Questions 2.1    Introduction Dear Student, In the previous lesson, we were concerned with various measures that are used to provide a single, representative value of a given set of data. This single value alone cannot adequately describe a set of data. Therefore, in the present lesson we shall study two more important characteristics of a distribution. First we shall discuss the concept of dispersion-variation and later the concept of skewness. 2.2    Learning Objectives After going through this unit, you will be able to •    Define and explain key concepts such as dispersion, absolute and relative measures of dispersion, skewness, and kurtosis. •    calculate and interpret various measures of dispersion. •    Assess the skewness and kurtosis of data distributions, understanding the implications of these characteristics on the shape and spread of data. 2.3 . Dispersion Two groups of students Show’ the same mean height, 64 inches. However, the height of the first group ranges from 58 " to 74 ", whereas that of the other, lies between 60" and 69". It is obvious, ‘that the heights of the first group are more scattered; whereas the second group shows a greater uniformity. In fig 2.1 (a), we have shown frequency curve A, with a greater spread or dispersion than frequency curve B, though both have the same mean. Fig. 2.1 (b) shows, on the other hand, two frequency curves having different means, but having, the same dispersion. Fig. 2.1 (a) Two Frequency Curves A & B, with different Means but same. Fig. 2.1 (b) Two Frequency Curves A & B with Same mean, but Different Dispersions Dispersion When we know the dispersion of a variate besides its central tendency, we may speak with greater confidence, about the dependability of the mean. You have all, I believe heard the story of a man who on an enquiry found that the mean or average depth of the river was 5 feet. Since he was 5'6" tall,” he decided on the basis of this information that he could safely cross across the river. Unluckily he was drowned, because he did not know swimming and he met, on the way, depths over 12 feet had he know about the dispersion of the river depth, he would surely have thought Better, not to risk wading through the river. Assume that two manufactures of fluorescent tubes claim an average life 2500 hrs. for their products, Further enquiry showed, that tubes; of company B lasted from 500 to 5,000 hours, whereas, those of company T, lasted a minimum of 1500 and a maximum of 3500 hours, it is obvious that products of company T have greater uniformity than those of B. Dispersion, therefore refers to the variability in the size of items. It indicates that the size of items in a series is not uniform i.e., the value of various items differs from each” other. If the variation is substantial, dispersion is said to be considerable and if the variaton is little, dispersion is insignificant. The term dispersion not only gives a general impression about the variability of a series, but also a precise measure of this variation. Generally in a precise study of dispersion, the deviations of size of items from a measure of central tendency are found out and then these deviations are averaged to give single figure representing the dispersion of the series. This figure, can be compared with, similar figure representing other series, and this it is possible to make a comparison between the averages and the dispersion of two or more series. Significance of Measuring Variation Measuring Variation is significant for some of the following purposes. (i)    Measuring variability, determines the reliability of an average by pointing out as to how for an average is representative of the entire data. (ii)    Another purpose of measuring variability is to determine the nature and cause of variatior in order to control the variation itself. (iii)    Measures of variation enable comparisons of two or more distributions with regard to their variability. (iv)    Measuring variability is of great importance to advanced statistical analysis. Sampling and statistical inference is essentially a problem in measuring variability. Absolute and Relative dispersion : If we calculate dispersion of a series relating to the income of a group of person in absolute figures, it will be expressed in the unit in which the original data or say Rupees. Thus when we say that in income of a group of persons is Rs. 120 p.m. and the dispersion is Rs. 30. This is called” Absolute Dispersion. When dispersion is, measured as a percentage or ratio of the average it is called Relative Dispersion. It is not expressed in the unit of original data, in the above example the average income would be Rs. 30 p.m. and the Relative dispersion 30/120 or 25.00 percent. 2.3.1    Measures of Absolute Dispersion We have several measures of dispersion available for use for different purposes. The most important of these are: (i)    The range, (ii) The 10-90 Percentile Range, (iii) The Quartile Deviation, (iv) The Mean Deviation, and (v) The Standard deviation. 2.3.1.1    The Range This is the simplest but a crude measure of dispersion. It is the difference between the values at the extreme items of a series. Suppose we are told that the average grade of two groups of students, A and B is 65 points. However, the highest and lowest grades of group A are 90 and 25 points respectively whereas those of group B, range from. 60 to 70 points i.e. only 10 points. Now the range of Group A 90 - 25 = 65 points and that of Group B = 70 - 60 = 10 points. The range, as may be see a, tells us only about the two extreme, values and we do not know anything more about the rest of the data whether they are concentrated around the mean or scattered widely. It is unfit for purpose of comparison, if the distributions are in different units. If range is divided by the sum of the extreme items the resulting figure is called the “ratio” of the range” or the “coefficient of scatter”. 2.3.1.2.    The 10-90 Percentile Range This measure excludes the lower 10 percent and the upper. It percent of the items and concentrates attention on the middle 80 per cent. Thus it steers clear of the, extreme values. It, however, does not use the value of all the items. Its only concern is the values of P90 and P10, and is not at all affected by the arrangement of values within and outside this range. It is obvious that the 1090 percent tile range. 2.3.1.3.    The Quartile Deviation or the Semi-Inter Quartile Range This measure of dispersion, as should be obvious from its name, is based on the values of the lower and upper quartiles. It is given by Q = Q Q (1) 2 In a symmetrical series, the lower and upper quartiles lie at equal distances from the median, so that the median ±Q “should en-compass 50 percent of the items. The quartile deviation like the 10-90. Percentile range is not affected by extreme values. However, this also fails to consider the values of all the items. 2.3.1.4.    The mean of Average Deviate A measure of dispersion that includes the variability of all the Stems is the mean deviation. It is the mean or average of absolute deviations from the mean of median (i.e. ignoring sign). Mean Deviation or - S |dx | o x =----- N (2.1) about mean Where fix is the mean deviation from the mean, dx are the deviations of the items forms x , it is the symbol for ignoring sign. Since we are interested in the amount of variability, i.e. in the distance of the deviations, the minus “signs” are diregarded when finding the mean variability. Mean Deviation or S | dx | OMed=—--- N (2.2) about Median where | dx |    =     deviations of the values from the median (ignoring the signs). Because the sum of the absolute value deviations is a minimum when taken round the median, sometimes, the mean deviation is computed in relation to the median. In practice, however, it is the mean which is generally used, and if the series is symmetrical, the resulting mean deviation will be the same; whether computed in relation to the mean or the median. Mean deviation for a frequency distribution is ox =Sfldxl (2.3) where fix    stands for the mean, deviation in relation to the mean. | dx | stands for deviations of the mid-points of class intervals, or the class average from the mean (signs ignored). f stands, for the frequency. Similarly, the mean, deviation in relation to the median for grouped data, will be , Lf |dx | S Med. =—--- (2.4) N where | dx | stands for absolute valise of deviation around the median. In case of normal distribution, 57.5 percent of the items lie within the range Y = Y . Where the distribution is moderately skewed, this will be approximately true. Let us compute the quartile deviation and the mean deviation for the monthly per capita consumer expenditure for rural India, as given in table below. Table 2.1 Consumer Expenditure in Rupees per person for 30 days by Monthly per capital. Expenditure Classes: Rural India | Midvalue X | Monthly Per Capita Expenditure Classes in (Rs.) | Percentage of persons in the Expenditure Class (f) | Average Consumer Expenditure (Rs.) | Cumulative Percentage of person under upper limit of group | \| dx \| From Median | f \| dx \| | |---|---|---|---|---|---|---| | 1 | 2 | 3 | 4 | 5 | 6 | 7 | | 4 | 8.00 | 7.28 | 6/62 | 7.28 | 1243 | 90.49 | | 9.5 | 8.11 | 13.79 | 9.72 | 21.07 | 6.93 | 95.56 | | 12 | 11.13 | 10.95 | 12.02 | 32.07 | 4.43 | 48.51 | | 14 | 13.15 | 11.12 | 13.94 | 43.14 | 2.43 | 27.02 | | 16.5 | 15.18 | 14.43 | 1650 | 57.57 | 0.07 | 1.01 | | 19.5 | 18.21 | 10.69 | 19.72 | 68.26 | 3.07 | 32.82 | | 22.5 | 21.24 | 73 | 22.51 | 76.09 | 6.07 | 47.55 | | 26 | 24.28 | 6.72 | 25.98 | 82.81 | 9.57 | 64.52 | | 31 | 28.34 | 8.09 | 30.09 | 90.90 | 14.57 | 117.31 | | 385 | 34.43 | 4.68 | 38.37 | 95.58 | 22.07 | 103.29 | | 49 | 43.55 | 2.32 | 50.74 | 97.30 | 32.57 | 75.56 | | *88-13 | 55 & above | 2.10 | 88.13 | 100.00 | 71.70 | 150.57 | | | All levers | 100.00 | 20.03 | | | 854.54 | * Source : Data for the columns 2—4 from NSS, 15th Round, July 59—June 60, No. 98 New Delhi 1965. Mid-value for this class is assumed to be equal to the class average. Q3 = 1 + -3f 3N 4 C 3N Now the upper quartile Q3 = value of 4 th items. i.e. = value of person 4 x 100 = 75% The percentage lies in the group Rs. 21-24. Hence l1, = 21, i = 3, f = 7.83, and C = 68.26. 33 a Q3 = 21 + — (75 - 68-26) 21 + — (6.74) .. = Rs. 23.58 N Similarly, Q1, the lover quartile = value of 4 th 100 i .e. 4 = 25th item. This lies in the group. Rs. 11-13. So l1= 11, i = 2, f = 10-95, and c - 21.07. 2 Hence Q1 = 11 +       (25-21.07) . 2 x 3.93 = 11 + 10.95 = Rs 11.72. Now quartile deviation, QD = Q - ^3 2 ^1 = 24.44. - 11.72 12.72 2    =Rs. 6.36. The median per capita monthly expenditure is the value of the 2 = 2 = 50th percentage person. This lies in the group Rs. 15-18. Here, L = 15, i = 3, f = 14.43, c = 43.14. „ „    , . i (N ? Median = L + “I ~- c I f \ 2    / 3 = 15 + 14.43 (50-43.14) = 85 + 1.43 = Rs. 16.43. Now mean deviation in relation to median 3 mad - 054.54 100 Rs. 8.55 app. Similarly, for this series, the mean is = 20.03 = x . For computing the mean deviation in relation to the mean we get the deviation etc. as follows: Work sheet table 2.1.1 for computing the Mean Deviation | Class midpoint | \| dx \| from 20.03 | Class Average | f | f \| dx \| | |---|---|---|---|---| | x(l) | (2) | x(3) | (4) | (5) | | 4 | 16.03 | 6.62 | 7.28 | 116.70 | | 95 | 10.53 | 9.72 | 13.79 | 145.21 | | 12 | 8.03 | 12.07 | 10.95 | 87.93 | | 14 | 6.03 | 13.94 | 11.12 | 67.05 | | 16.5 | 3.53 | 16.50 | 14.43 | 50.9’4 | | 19.5 | 0.53 | 19.72 | 10.69 | 5.67 | | 22.5 | 2.47 | 22.51 | 7.83 | 1862 | | 26 | 5.97 | 25.98 | 6.72 | 40.12 | | 31 | 10,97 | 30.09 | 8.09 | 88.75 | | 38.5 | 18.47 | 38.37 | 4.68 | 86.44 | | 49 | 28.97 | 50.47 | 2,32 | 67.21 | | *88.13 | 63.10 | 88.13 | 2.10 | 143.08 | *Mid-point of the last class assumed Sf | dx | = 917.65 to be equal to the class average. Sx ; using mid-points of classes = Sf | dx | 917.65 ——— - , ARs. 9.18 we can clearly see that the N 100 mean deviation from the Median (Rs. 8.55) is less than that from the mean (Rs. 9.18). The mean deviation for grouped data is very rarely used. We therefore, take up the discussion of the most important measure dispersion viz. The standard deviation. 2.3.1.5 The Standard Deviation The standard deviation is regarded as the most important measure of dispersion because this has desirable mathematical properties. We overcome the people of the negative sign are deviations, not by ignoring it, but by squaring the deviation. Standard Deviation is the square-root of the arithmetic average of the squares of the deviation measured from the mean. For ungrouped data we compute the measures as follows : CT 2 sdx2 N S(x1 - x )2 N (3.1) where o2 stands for the variance (which is the square of the standard deviation and dx stands for (x1- x ). Hence standard Deviation c (read as ‘sigma’) - ^< S(x1 - x )2 N >-^ Sdx2 N (3.2) Table 2.1.1 Computation of standard deviation for Annual rainfall in Himachal Pradesh | Year | Annual Rainfall (cms) x | dx from x = 138.5 | dx2 | X2 | D from xa = 140 | D2 | |---|---|---|---|---|---|---| | (1) | (2) | (3) | (4) | (5) | (6) | (7) | | 1950 | 146 | + 7.5 | 56.25 | 21316 | + 6 | 36 | | 1955 | 161 | + 22.5 | 506.25 | 25921 | - 21 | 441 | | 1960 | 120 | - 18.5 | 342.25 | 14400 | - 20 | 400 | | 1961 | 161 | + 22.5 | 506.25 | 25921 | + 21 | 441 | | 1965 | 109 | - 29.5 | 870.25 | 11881 | - 31 | 961 | | 1966 | 149 | + 10.5 | 110.25 | 22201 | + 9 | 81 | | 1967 | 157 | + 81.5 | 3421.25 | 24469 | + 117 | 289 | | 1168 | 105 | +33.5 | 1122.25 | 11025 | - 35 | 1225 | | Total | 1108 | + 33.5 | 3856.00 | 15731 | +74 - 86 = 12 | 3874 | Source: Statistical outline of Himachal Pradesh 1970, pp. 92-93 approximation to, cms, done by the author. _ Xx  1108 x-~ -     = 138-5 cms. N8 ^ 2' "N')=y385600]P (482-00) We pointed out in the last lesson that Ed2 is a minimum when taken round the arithmetic mean. Therefore, the standard deviation is always computed in relation to the mean. The steps in the above computation are : (i)    Determine the deviation dx of each item from the X i.e. find x1 -x ,x , (ii)    Square these deviations i.e. compute each (x1- x )2, (iii)    Sum these squared deviations i.e. compute E(x1- x )2, (iv)    Divide, this sum by N. This gives us the value of the variance i.e. c2. (v)    Take the square root of the variance. The above procedure involves computation of deviation (x) for every item, and this would be quite lengthy and cumbersome, if the number of items, N is quite large. The value of standard deviation can also be found out, without doing this procedure, by means of the formula : | 2      5x 2 <5 a = \--- 1 N U _ J[157314 = 1    8 | y NJ/                    (3-3) - (138.5)4 | |---|---| = (19664.25 - 19182.25) = (482.00 = 21.95 cms. It is obvious from the calculation of x2 in col. 5 of Table 2.2.1 that this method is more suitable to machine calculation and simpler. By combining formulas (3.2) and (3.3), we get 2 2( Xi — x) N Ex2 -{(Ex)2/N} (3.4) N or a2 Ex2 — {Ex}/ N N = A 157314 — (11082/8) 8 157314 — (12227664/8) 8 = A 157314-153458 8 We have a short-cut method also available, which allows us to compute the value of the standard deviation, by taking the deviations from an assumed; mean, rather than, from the true mean and make the necessary correction. This formula is: S.D. = a = A a = A ΣD2 — N (3.4.1) LD2-{(LD)2/N} N (3.5.2) where D stands for deviation of the items from an assumed, mean i.e. D = xf - x0 . Let us choose 140 cms as our assumed mean and calculate the standard deviation, as shown in cols. (6) and ,(7) of Table 2.2.1. Now according to formula 3.4.1., we have a = A ΣD2 N = (484.25-2.25) A 3378 8 = 3(484.25-2.25) = (482.00) = 21.95 cms. Similarly, with formula 3.4.2, we have °= A id :(id) \: N = A 3874 - {(-12)2/8)} N = 43874 —181=42856 U (482) I 8I 8 = 21.95 cms. The standard deviation for frequency Distribution For a frequency distribution, where the values of the individual items are not known, such as in Table 2.1, a formula that gives the valve for the, standard deviation of the distribution is as follows : °=^[ ff \ N J = A ' f^ J (3.5) where dx = x1 - x , stands for the deviation of the mid values (x1’s) for mean. It may be noted that x1’s here stand tor class mid values as distinguished individual values for ungrouped data. Worksheet Table 2.3.1 for computing the Standard Deviation for Consumer’s Expenditure Data of Table 2.1 | Monthly per capita expenditure classes | Class midvalues (x) | Percentage of persons in expenditure true class (f) | Deviations for the mean = 19.74 dx2 (dx) | fdx2 | |---|---|---|---|---| | (1) | (2) | (3) | (4) | (5) | (6) | | 0-8 | 4 | 7.28 | -15.74 | 248.27 | 1847.41 | | 8-11 | 9.5 | 13.79 | -10.24 | 104.86 | 1146.02 | | 11-13 | 12 | 10.95 | -7.75 | 59.91 | 656.01 | | 13-15 | 14 | 11.12 | -5.74 | 32.95 | 366.40 | | 15-18 | 16.5 | 14.43 | -3.24 | 10.50 | 151.52 | | 18-21 | 19.5 | 10.69 | -0.24 | 0.06 | 0.64 | | 21-24 | 22.5 | 7.83 | +2.76 | 7.62 | 59.66 | | 24-28 | 26 | 6.72 | +6.26 | 39.12 | 263.36 | | 28-34 | 31 | 8.09 | +11.26 | 126.79 | 1025.73 | | 34-43 | 38.5 | 4.68 | +18.76 | 351.94 | 1647.08 | | 43-55 | 49 | 2.32 | +29.26 | 856.15 | 1986.27 | | 55 & above | 88.13 | 2.10 | +68.39 | 4677.19 | 9822.10 | | All levels | | 100.00 | | | 9232.20 | * Mean compound through the use of class in d - point = 19.74. * Class mid-value assured to be equal to class average, a=^ 'f^ / 19232.28 A = ^ (192.32.20) I 100 J = Rs. 13.87 Computation of the standard deviation, through this long method is extremely laborious and cumbersome. We may therefore use deviations from an assumed mean and make the necessary correction for this by the use of the formula. a = ^ f2 N or A S/D2 -{S/D)2}/N N (3.6) where D = x1 = xa Let us compute the standard deviation by the short-cut method 3.6 for the foregoing frequency distribution. Worksheet Table 2.3.2 for combating the standard Deviation for Consumer’s Expenditure Data of table 2.1. (by short-cut method) | Class midvalues (x) | Frequency (f) | Deviations from x = 20 (D) = x - xa | f.D (2) x (3) | f.D2 (5) -(3) x (4) | |---|---|---|---|---| | (1) | (2) | (3) | (4) | (5) | | 4 | 7.28 | - 16 | -116.48 | 1863.68 | | 9.5 | 13.79 | - 10 | -144.80 | 1520.40 | | 12 | 10.95 | - 8 | -87.60 | 700.80 | | 14 | 11.12 | - 6 | -66.72 | 400.32 | | 16.5 | 14.45 | - 3.5 | -50.05 | 176.75 | | 19.5 | 10.69 | - 0.5 | -5.34 | 2.67 | | 22.5 | 7.83 | + 2.5 | +40.38 | 248.93 | | 26 | 6.72 | + 6 | +40.32 | 241.92 | | 31 | 8.09 | + 11 | +88.99 | 978.89 | | 38.5 | 4.68 | + 18.5 | +86:58 | 1601.73 | | 49 | 2.32 | + 29 | +67.28 | 1951.12 | | 88.13 | 2.10 | + 68.13 | +143.07 | 9747.36 | | | 100.00 | | - 471.44 + 445.82 | 19234.59 | | | | | - 25.62 | | G =A 2 - (W ] __N— N 19234.59 -I 25'62 V 100 100 v                       J = A 19234.59 -T 65628 Ì V 100 J 100 19234.59 - 6.56 100 A= | 19228.03 | = A (192.28) = Rs.13.87 V 100 J                             . A further implication of computation can be achieved by taking the deviation in tens of the com-mon class interval or common factor. The formula then, becomes a = i x 3 f2 - (W? N N where D’ = D . i i = common class-cionterval of common factor. Example: Table 2.4.1 for computing Standard Deviation | Percentage of mark | Midvalue (x) | Number of Student (f) | D’ =(x= Xa )/i X = 05x i = 20 a | fD’ | fD‘2 | |---|---|---|---|---|---| | 0-20 | 10 | 7 | -2 | -14 | 28 | | 20-40 | 30 | 15 | -1 | -15 | 15 | | 40-60 | 32 | 0 | 0 | 0 | | | 60-80 | 70 | 12 | 1 | 12 | 12 | | 80-100 | 90 | 9 | 2 | 18 | 36 | | | | | | -29 | | | | | | | +30 | 91 | | | | | | =1 | | | | [ 91 - (1)21 | |---|---| | a = 20 x A | ----100 100 L J | = 20 x 90.9999 J' 100 J = 19.08 marks. Formula 3.5 can also be ‘written in the form G = 3 f2 ( f \ ^^^^^^^^^^^^^^^^^^^™    ^^^^^^^^ I ^^^^^^^^^^^^^^^^^^^^^^^™ I N k N ) (3.8) Properties of the standard Beyiatioa. The standard deviation and its square viz. Variance are, by fax, the most important and roost frequently used absolute measures of dispersion. Its use in sampling for determining the areas under the normal, curve and for various types of skewed distribution is extremely common. It is also used in testing the reliability of certain statistical measures, in correlation and in business cycle analysis. 1.    The standard deviation may be defined as : a = A 1 k N ) where a is an average besides the arithmetic mean. Of all such c’s, the minimum is that for which a = x. As. stated in the last lesson, the sum, of the squares of the deviations from the mean x , is the minimum. Hence, the definition of the standard deviation is a = A l^(^^ ) For the consumption expenditure distribution of Table 2.1 X± o =19.744 ± 1370 is Rs. 6.04 and 33.44. To find out the percentage of persons, lying between these two limits, we first determine the percentage included in the range two limits, we fest determine the percentage included in the range Rs. 6.04 and Rs. 8 (the upper limit of the first group) and in the range, Rs. 28 and Rs. 33.44. The percentage of persons in the classes, 8-11, 11-13, 13-15,15-18,18-21, 21-24, and 24-28, all lie within the limit of X ± o or: Assuming regular distribution of frequencies in the classes, 0 to 8, and 28 to 34, the estimated, frequency lying between, the specified limits are: 8 - 6.08 Rs. 6.04 and 8 =------- x 7.78 = 1-75% app. 8 33.4 - 28 and Rs. 28 and Rs. 33.4 =             x 8.09 = 7.28% app. 6 — (34 — 28) Hence, the percentage between Rs. 6.04 and 33.44 = 1.75 + 13.79 + 10.95 + 11.12 + 14.43 + 10.69 + 7.81 + 6.72 + 7.28 + 84.56%. Similarly, the percentage of persons with per capita monthly consumption expenditure, lying within the range X + 2o i.e... Rs. 19.73 + 2 (13.70) - Rs 47.13 an 0*, may no calculation be found to be 96.28% and those within X ± 3o, 99.82%. It is obvious from these observed percentages that our distribution diverges sharply from a normal curve and is highly skewed. When we study the normal curve, we shall be interested in determining the limits within which 90%, or 95% or 99% of the observation lie, or rather in the proportion lying beyond certain specified limits, such as 1 percent or 5 percent etc. We can ‘always translate the difference between the mead and an individual value into units of standard deviations. We say the deviation X1—X have been standardized of normalized. In general, when the variable X is divided by its standard deviations, X: X1: X2 ........ Xn a a a a We say the variable X has been standardized. These deviations in units of standard deviations X1 - X i .e. are called Z’s (more about these later). a The standard deviation of the distribution of a standardized variable: X: X1: X2 ........ Xn aaa  a is unity, in theoretical statistics, distribution with mean = 0 and standard, deviation (or variance.) = 1 (called unit distribution are often used to facilitate analysis). 3.    If we have two set, consisting of N1 and N2 numbers (or two frequency distribution with total frequencies’ N1 and N1) have variances given by o12 and a22 respectively and the same mean = X then the combined variance (o2) of both sets (or both frequency distribution) is given by 2 Nσ+Nσ σ=   11   22 (i)           N= (N1+N2)                               ..... (3.9.1) N1                     N2 or N σ2 = = ∑(X1i - X1) + ∑(X2i - X2)            ..... (3.9.2) ii (ii)    W here, however, N1 , X 1 and σl and N2, X2 and σ2 are given of two Sets of variates, the variance for the composite set is given by : Nσ2 = N1 (σ12 + d12 ) + N2 (σ22 + d22 )                ..... (3.9.3) where d1 = X 1 - X and d2= X2- X, and where X is’ combined mean for the two groups : The combined variance σ2 can also be computed by the following methods. (iii)    Nσ2 = N1 (σ12 + X12 ) + N2 (σ22 + X22 ) - NX2)        .... (3.9.4) N1N2            2 (iv)    Rσ2 = N1σ12 + N2σ2 + (N1+N2) (X1-X2) .... (3.9.5) These results can be generalized to 3 or more sets. Sheppard’s correction for variance Sheppard’s correction is intended to make adjustments for error due to grouping, of data into classes (grouping error) Corrected σ2 = Variance from grouped data C2/12.          ...... (3.10) where C is the class interval size. The correction C2/12 is used, for distributions of continuous variables where the ‘’tails” taper gradually to zero in both directions. There is lack of agreement among statisticians whenand where Sheppard’s corrections should be applied, because of the belief that they often to over correct and thus replace old errors, by new errors. Hence, great caution is needed in their use. Empirical Relation between Measures of Dispersion For moderately skewed distribution, the relationship between mean deviation, quartile deviation and standard, deviation is as under. 4 δx =  σ 5 4 and Q.D = 5 σ These relationships result from the fact that for the normal distribution, the mean deviation and the quartile deviation are equal to 0.797 σ and 0.6745σ, respectively. 2.3.2 Measures of Relative Dispersion We have so far been discussing measures of absolute dispersion, all of which are expressed in terms of the units rupees centimeters etc. When we want to compare the dispersion of spread of two or more series, it may not generally be appropriate to use such a measure. We may be faced with three possible types of situations while making such comparisons: (a)    The series sought to be compared may be expressed in terms of the same units, and the means may be equal or nearly, equal. Hence absolute dispersion measures can servers fairly adequately for comparing the series. (b)    The series to be compared may be expressed in terms of the same units, but the means may be different. Assume that two brands of electric bulbs have the following mean life and standard deviation : | Company A | Company B (in hours) | Company C | |---|---|---| | Mean life | 1500 | 1800 | 2500 | | Standard | 50 | 90 | 100 | | deviation | | | | Here the lowest absolute dispersion measure is 50 i.e. of Company A. However, we cannot say, for certain that the life of company. A’s electric bulbs is relatively’ less spread or scattered. Here measures of relative dis-persion are a better guide. Co-efficient of Variation V = — x 100 x .... (3.11.1 (expressed as a percentage) The co-efficient of variation, expressed of a fraction is also sometimes, called the Co-efficient of Standard deviation and’ is equal to . x Thus, for the relative dispersion of the three/ brands of electric bulbs, we have V = 5l x 100 = -50- x 100 = 3.33%. A   x11500 = 52 x 100 = -90- x 100 = 5%. x21800 V = ^3 x 100 = 50 x 100 = 4%. c x32500 It is evident that the series A is relatively less spread and hence, more concentrated around its mean. (c)    The series to be compared may be expressed in different units, and we cannot therefore, compare the standard deviations. Similar to the co-efficient of standard deviation are the co-efficient of quartile deviation and mean deviation. Co-eff. of Quartile Deviation Q3-Q1 Q3 + Qi (3.12) Sx    ¿Med .... (3.13) and Co-eff. of Mean Deviation or = or X    Med Thus for any absolute dispersion measure, we can find out a relative dispersion measure, Absolute Dispersion Relative Dispersion = Average used Xh-X1 Thus Co-eff. of Range =            X + X             .....(3-14) where Xh is the highest value of an item and X is the lowest value of the items in the series. Co-efficient of 10-90 percentile Range = P -P 90    0 P +P, 90 + 10 (3.15) Diagram Percentage of Area under the Normal Carve Self-Check Exercise-1 Q1. Define Dispersion. What are the various measures of dispersion? Q2. Calculate the standard deviation of the following series: 7,    10, 12, 13, 15, 20, 21, 28, 29, 35 Use the assumed mean method. 2.4 Skeweness Skewness is the degree of asymmetry or departure from symmetry of a distribution. Measures of skewness not only indicate the magnitude of skewness but also its distribution, A distirbution is said to be skewed in the direction to the extreme values or in the direction of excess toil, for a frequency curve. Most of the skewed curves in social sciences are skewed to the right. It is only very rarely that we meet curves skewed to the left and even more rarely, do we find series characteristically skewed to the left i.e. negatively skewed. A large number of series and frequency distributions ate, however, characteristically skewed to the right (or positively skewed). Frequency distributions of wages, salaries, incomes, earnings, sales, output, consumption, and of a number of other socio, economic and business variables are likely to be sharply stewed to the right. Diagram Fig. 5.3 Croxten, Cowden and Klein cite the distribution of ages at death of 371 America inventors, which hap-pened to be characteristically skwed to the left. The plausible reasons cited are : (i)    “Younger men do not often have enough inventions to their credit to be classified as “inventors” (ii) “a time factor is present—almost one fifth of the inventors included in this study were born before 1800.” We have seen in the preceding lesson that the per-centage of the extreme, values does not affect the mode, the median is affected by their position only and the arithmetic mean in influenced by their size. Karl Pearson used these characteristics of the mode and the mean to measure skewness. An absolute measure of skewness, then, would be Skewness = Mean — Mode                    ...... (3.16) Since, in a moderately skewed distribution of a continue is variable, Pearson showed that the median travels 2/3 the distance from the mode towards the mean the value of mode could be written as M0 + X - 3( X -Med.) = 3 Median - 2 Mean. Substituting the expression so the mode in the measure of skewness (3.16) we get Absolute Skewness = X - [ X -3 ( X - Med.)] = ( X - Med) = (Mean - Median)                ...... (317) However, measures of absolute skewness suffer, from certain obvious defects. First, these measures are expressed in terms of the unite of the problems. Thus, heights would be measured in centimeters and weights, in kilograms. We cannot, obviously, compare the skewness in the distribution of Heights with that in the distribution of weights without bringing them to a common denominator. This can be achieved by dividing the absolute measure by a measure of dispersion, such as the standard deviation. Hence, Sk = X-M0 σ (3.18) Since, in income distributions, mode is only an approximation, M can be more satisfactorily located and hence used for the purpose. 3(X - Med) .... (3.19) S = kσ The above two measures are called, Pearson’s first and second Co-efficient of skewness. Similarly, we know that in a symmetrically distribution, median lies exactly half-way between the two quartiles, as also between the 10th and 90th percentiles. Hence, when the median does not lie in this position, some asymmetry or skewness is surely present. Based on this position of the median vis-a-vis the quartiles and the 10th & 90th percentiles, we have quartile and percentile coefficients of skewness. These are Sk (Q3 -Med)-(Med-Q1) and S k Q3-Q1- 2Med. =    Q3 - Qi P90 -P10- 2Med. p -P 90     10 .... (3.20) .... (3.21) However, these measure of skewness are not as satisfactory as Pearson’s ‘co-efficient, because they suffer from obvious deficiencies of the quartiles and percentiles. An important measure of skewness used is the third moment about the mean, expressed in dimension less form as follows : m3 Moment Co-efficient of Skewness = a =   . 3   ^3 m3 (7m)                        (3-22) where m2 and m3 are the second and third moments about the mean respectively. Another measure of skewness is sometimes given by β1 = α32. For perfectly symmetrical curves, such as the normal curve, α3 and β are zero. Ginni’s Mean Difference: Carrado Ginni has suggested alternative method of studying dispersion. The method is : D Ginni mean Difference = nq Sk = 3(x - Med) a Nq = Number of difference = l/2 n(n—1), Following Example illustrate the method. Example: Find Ginnis mean differences from the following items: X :      8      10     12     14     16 Solution : Given items ate 8, 10, 12, 14, 16 | 16-8=8 | 14-8=6 | 12-8=4 | 10-8=2 | |---|---|---|---| | 16-10=6 | 14-10=4 | 12-10=2 | | | 16-12=4 | 14-12=2 | | | | 16-14=2 | | | | | 20 | 12 | 6 | 2 | Now total of alt differences 20 + 12 + 6 + 2 = 40 11 Total Number of differences = 2 n (n—1) 2 5 (5—1) = 10 D 40 Gini’s Mean difference =     =    = 4 Ans. nq 10 Moments: If X1 X2   Xn, are the values assumed by the variable X, then the rth moment is defined as. X1 + Xr2r +......+ Xur _ EXur (3.23) N      = N The first moment (i.e. when r =1) is the arithmetic mean X. The rth moment about the mean X is defined as nror mr = E(X - X)r E(dx)r NN (3.24) where Dx = X-Xa is the deviation of X from Xa (the assumed mean), if X = 0, (2.25) is often called, rth moment about zero. The following relations exists between the moments, about the mean, mr or pr and moments about an arbitrary origin, vr and pr m1 = n1 = 0                                                     (3.26.1) m2 = ~- = v2 = v12                                               (3.26.2) m3 = n3 = v3 = 3v1 v2 +2v1                                   (3.26.2) where v1, to v2 and v3 are the 1st, 2nd, 3rd moments about an arbitrary origin m2 or n3 is a measure of absolute skewness. The measure of relative skewiess is 2 n β = (a2) = 33          ..... (3.27.1) n 2 ^3 ai = ^ = ^ (<)      (3.27.2) For symmetrical series β1 = O. The greater the value of β1 the more skewness there is in the series. Self-Chek Exercise-2 Q1. Define a) Skewness b) Moments 2.5 Kurtosis Kurtosis is the degree of peakedness of distribution, usually taken in relation to a normal distribution. A distribution having a relatively nigh peaked as the curve of Fig. 5A (a) is called leptokurtic while the curve of. Fig. 5.4 (b) (c) which is flat topped is called playkurtic. The normal distribution, Fig. 5.4 (b) which is neither, very peaked nor very flat topped, is called mesokurtic. *Kurtic Means Jump backed : thus hamped or un-normal. Lep to means slender narrow, Platy means broad, wide flat and Meso means in the middle. (c)LmWJJRHC          WMSSORUBTIC (c) PLATYEVETK? Fig 5.4 (a) Fig. 5.4 (b) The degree of kurtosis present in a series, may be measured by making use of the fourth moment aboutme mean. Sdx4 n4 or m4 = (3.28) Where dx stands for the deviation of the items from the mean or for a frequency distribution: Sfdx4 n4 or m4 = N~ The fourth, moment about the mean, can be expressed, in terms of the moments about any arbitrary origin, as under (in terms of class interval units). n = Sf (d ')4 4 N ^^^^a Sf (d')   Sf (d ')3 4 ------ X ------- N N + 6 (fYI ^^H 3 fw2 Y V N ) However n4 is an absolute measure of kurtosis. The moment Co-efficient of kurtosis is given by n 4 _ n4 P or a = 4 =   2                                   (3.30) 44 crn2 For the normal distribution β2 = 3 For the reason the kurtosis is sometimes defined which is positive for a leptokurtic distribution, negative for a platykurtic distribution and zero for the mesokurtic or normal distribution. Alternately, we can put this criterion as Type of Curve                        Valueof Leptokurticβ Mesokurticβ Platykurticβ k0 another measure, of kurtosis, which is some-times used, is based on both quartiles. and precentiles and is given by n4. k (pronounced as Kappa) = Q P -P. 90     10 (3.31) Where Q stands for the quartile deviation. This is sometimes referred to as percentile coefficient of kurto-sis. For the normal distribution, it has the value 0.263. Example 1. Let us compute the moments, moment coefficient of skewness and kurtosis for our data of percentage marks of students in Economics in Table 5.4. 75 + 30 71 + 84 + 15 283 m1' or v1 m2' or v2 m3' or v3 m4' or v4 sfD       1 N^~ x i = — i = 0.13 i = 0.13i. w N f3 N SfD’4 N X 91 i2 = 75 i2 = 1.213i2. X i3 = x i4 = 13 75 i3 = 0.173i3. 283 75 i4 = 3.77i4. | Percentage Marks of Students Mid- value X | Number of Students f | D' | fD'1 | fD'1 | fD'1 | fD'4 | |---|---|---|---|---|---|---| | 10 | 7 | -2 | -14 | 28 | -56 | 115 | | 30 | 15 | -1 | -15 | 15 | -15 | 15 | | 50 | 32 | 0 | 0 | 0 | 0 | 0 | | 70 | 12 | + 1 | +12 | 12 | + l2 | 12 | | 90 | 9 | + 2 | +18 | + 36 | +72 | 144 | -29         91         -71 Where i stands for the size of the class interval, 20, in the present case, m1 or np = 0. m2 or n2 = v2 - v12 = [1.213 — (013)2] i2= 1.213 i m3 or -; = v3 - 3vv + 2v23 = (0.173-3(0.13) (1.213) + 2(.013)3] i3 = 0.173 - 047 + .000] i3 m4 or n4 = v4 = 4v1v3 + 6v12 v2 - 3v14 = 3.738 i4 ^ Æ = , _ n3 _   0.126 i3 _ .07 '3“ TÈ ’ 7Ï12ÏW} A = «4 = n4 _ 3.738 " 1.471(i2)2 _ 2.565 The values of α3 and α4 show mat the distribution is veryslightly skewed positively and is tending to be platykurtic. 1.    Before the representatives of a measure of, central tendency can be assessed, the dispersion or spread of the data needs to be known. 2.    Of the Various measures of absolute dispersion, the standard deviation is, by far, the most important, and most frequently used. (i) . The range = Xn — Xi i.e. the difference between, the highest and lowest values of the Variable xi is a very crude measure and uses only two extreme values-the lowest and the highest. (ii) . The 10— 90 percentile range = P90- P10, though an improvement upon the rang, suffers broadly from the same defects as the range.” (iii) The quartile deviation, Q _ Q Q 2 encompasses the middle 5% of items. Like the 10—90 percentile range, it is not affected by ‘extreme values’. ‘However,’ like the range the 10 - 90 percentile range, it fails to consider the values of all, items. (iv) The average or mean deviation. ô = Y I dv I     Yf\dx I ------or------- N or N is usually computed in relation to the mean, even though, S | dx | or S | fdx | i.e the sum of deviations, neglecting the sign is the minimum in relation to the median, Despite the feet that it uses all the items in & series, it is very rarely used from frequently distributions. (v)    The standard deviation is the square root of fee mean of the squared deviations of the values from the arithmetic mean. A number of formula are ‘available for computing the standard deviation, o as follows : (a)    The formula by long method is, < S(x x2)|_ ^f^dx2) I N ( _ f ) J I N ) (b)    The formula for the short-cut method (especial suited to machine calculation) is | | V 2 ( YX V | |---|---| | | Sx2 - --- | | I (Sx2 ^ (SXV | I N ) | | 41 N H N J J or | N | | | J | (c)    The formula for the short method using deviations (D) from an arbitrary or assumed mean Xa is YD2 (YD I2' N N I 5 f or YD2 -DD^ k N N (d)    The variant of (c) above, especially when the distance between successive values of the variable are equal and greater than of less than 1, uses equal distance as common interval (i) and the deviations from the assumed mean, in terms, of this common interval(i); D = D i or i X^< i X^< YD'2 (YD V' N YD'2 -------------------------1 N YD' V' k N N For grouped data these deviations are the mid values of the classes from the true or assumed mean, and frequencies are introduced and the ‘formula are as follows (a) j = JWzX) 2 L n(=v) J ■(f-' k N J 2 (b) j = Jf2 N 2 - or N 2 (c) j =JYD N 2 - or (d) (J= i Xjf N - 2 j^ N | | 2 W2-RD-1 | |---|---| | i x\ | V N ) | | or | N | | | [                           J | 3.    The Varaince = a2. 4.    Properties of the standard deviation are: (i)    The sum of squared deviations is the minimum when the deviations are computed from the arithmetic mean, rather from any other origin. (ii)    For normal distribution the following percentage of observations lie within specified ‘multiples’ of standard deviation from the Mean, in both direction; | | Range | Percentage of item range of limits | |---|---|---| | X | X = o | 68.27 | | | X = 2o | 95.45 | | | X = 3o | 99.73 | For moderately skewed distributions, these percentages aero approximately valid. This property helps us in computing the approximately valid. This property helps us in computing the actual percentage of, observations within specified standardized units range, and then compare these with the expected or theoretical frequencies for the normal curve, and thus find out; how far the distribution departs from the normal. (iii)    The deviations of the individual values from the mean i.e. X, X have been standardized by dividing o, i.e. Z1 = X1 X a (iv)    We have can compute the combined variance (or standard deviation), if we know the a, s of the various sets, and the number of items in each set. 5.    Sheppard’s correction for variance of grouped data is a —C2/12, where C is me size of the class intervals. This correction is however; a subject of controversy regarding its usefulness. 6.    For moderately skewed distributions, a X = 4 a 5 Q=3a 7.    Measures of absolute dispersion fail to guide us, if two or more series have different means or if these series are expressed in different units. In such cases, we have to find measures of a relative dispersion, the most important among these being the coefficient variation; V =    expressed X as a fraction or percentage. Percentage In other measure of relative dispersion, too, we divide the measure of absolute dispersion by some sort of average. Absolute dispersion Relative dispersion = Average used 8.    Skewness is the degree of departure, from symmetry of a distribution: it may be positive or negative. Most distributions in social sciences are skewed to the right. The coefficients of skewness are given by the following formulae: Karl Pearson’s first co-efficient of skewness = X-M0 σ (ii)    Karl Pearson’s Second co-efficient of skewness 3(X - Med.) especially, where σ mode is not defined, and is only an approximation. (iii)    Quartile co-efficient of skewness = Q3-Q1 Q2+Q3. (iv)    Percentage co-efficient of skewness P -P. 90     10 . 90 + 10 ΣXr 9. (i) The rth moment of variable X is N the first moment being the mean. (ii) The rth moment about the mean X, is defined as mr = πr = Σ(X-X)r N Σdxr N (iii) The rth moment about any origin, is X vr or π1r = Σ(X-Xa)r N ΣDrx N moments about the mean can be expressed in terms of moments about an arbitrary origin. 10.    Kurtosis is the degree of peakedness of a distribution. Distributions may be leptokurtic (with narrow peaks); mesokurtic (with middle or intermediate peak) and platykurtic with falt peak). Kurtosis is measured by the following co-efficient π4 2 If β = 3 the distribution is normal, σ2     2 π4 (i)    Moment co-efficient of kurtosis, β or σ = 4 = 2    4σ if β2 < 3, it is leptokurtic and if β2 < 3, then it is platykurtic. Q (ii)    Percentile measure of kurtosis. k = P90 + P10 For normal distribution, k = 263. Self-Check Exercise-3 Q1. Define Kurtosis. Q2. With the help of graph define a)     Leptokurtic b)    Mesokurtic c)     Platykurtic 2.6    Summary 2.7    Glossary •    Coefficient of Variation: A relative measure of dispersion that is the ratio of the standard deviation to the mean, expressed as a percentage. It allows for comparison of variability between data sets with different units or scales. •    Dispersion: A statistical measure that describes the spread or variability of a data set. It indicates how much the data points differ from each other and from the central value. •    Kurtosis: A measure of the “tailedness” of the probability distribution of a data set. High kurtosis indicates heavy tails and a sharp peak, while low kurtosis indicates light tails and a flatter peak. •    Measures of Absolute Dispersion: These are measures that provide the actual magnitude of dispersion in a data set, such as range, quartile deviation, mean deviation, and standard deviation. •    Measures of Relative Dispersion: These measures provide the spread of data relative to the central value, often expressed as a percentage, such as the coefficient of variation. •    The 10-90 Percentile Range: The range between the 10th and 90th percentiles of a data set, which excludes the extreme values and provides a more robust measure of dispersion. •    The Mean of Average Deviation: The average of the absolute differences between each data point and the mean of the data set. It provides an average of how much the data points deviate from the mean. •    The Quartile Deviation or the Semi-interquartile Range: Half of the difference between the first quartile (Q1) and the third quartile (Q3) values in a data set. It measures the spread of the middle 50% of the data. •    The Range: The difference between the maximum and minimum values in a data set. It is the simplest measure of dispersion. •    The Standard Deviation: A measure that indicates the average amount by which data points deviate from the mean. It is the square root of the variance and is widely used in statistical analysis. •    Skewness: A measure of the asymmetry of the probability distribution of a data set. Positive skewness indicates a distribution with a long right tail, while negative skewness indicates a distribution with a long left tail. 2.8    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 2.3. Answer to Q2: 8.76. Self-Check Exercise-2 Answer to Q1. Refer to Section 2.4. Self-Check Exercise-3 Answer to Q1. Refer to Section 2.5. Answer to Q2. Refer to Section 2.5. 2.9    References/Suggested Readings 1.    Stockton, J. R., & Clark, C. T. (1975). Introduction to business and economics statistics. Southern Western Publishing Co. 2.    Merrill, W. C., & Fox, K. A. (1970). Introduction to economic statistics. Wiley. 3.    Jain, T. R., &Aggrawal, S. C. (2020). Business statistics. V K Publications Pvt. Ltd. 2.10    Terminal Questions Q1. Discuss the various measures of dispersion.\ Q2. Why is standard deviation the most widely used measure of dispersion? Q3. Calculate arithmetic mean, median, mode and standard deviation for the following series: | Daily Wages (Rs): | 24.5-34.5 | 34.5-44.5 | 44.5-54.5 | 54.5-64.5 | 64.5-74.5 | 74.5-84.5 | |---|---|---|---|---|---|---| | No. of Workers: | 4 | 24 | 62 | 86 | 96 | 100 | ***** UNIT-3CORHELAHON AND REGRESSION STRUCTURE 3.1    Introduction 3.2    Learning Objectives 3.3    Correlation Self-Check Exercise-1 3.4    Degree of Correlation Self-Check Exercise-2 3.5    Methods of Measurement of Correlation 3.5.1    Scatter Diagram 3.5.2    Correlation by Graphical Method 3.5.3    Coefficient of Correlation 3.5.3.1    Karl Pearson’s Coefficient of Correlation 3.5.3.2    Short-cut Method 3.5.3.3    Calculation of Coefficient of Correlation in Continuous Series 3.5.3.4    Coefficient of Correlation by Rank Self-Check Exercise-3 3.6    Regression Lines 3.6.1    Why There Are Two Regression Lines 3.6.2    Properties of Regression Coefficient 3.6.3    Limitations of the Theory of Linear Correlation Self-Check Exercise-4 3.7    Summary 3.8    Glossary 3.9    Answers to Self-Check Exercise 3.10    References/Suggested Readings 3.11    Terminal Questions 3.1    Introduction Dear Friends, We have so far been studying characteristics of series involving a single variable, consumption expenditure, heights, weights etc. In this lesson, we are going to study a technique, very frequently used in economics and business, research, and by applied statisticians, namely, that of correlation and regression analysis. In the early days Correlation analysis found widespread use in biological problems,’ but subsequently, it has been extensively used in economics, agriculture, and many other fields. 3.2    Learning Objectives After going through this unit, you will be able to •    Understand and Describe Key Concepts of Correlation and Regression. •    Apply Methods to Measure and Interpret Correlation. •    Understand and Utilize Regression Analysis Techniques. 3.3    Correlation The term correlation or co-variation indicates the ‘relationship between two such variable in which, with changes in the values’ of one variable, the values of the other variables also change. For example, we may be interested in finding out the relationship between the number of years of education completed and the income of adult males in a community or we may be interested in relating the crime rate, with the incidence of employment. In some of these problems, we are often not only to describe the nature of relationship between file two variables so that we can predict (or estimate) the value of one variable, if we know the value of the other. For instance, we may want to predict a person’s future’, income from his, education. ‘When we are, principally, interested in them exploratory task of finding out which variables are related to a given variable, we are likely to be mainly interested in measures of degree of relationship such as, correlation co-efficient. On the other hand, once we have found the significant variables’, we are more likely to turn our attention to regression analysis. We may take up, correlation first and’ shall take up regression, in the next lesson. Correlation can either be positive or negative. When the values of two variables move in the same direction so that an increase in the value of one variable is associated with an increase in the value of the other variable also, and a decrease in the value of one variable is associated with the Decrease in the value of the other variable also, correlation is said to be positive or direct if the values of two variables move in different directions in such a way that with an increase in the value of one variable the value of the other variable decreases, and with a decrease in the value of the variable the value of the other variable increases, correlation is said to be negative. In exact sciences, the study of correlation is easier because mathematical relationship can be established between values of two variables on the basis of experiments. The effect for example, of heat on density of air can be reduced to a mathematical formula disclosing the relationship between these two variables. Economic and social data are affected by a large number of causes. We cannot determine how much a particular cause contributes to a given effect. Increase in price will lead to increase in supply, and vice versa but supply is affected by many other factors also. We cannot make the other factors inoperative for the time being as we can do in experimental methods. Therefore, measurement of correlation is relatively more difficult in social sciences. Even, when correlation has been established and measured, it does not mean that each and every item would confirm to the same pattern. Tall fathers will have tall sons on the average, but some tall fathers may have short sons. Self-Check Exxercise-1 Q1. Define Correlation. Q2. What is the significance of studying correlation? 3.4 . Degree of Correlation : When changes in the two Mated variables are exactly Proportional, there is perfect correlation. Correlation-will-men be said to be linear. If a 10% increase in price is every time accompanied by a 15% rise, in supply, the relationship between these two variables is linear and of the type y = a + bx. The Values of such series when plotted on a graph paper will fall exactly on a straight line. When such a straight line rises from left to right upward, correlation will be perfect and positive. When the straight line falls from left/to right downwards,’ correlation is perfect but negative. Self-Check Exercise-2 Q1. What do you mean by degree of correlation? 3.5    Methods of Measurement of Correlation But such perfect linear-correlation are fare in social sciences. If changes in the two variables, are in the same direction but not in the same proportion, the correlation is positive but less than perfect if changes are in the opposite direction (and not in the same proportion), the correlation is negative and limited. Correlation between two series, may be studied by one of the following methods. 1.    Scatter diagram. 2.    Correlation graph. 3.    Coefficient of Correlation. 4.    Correlation table. The third method of computing a coefficient is the most important method. 3.5.1    Scatter Diagram: Diagrams and graphs can be drawn to have an idea about the correlation, between two variables. Suppose we are given figures of ages of hundred wives and husbands. If the age of husbands is represented by x and the age of wives by y, we shall have hundred pairs, of x and. y values. Now plot variables, (age of husbands) on the horizontal axis and y, (age of wives) on the vertical, axis. We shall obtain, one point for each set of two values. When, we have plotted data about, these hundred wives and husbands we shall get hundred such points. On the horizontal scale each point represents the age of the husband and at the vertical scale, the age of wife. The diagram showing, these, hundred points would be called a Scatter Diagram. +reCORSMW -W CORRELATION X2TKXXMXX ^tt^xaf«KU; J HMiftstr#«« f Nd CORRELATION If these points show definite tendency or trend upwards or downwards; the two variables are correlated. If the points lie around a diagonal band rising from left bottom to the right top the correlation is positive. If the points are scattered around a diagonal band from left to top to right bottom, the correlation is negative. Closer the points are to the diagonal higher is the degree of correlation, if the points, of scatter diagram be exactly on straight line, the correlation is perfect. If the plotted points scattered widely over the whole graph so that no evidence of definite, trend or tendency in any direction is available, the correlation does not exist. A scattered diagram can only provide rough idea, about the degree of correlation and its direction i.e. whether the correlation is positive or negative. It does not provide any exact, measure of correlation between two variables. 3.5.2 . Correlation by graphic Method Another way of detecting positive or negative correlation and also its extent is to draw correlation graphs and read the direction of the curves. You have to be very careful, while, drawing a correlation graph about the choice of the series and the base line. They should be so chosen dial the averages of the two variables do the vertical scale should be as close to one another as possible. If the curves or graph representing two variables move in the same direction in the same ranges, the correlation is positive; if they move in opposite directions, the correlation is negative. And if the two curves show no’ particular pattern, in their movements, sometimes moving in the same direction and other times in opposite direction, sometimes one rising or falling and the other remaining constant correlation does not exist at all. 3.5.3    Coefficient of Correlation Correlation exists in various degrees. It is perfect and positive when an increase (or decrease) in one variable is always followed by a corresponding and proportional increase (or decrease) in the other variable. Correlation is perfect but negative if an increase or decrease in one variable is always, followed by a corresponding and proportional decrease (or increase) in the other variable. Correlation; will be non-existent or zero if changes in one variable cannot be associated at all with changes in other variable. In between positive perfect correlation and no correlation; there shall be different degrees of positive correlation, and similarly in between perfect negative correlation and no correlation, there shall be different degrees of limited negative correlation. Coefficient of correlation is calculated to study the extent or degree of correlation between two variables. Generally Karl Pearson’s Coefficient ‘ of Correlation is used. This Coefficient varies between + 1 and —1. Perfect positive correlation is indicated by + 1. Perfect negative correlation by —1 and complete independence by 0. Limited degrees of positive correlation are indicated by values lying between 0 and +1-and limited; degrees of positive correlation by values tying between 0. and —1. All this, is true also of Spearman’s Coefficient of Rank Correlation. 3.5.3.1.    Karl Pearson’s Coefficient of Correlation: The formula devised by Karl Pearson, the great biologist and statistician, measuring coefficient of correlation between two variables is the most satisfactory. The formula is as under: Σdx dy n σ1σ2 ..........3.1 Here r stands for the coefficient of correlation, dx for the deviations ,of values of, x variables from the arithmetic mean of x series, dy for the deviations of y values from the arithmetic mean of y series, n for the number of pairs of observations, σ1, for the standard deviation of the x-series and σ2 for the standard deviation of the y-series. In some books, you will find this formula written as: r = Σxy n σ1σ2 Here x and y are the same things as dx and dy in formula 3.1 above Thus the coefficient of correlation between variables is obtained by dividing fee sum of the products of the corresponding deviations of the various items of two series from their respective mean by the product of their standard deviations and the number of pairs of observations. Σdxdy n σ1σ2 is the basic form of Pearson’s formula. All me number forms are derived from this fundamentals form. Example 1. The following table shows the marks obtained by ten students in Economics and Statistics. Find the coefficient of correlation; Marks in Eco. 78    36    98    25    75    82    90    61    65    39 Marks in Stats. 84    51     91     60    68    62    86    58    53    47 Solution: | | I | II | III | IV | V | VI | VII | |---|---|---|---|---|---|---|---| | Sr. No. | Marks In Eco. x | Deviation fromdy=(65) dx | d2x | Marks in Stats y | Donation from =(66) dy | d2y | dxxdy | | 1 | 78 | + 13 | 169 | 84 | 18 | 324 | + 234 | | 2 | 36 | - 29 | 841 | 51 | -15 | 225 | + 435 | | 3 | 98 | + 33 | 1089 | 91 | + 25 | 624 | + 825 | | 4 | 25 | - 40 | 1600 | 60 | - 6 | 36 | + 240 | | 5 | 75 | + 10 | 100 | 68 | - 1 | 4 | + 20 | | 6 | 82 | + 17 | 289 | 62 | - 4 | 16 | - 68 | | 7 | 90 | + 24 | 625 | 86 | + 20 | 400 | + 500 | | 8 | 72 | - 3 | 9 | 58 | - 8 | 64 | + 24 | | 9 | 65 | 0 | - 0 | 53 | - 13 | 169 | + 0 | | 10 | 39 | - 26 | 676 | 47 | - 19 | 361 | + 494 | 0           5398     = 660                  =0 2224 =2704 Σx = 650                      Σd x 2     Σy                    Σd y 2 Σdxdy Marks in Economics (x) and marks in statistics, (y) are given in columns I and IV respectively. 650 10 = 65 Σx A.M. of x series is = n 660 10 = 66 Σy A.M. of y series is n Standard deviation of x series - I ÌYdx2 ì   J5398^ a V ----- X X (539-8) ^ n J \ 10 J Standard Deviation of y series - a, _ V i^dy2-ì = V [2224ì = V(222.4) ^ n J      \ io ddx/v _ r ~ n CT1CT2 " 10(539.8)x(222.4) 27042704 = 10 x 23.2x 14.7 = 3456.8 = 0’78 approximately- The algebraic sign of r will be the same as that of Ydxdy.n, the number of pairs of observation cannot, be negative, o1 and 02 cannot be negative. The denominator of ddx.dv ii ya2 cannot be negative. If therefore, Ydx dy is positive, r will also be positive if Sdxdy is negative, r will be negative. If can be proved that r or Zdx.dy cannot exceed the numerical value of 1. We need not take na1a2 the trouble of proving this. But we must always be aware that 1 (— l) correlation is negative) is the maximum or minimum value that r can have. If in a problem your r comes to be greater than 1, it simply means that your calculations have been wrong somewhere. The formula (3.1) above is ddx.dy r =------ nala2 If values of o1 and 02 inserted, the formula becomes r = Ydx.dy (^dx2 ^ (^dy2 Ì IJ! n J (3.2) In the denominator, the two n’s within the under root sign’s will cancel with the one a outside. Formula (3.2) will then be reduced to _ ddx dy (3.3) (^dx2 ) x J(^dy2 ) From formula (3.3) it should be clear that we need not find standard deviations in order to calculate r. Let us find r in Example 1 by applying formula. (3.3). _      Ddx.dy yl(Ddx2 ) x 7(Ddy2 ) 2704         2704 r = ^^=^=--=---- 7(5598) x 7(2224)   3456.8 = 0-78 approximately exactly the same as obtained earlier. 3.5.3.2.    Short cat Method : In example (1) above arithmetic means of both x and y variables are complete numbers and, therefore, finding deviations of value from these arithmetic means and processing, these deviations, mathematically is not difficult. If however, arithmetic means are in fractions finding deviations and processing them becomes a tedious job. In such cases, certain short-cut-methods can be used where assumed average is used for calculating the coefficient of correlation. This is what we have done for calculating standard deviation, in lesson-2. One of the following shortcut formula can be used. r _ZDx.Dy -XDx.Dy/n ^1^2 Here, XDxDy = sum of products of deviations from assumed means, SDx = Sum of deviations of x values value from assumed mean X a X Dy = sum of deviations of y valises from, assumed mean ya o1 = a standard deviations of X series, 02 = as standard deviations of y series, If in formula (3.4), we insert the formula of o1 and 02 the formula becomes r = DDx.Dy / n - (DDx / n)(DDy / n) | | ^DDx2 (DDx)2 ' nn | i V | \DDy2 (DDy )21 . n         n | |---|---|---|---| DDx.Dy -DDx DDy / n) Ux 2 - (^ Uy 2 - (^ (3.4) r= > x > (3.5) n n r      nD Dx.Dy -DDx DDy / n )_______ [>Dx2 -(DDx)2}x J^DDy2 -(DDy)2}]                     (3-6) (Work out yourself how - outside the square root signs in the numerator has been cancelled). Formula (3.1) is the most convenient and most commonly, used from of the formula for calculating Persons correlation coefficient. You will recall that Xa = Xa + YDx n where x is, is arithmetic mesh of x values xa is the assumed mean and YDx is the sum of deviations from assumed mean. SDx _ -            _ - n = x — Xa = or ^Dx = n ( x - Xa )• Similarly, SDy = n ( y - ya ) Substituting these values of Dx and Dy in formula (7.5), we get r = Y,DxDy -n(x -xa /n)(x -xa)/n 2 ID2 - [n(x — xa )] [DDy 2 - [n(y — ya ) k             n ) k XDx.Dy - n(x - xa ) +(y - ya) 7|sDx2 - n(X - xa)2 J7 |w2 - n(y - ya)2 J Example II: Calculate the co-efficient of correlation between the values of x and y given below; x     78    89    96    69    59    79    6861 y     125   137   156   112   107   136   123108 Solution: | x | y | (x-69) =Dx | (y— 112) =Dy | Dx2 | Dy2 | DxDy | |---|---|---|---|---|---|---| | 78 | 125137 | 920 | 13 25 | 81400 | 625 | 117 500 | | 96 | 156 | 27 | 44 | 729 | 1936 | 1188 | | 69 | 112 | 0 | 0 | 0 | 0 | 0 | | 59 | 107 | -10 | - 5 | 100 | 25 | 50 | | 79 | 136 | -10 | 24 | 100 | 576 | 240 | | 68 | 123 | - 1 | 11 | 1 | 121 | -11 | | 61 | 108 | - 8 | - 4 | 64 | 16 | 32 | | ZDx = 47 | SDy | SDx2 | XDy | SDxDy | |---|---|---|---|---| | | = 108 | = 1475 | = 3468 | = 2116 | Assumed mean of x series = 69 Assumed mean of y series =112 Applying Formula (3.5) r = ZDx.Dy -^Dx. ZDy / n ^Dx2 - (ZDx )2 A n ^Dy2 -k (W1 n ) r = 2116 - (47 X108/8) 1475 - k « 1 3468 - X k nr 1 8 J r = ,   11852   , = 954 (2591)(16080) Simplification of the terms here is a tedious job. You must have lot of practice if you want to avoid, waste of your precious time in the examination. You can also use logarithms for this simplification. If zero is taken as the assumed mean for both x and y variables, the deviations will be nothing but the original x and y values themselves. Formula. (3.5) will then become. r = Sx v   Sx Sy Sxy n . 2 (Sx )2 n v sy 2   (Sy )2 (3.7) n With this formula we can find co-efficient, or correlation if we are given : (1)    The number of pairs of observations. (2)    Sum of the X values and sum of Y values. (3)    Sum of the squares of X values and sum of the squares of Y values. (4)    Sum of the products of values of X and Y. variables. 3.5.3.3.    Calculation of Coefficients of correlation in continuous series. If the values of the two variables are grouped and the frequency of different groups are given two way tabulation is necessary in order to find out the coefficient of correlation. Suppose the two variables are marks obtained in Economics (x) and marks obtained in Statistics (y) and they have, been grouped in class intervals. They are given below: r= SfDxDy - (SfDx. SfDy )/ n (SfDx)21 SfDy2 - (SfDy )21 | y Marks In Statistics I | X Marks in Economics | Total fy | |---|---|---| | 5—15 | 15—25 | 25—35 | 35-45 | | 0—10 | 1 | 1 | — | — | 2 | | 10—20 | 3 | 6 | 5 | 1 | 15 | | 20—30 | 1 | 8 | 9 | 2 | 20 | | 30—40 | — | 3 | 9 | 3 | 15 | | 40—50 | — | — | 4 | 4 | 8 | | Total fx | 5 | 18 | 27 | 10 | 60 | We want to find the coefficient of correlation between marks in Economics and in Statistics of these 60 boys. Formula (3.5) will assume the following shape when frequencies are also given This is the most convenient formula for use in grouped data: In the data given above, we can find S/Dx, ^fDy ^fDx2 SfDy2- Please find out these values; Remember that deviations are being taken from assumed means/If you. have any doubt anywhere, revise the techniques of finding standard deviation. Understand this table very clearly and carefully. In the cell against 0—10 and below 5—15 groups, “we find the-value 1. This is the frequency. It means that there is 1 student whose score in Statistics is between 0—10 and in Economies’ between 5—15. Similarly, for all other entries in these cells. Examples III: Calculate coefficient of correlation from the grouped data given above : Solution : | MARKS (ECO.) X ^ MIDPOINT Dx MARKS (STAT.) MIDD’x Y ^           POINT | 5-15 | 15-25 | 25-35 | 35-45 | | |---|---|---|---|---|---| | 10 | 20 | 30 | 40 | | -10 | 0 | 10 | 20 | | -1 | 0 | 1 | 2 | Total fy | fD'y | fD' y2 | | 0 - 10 | 5 | -20 | -2 | 1 | 1 2 | - 0 | - | 2 | - 4 | 8 2 | | 10-20 | 15 | -10 | -1 | 3 | 8 3 | 5 0 | 1 -5 -2 | 15 | -15 | 15 -4 | | 10-30 | 25 | 0 | 0 | 1 | 3 0 | 9 0   0 | 2 0 | 20 | 0 | 0 0 | | 30-40 | 35 | 10 | 1 | — | 3 | 9 2   2 | 3 6 | 15 | 15 | 15 15 | | 40-50 | 45 | 20 | 2 | — | — | 4 8 | 4 16 | 8 | 16 | 32 24 | | | Total fx | 5 | 18 | 27 | 10 | 60 | 12 | 70 37 | | fD'x | -5 | 0 | 27 | 20 | 42 | | | fD'x2 | 5 | 0 | 27 | 40 | 72 | | fD'xD'y | 5 | 0 | 12 | 20 | 37 | SfD 'x = 42      f 'x = 12 f 'x2 = 72      If) 'y2 = 72 SfD ’xD y = 37          x = 60 Formula (7.8) for grouped data is ZfD' xD y - (f x. ZfD' y ) / n YfDx2 - f x)2 I                  n I fy2 - fx)2 i nI D' x = —; D' y = Dy iJ where i and j are common factors in the series of Dx and Dy. All the values required for this formula have been calculated in the lengthy table above, x values are given, horizontally, and y. values vertically. Step deviations of the x values are given us — 1, 0, 1, 2 and of the y values as — 2, —1, 0, 2. Once the step deviations are taken, farther adjustments at any stage are not necessary. This is true even if common factors in the step deviations of x and y variable are unequal. Frequencies corresponding to step deviations of y variable occur in in the row marked fx immediately below the calls. To find fD’x, multiply the frequency (fx) by the corresponding step deviation given in fee row immediately above the frequency cells. fD’x values are given in the row below fx. To findfD 'x2 simple multiplefDx by D'x. Summations of fDx andfD'x2 gives us YfD'x. Frequency corresponding to step deviations by y variable are given in the column fy immediately after the frequency cells. By multiplyingfy by step deviations we can getfD ' y and multiplyingfD ' y further again by D 'y we get value fD 'y2 their Summations gives YfD 'y2 and YfD 'x2. n is the sum of frequencies. ZfD'xD'y still remains to be found f D'xD 'y is calculated in each frequency cell. For example, in the frequency cell against D 'x = — 1 and D 'y = — 2, frequencyf = 1, :. fD 'xD 'y = 1 and D 'y = — 1 f = 5, : fD'xD'y as 5 (1) (—1) = —5. All values offD 'xDy written in the frequency cell have been underlined in order to distinguish them from the frequencies. These fD 'xDy values have been added in the last of columns as well as last of rows. This sum (whether obtained vertically from the column or horizontally from the row) will be SfD ’x D' y. In this example SfD 'xD ' y = 37. This is a lengthy table. You will have to exert considerably to understand it at once. Its sheer length and breadth should not frighten you into leaving it altogether. From the table : n = 60       fD'y = 12 fD 'x = 42    fD'y2 =70 fD 'x2 =72     fD'x D 'y = 37 r = SfD' xD y - (f' x. f' y) / n fx2 - (fxL <           n ) W y V 2 (fy)2 Ì n I 37 - 42 x 12/60 72 - (42)2 A 60 70 -k (12)21 60 ) 37 - 8.4 = .533 28.6 53.66 Probable Error of the Coefficient of Correlation The exact value, of various types of errors will be understood when we study sampling methods. We can however, get a working idea of the Probable Error of the correlation coefficient at this stage also. Suppose out of a student population of 100,00 we select 100 students at random and compute the correlation coefficient between heights and weights, Suppose this coefficient is r = 8 and suppose further that we select 200 more, random. Samples of 100 students each from this population, if we calculate correlation coefficients in respect of each of these, samples, there will be certain limits within which these coefficient will probably lie. “Probable error of the coefficient of correlation is the amount which if added to and subtracted from the mean correlation coefficient gives, amounts within which the chances are even that a coefficient of correlation from the series selected at random will fall.” The formula for P. E. of Karl Pearson’s correlation—coefficient is 1 - r2 . P E, or r = 0.6745 n Where r is the coefficient of, correlation arid n is the number of pairs of observations. If in a particular sample of 100 students, the correlation between heights and weights is 0.8 the probable error will be 0.6745 x 1 - r2               1 - 82 n = 0.6745   (100) 1 - 64 = 0.6745 10 (0.36) = 0.6745 10 = .24282/20 The coefficient of correlation in this case will, therefore, lie between and can be written as r = 0.8 ± 24282. If more samples were taken and their correlation coefficient computed, they will probably lie within 8 ± 024282, i.e. between .715718 and .824282. 3.5.3.4    Interpretation of Coefficient of Correlation Coefficient of correlation for a given’ pair’ of variables lies between + 1 (perfect positive correlation) and —1 perfect negative correlation. It may be asked whether the coefficient of correlation between two variables is “significant” or not. In order to answer this the following points may be kept in mind, they are based on the use of probable error of the coefficient. 1.    If r the less than its probable error, it is not at all significant, i.e. there is no evidence of correlation. 2.    If r is more than six times its probable error, correlation is significant. 3.    If the probable error is small and r is 5 or more, correlation decided exists and r is significant. 4.    Even if probable error is small but r is less than 3, conflation should be considered are not marked and r should not be considered as significant. You can now take tip few illustrations and decide for yourself whether r is significant or not. Coefficient of Correlation by Rank Differences Method of rank differences, is used to calculate correlation coefficient in, those cases where it is possible to arrange the various items of a series in a serial order, but the quantitative measurement of values is impossible or difficult or unpractical. Many attributes are incapable of direct measurement for example; intelligence, honesty beauty character etc. But the observer may be in a position to arrange items in a serial order and to assign ranks to different items. This method can be used also in such, places where dependable measurements cannot be obtained because of lack of finances or absence of investigation. This method involves easier mathematical calculations and it may be preferred for flu’s reason also. The next two examples will clarify this method. Example VI : Calculate coefficient of rank correlation from data of Example II above. Solution: | I x | II Ranks | III y | IV Ranks | V Difference or Ranks (d) | VI d2 | |---|---|---|---|---|---| | 78 | 5 | 125 | 5 | 0 | 0 | | 89 | 7 | 137 | 7 | 0 | 0 | | 97 | 8 | 156 | 8 | 0 | 0 | | 69 | 4 | 112 | 3 | 1 | 1 | | 59 | 2 | 107 | 1 | 1 | 2 | | 79 | 6 | 136 | 6 | 0 | 0 | | 68 | 3 | 123 | 4 | — 1 | 2 | | 57 | 1 | 108 | 2 | — 1 | 1 | | | | | | Σd = 0 | Σd2= 4 | Values of x variable appear in column I and y variable in column III. Our first task is to assign to x and y values. (In same problems, not values but ranks are given. There we do not have to assign the ranks). Start with x values. We, can assign rank 1 either to the smallest or the largest value. We have assigned rank 1 to the smallest value 57, The next large value is 59 which has been given rank 2. Next large” value is 60 which is given rank 3. The process is continued till we reach the largest value 97 of the series which has been given rank 8. The ranks of x values appear in column II. The process is repeated in y series, starting with the smallest “value 107 rank (f) and going up the largest value 156 (rank 8), y rank appear in column IV above. Next we find the differences between rank of x and y written in column V as (d). It makes no difference whether x ranks are subtracted from y ranks or vice-versa; y ranks from x ranks. The sum of these differences will be zero. In column VI are given the squares, of these rank difference d2Sd2 = 4. Coefficient of rank correlation or 6Ed2 e = 1 - / 2 n : [Ed2 = 4, n = 8] n(n -1) L                J 6 x 4 _ 1 - 24 e = 1 - 8(64 -1) =    504 _ 480 _.95 504 Rank correlation coefficient is also called Spearman’s correlation coefficient. The numerical value of rank correlation coefficient need not be the same as that of Karl Pearson’s coefficient for the same data. Example V. Calculate the coefficient of rank correlation for the following data: x     48    33    40    9     16    16    65    24    16    57 y      13     13    24    6      15    4      20    9      6      19 (PU 1964) | I X 48 | II Ranks 8 | III V 14 | IV Ranks 5.5 | V Difference of ranks (d) | VI d2 | |---|---|---|---|---|---| | 33 | 6 | 13 | 5.5 | 2.5 | 6.25 | | 40 | 7 | 24 | 10 | .5 | .25 | | 1.6 | 1 | 6 | 2.5 | -3.0 | 9.00 | | 16 | 3 | 15 | 7 | -1.5 | 2.25 | | 65 | 3 | 4 | 2 | -4.0 | 16.00 | | 24 | 10 | 20 | 9 | 2.0 | 4.00 | | 16 | 5 | 9 | 4 | 1.0 | 1.00 | | 57 | 3 | 6 | 2.5 | 1.0 | 1.00 | | | 9 | 19 | 8 | 5 | .25 | | | | | | 1.0 | 1.00 | Ed = 0 Ed2 — 41.10 A problem arises here in the assigning, of ranks because some values both in x and y series appear more than once: This is how the problem is resolved. Let us start with x values. The ‘smallest value 9 is given rank 1. The next higher value 16 occurs three times. If these three values of 16 had differed from one another, they would have been given the ranks 2, 3 and 4. Now that they don’t differ, average of the ranks 2.3 and 4 i.e 2 + 3 + 4 5 or 3 rank is assigned to all the three 16's. The next higher value 24 is given rank 5 Because ranks, 2, 3, 4 have been exhausted on 16’s. Therefore, whenever ties occur, the items are given the average of the ranks, they would have received if they had slightly differed. In y series, rank l is given to 4, the smallest value. After 4 comes 6, two times, if the two times had differed slightly, the rank assigned would have been 2, and 3, Average of 2 and 3 is 2.5. Therefore both’s receive the rank 2.5 is given rank 4, 13 occurs two times therefore both 13's 5 + 6 = 5.5 and so on. receive the ranks rece ve e ran s 2 You would have noticed that x rank appear in column II and y ranks in column IV. Rank differences are given in column V, squares of the ranks : d2 in column VI. Here n = 10 Sd2 = 41 Due to common ranks, coefficient of correlation has to be modified and the following, is used: e = 1 - 6Sd2 + S 112[(m3 — m)] n(n2 -1) (3.11) Here m stands for the number or times that a value has been repeated. In our example above three values, 16, 6 and 13 have occurred repeatedly, therefore 1/12 (m2- m) will be added three times. Since 16 repeats 3 times, 6 two times and 13. two times, the value of m will be successively 3,2 and 2. e = 1 - 6Sd2 +1/12 S[(m3 - m)] n(n2 -1) 6[41 +1/12(33 - 3) +1/12(13 - 2) +1/12(23 - 2)] = 1 —                    10(102 -1) =1 - 6HH 10 X 99 262 726 = 1--=--- 990 990 = + 73 A lot of practice is needed to master the techniques of calculating Karl Pearson’s, and Spearman’s coefficient of correlation. The independent variable the other as dependent variable.’ After estimating the relationship between two variables one can predict the most likely values of dependent variable on the basis of given values of independent variable. Self-Check Exercise-3 Q1. What is Scatter diagram and how it is useful in the study of correlation? Q2. Define the Karl Pearson’s Coefficient of correlation. Q3. From the following data, calculate Karl Pearson’s Coefficient of correlation: | X: | 2 | 3 | 4 | 5 | 6 | 7 | 8 | |---|---|---|---|---|---|---|---| | Y: | 4 | 7 | 8 | 9 | 10 | 14 | 18 | 3.6 Regression Lines A regression line is used to show the functional relationship between the variables and y. It is a line of “average relationship” as it ‘shows the relationship between x and on the average. If we have a problem of estimating y variable for some given values of x variable the requirement will be to construct a mathematical : equation of the form (y = a + bx). Such that y is dependent variable and x is independent variable. Such an equation known as regression equation of y on x can be used to estimate the most likely value of y. for some given value of x. Similarly, the problem which involves, the estimate of most probable or mean value of x for some given value of y will require the construction of regression equation x and v. This shows that from a set of observation of X and Y we can form two regression equation. This can further be illustrated with the help of the given ‘example’ Suppose, we want to study the relationship between the demand for wheat, and the, price of wheat in a certain market for a specified period. If our purpose is to estimate the demand for wheat corresponding to some arbitrarily selected price level’s then the regression line of wheat on price will ignore’ the variation of demand from one buyer to the other at the same price. Corresponding to this particular price, we may notice that some buyer is demanding larger quantity than some other corresponding’ each given price level. We may provide different demand levels. But the regression line of demand for wheat and price of wheat will give an estimate, of demand for wheat most likely to be bought at price, i.e. demand on the average. It shows what happens on the average to demand total variation in the demand for wheat may be split up into two parts. One part is explained by the regression line while the other part is not explained.’ 3.6.1    Why there are two regression lines ? Although both regression equation involve x and y, the two equation cannot be used interchangeably neither can be employed to predict both x and y. This is an important fact which the student must bear in mind constantly. The first regression equation y = ray/ax x x can be used only when y is to be predicted from a given x (when y is the dependent variable). The second regression _ ox equation r = fy x y can be used only when x is to be predicted from a known y (when x is the dependent variable). In summary, there are two regression equations in a correlation table, the one through me means of the columns and the other through the means of the rows. This is always true unless, the correlation is perfect. _ °x When r = 1.00 y = r =    x x becomes y = x or y on = xoy. Moreover when r = 1.00, oyo onox x = r oy y = y becomes x = r Oy X y or x °y = °xy In short when the correlation is perfect, the two regression equation are identical and two regression lines coincide. Regression Coefficients The least square regression equation, of y on x is written as y = a0 + b0 x. where a, and b, are two constants and can be obtained from the normal equations. Regression Coefficients The term regression coefficient is me name of the slope of regression line. Since there are two regression lines there will be two regression coefficients. The slope of regression line, of y on x is represented by the regression coefficient of y on x and has been denoted by the symbol b. Another convenient symbol is by x and it measures the change in y corresponding to civil change in x. When deviations are taken from the means then by Ixy   Ixy   roy x = v 2 =    2 = lx   nox   ox Similarly the regression coefficient of x on is the slope of the regression line of slope of the regression line of x on y. (2)    Regression Equation of x on y: Regression ‘Equation of x on y is x - x = b1 (y-y) x - b1 y x = V1/ ^ y 7 The alternative method is based on the assumption that deviations x and y are taken from the means of X and y respectively Regression Equation of x on y ox     _ ■x ■ x = r oy (y- y) and Regression equation of y on x is -     ly     _ y- y = r 1x2 (x - x )=r from these we get = (1y ) (1y2) (1x ) (1xy ) a1 =     Nly2 - (ly )2 X= b1y b = NSxy - (Sx) Sy 1   N(Sy 2) - (Sy )2 Alternative Method : Regression equations of y on x and of x on y can also be written in an alternative way. (1)    Regression equation of y on x is y - y = b0 (x - x ) or Y = b0x Y - y = y X - x = x , - Sxy 0 Sx2 a0 = Y = b0X y = b0x becomes ( Sxy A y =     J x SY a = a0N+b0 SX SXY = a0 SX+b0 SX2 from these we have (SY) (SX2)-(SX) (SXY) a0 =      NSX2 - (SX)2 or     a0 = - b0 Y N(SXY)-(SX) (SY) b0 =    N (SX 2) - (SX )2 a0 is Y intercept and b0 is the slope coefficient. Regression Equation of X on Y The least sq regression equation of X on Y is written as : X = a1 + b1 Y where a1 and b1 are constants and can be determined By the following normal” equations. SX = a1N+b1 SY SXY = a1 SY+bl Sy2 it is denoted, by b1 and another convenient symbol is bxy. As noted above Example 1 Sxy bxy =    2 S Y | X | 2 | 3 | 4 | 5 | 6 | |---|---|---|---|---|---| | Y | 3 | 5 | 8 | 7 | 12 | Suppose me following data are observed We shall use these data to illustrate the procedure of working calculation. We want to form the line yc = a0 + b0 X by the method of least squares Following table shows the calculations of various quantities needed for the estimation of a0 and b0. | X | Y | z | y | xy | x2 | y2 | |---|---|---|---|---|---|---| | 2 | 3 | -2 | -4 | 8 | 4 | 16 | | 3 | 5 | -1 | -2 | 2 | 1 | 4 | | 4 | 8 | 0 | 1 | 0 | 0 | 1 | | 5 | 7 | +1 | 0 | 0 | 1 | 0 | | 6 - | 12 - | +2 - | 5 - | 10 - | 4 - | 25 | | 20 | 35 | Sx = 0 | Sy =0 | 20 | Sx2 = 10 | Sy2 = 46 | | X = | SX n | = 20 = 4 5 | |---|---|---| | Y = | Sy = N | 35 = — = 7 5 | | bo = | Sxy Sx2 | = 20 = 2 10 | a0 = Y- bX= 7 - 2(4) = - 1 Yc = 2X - 1 if our aim is to form regression line Xc = b1 + b1 Y then we have to estimate a1 and b1. sY2 = 46 Sxy  2010 ...... bi  Sy2   4623 a1 = X-b1 = Y = 4-10 x 7 = 23 92 - 70 23 22 23 22 Xc = 23 + 10 23y Important Remarks (1)    Both the regression lines pass through X and Y . In example above. Y = 2X - l .................. I 22   10 X = 23 + 23 Y II | | 22 | 10 | |---|---|---| | From II | X = 23 | + 23 (2X + 1) | | | 22 | 20  10 | | | x= — | +     - | | | 23 | 23 23 | | | 12 | 20 | | | x= — | + | | | 23 | 23 | 23X = 12 + 20X 3X = 12 X = 4 Putting X = 4 we get Y = 8 - 1 = 7 Hence the point (4.7) Satisfies both the regression equations, Thus X = 4 and Y = 7. (2)    As is clear from above when two regression equations are given, then solving them simultaneously for X and Y will give mean value of X and mean value of Y. (3)    Making predictions is an important aspect of regression analysis. When we are to predict Y given X, then the most likely value of Y is to be predicted by using regression equations in which Y is the dependent variable and X is independent variable. Similarly when the problem is to predict X given Y, then regression equation of X on Y is to be used. Suppose we are to find most likely value of Y when X is 12. For this, we shall use regression equation Y0 = 2x (12) —1. From this Y0 = 2 (12) -1 = 23 (4) . If two regression equations are given then we can find out correlation coefficient by calculating the Geometric mean of two regression coefficients. In the above example Y0 = 2x -1 I 2210 and   X0 = 23 + 23 Y II From (I) b0 (=byx) = 2 10 From (II) b1 (=bxy) = 23 10 r2 = 2 × 20 23   93 3.6.2 Properties of Regression Coefficient (1)    The geometric mean between regression coefficient is coefficient of correlation Proof Sxy bo = byx = Sx Sxy bi =bxy = Ëy2 ^xy   ^xy byx bxy = ^x^ x sy 2 (Sxy)2 = Ex2 Sy2 = r2 Sxy = r = V (Sx 2)(Sy2) This proof can also be given as, byx bxy = ray ray ox   oy = r2 r = + byx.bxy Important Remarks (a)    If bxy is positive then byx will also be positive. (b)    If byx is negative, then bxy will be negative. (c)    Both regression coefficient, must have the demo sign. If byx and bxy are both positive, then r will be positive and when byx and bxy are negative then r win also be negative. (d)    bxy, byx < 1 2.    If one regression coefficient is greater than unity then other regression coefficient must be less than unity. Proof: Let bxy > 1, then we are to show that byx < 1. 1 If byx > 1 then bxy < 1 Now byx bxy < 1 1 or bxy < , byx 1 From (1) byx < 1 :. bxy < 1 2.    Arithmetic mean of byx and bxy is equal to or greater than coefficient of correlation. Proof : We are to prove that byx + bxy 2    - r byx + bxy if ----2----> r than byx + bxy > 2r or byx + bxy < 21 ± Jbyx.bxy or (byx + bxy ) ± 2   byx.bxy > 0 or (\byx + bxy)2 < 0 which is always, true. byx + bxy Hence ----> 2 or byx + bxy > 2r rayrax or — +> 2 oxay oyox or — +> 2 oxoy 22 oy + ox or> 2 oxoy or ay2 + a2 > 2 oxoy or oy2 + ox2 - 2 oxoy 0 or (oy - xo)2 > 0 which shows that Regression coefficients, are independent of origin but not of scale. Example (II) Using the following data | | X series | Y series | |---|---|---| | Mean value | 12 | 10 | | Standard deviation | 4 | 3 | Coefficient of correlation X and Y is 0.8 (a)    Form two regression lines. (b)    Predict most likely value of Y when X is = 11 and most likely value of X when Y is = 13. Sol. (a) Let the least square regression line of Y on X be | | yc = a + bx ^xy   rm-   0.8 x 3 b = v 2 =     = j = 0-6 Lx    ox     4 a = Y - ax = 10-0.6 (12) = 2.8 Hence yc = a + bx yc = 2.8 + 0.6 X | |---|---| Let the least square regression line of X on Y be | | Xc = c + dY Lxy    ox  0.8 x 4   16 = r =       =    = 1.07 Ly     oy    3     15 c = X - dY = 12 - 1.07 (10) = 1.3 xc = 1.3 + 1.7 Y | |---|---| (b) Most likely value of Y when X is 11 | | For this we need, regression line of Y on X which is Yc = 2.8 + 0.6 X Putting X = 1l we get Yc 2.8 + 0.6 (11) = 2.8 + 6.6 = 9.4 | |---|---| Most likely value of X when Y is 13. For this prediction, we need regression equation of X on Y which is | | Xc = 1.3 + 1.07y when Y is 13 = 1.3 + 1.07 × (13) = 15.21 | |---|---| Example (iii) Two regression equations are given as | (a) (b) (c) | x - 4y = -13 9Y - X = 53 and ax = 12 Find mean value of X and Y Coefficient of Correlation Standard deviation of Y. x - 4y = -13 ..........I -x +9y = 53 ............. II | |---|---| Adding I and II 5y = 40 y = 8 y = 8 Similarly x = 19 (b)    Coefficient of Correlation Assume that X - 4y = - 13 is regression line of Y on X and aY - X = 53 is regression line of X on Y. From Y on X equation X - 4y = - 13 -4y = - X - 13 4Y = X + 13 Y = 1x + 13 44 1 Hence byx = 4 From X on Y equation 9Y - X = 53 - X = - 94 + 53 X = 9Y -53 Hence bxy = 9 Now bxy byx = 4 × 9 > 1 This shows our assumption is wrong. Therefore we take x - 4y = - 13 as X on Y and 9Y - X = 53 as Y on X 1 bxy = 4 and byx = 9 r = 2 3 r = 667 (c)    Standard deviation of Y. Now bxy = 4 ox r = 4 °y r g = 4g xy ox ■■Gy = r 7 2 12 × 34 = 2 σy = 2 3.6.3 Limitations of the Theory of Linear Correlation (i)    Correlation analysis suffers from serious limitations as a technique for the study of economic relationships r =    Σxi yi Σx2 Σy2 The above formulae for r is applicable only when the relationship between the two variables is linear. However, two variables may be strongly connected with a non linear relationship. The students should also note it well that zero consolation and statistical independence of two variables (x and y) are two different things and they should not treat, them as one and the same, thing. Zero correlation implies zero. covariance of x and y so that r =    Σxiyi    = 0 Σx2 Σy2 Statistical independence of x and y. implies that the probability of xi and yi occurring simultaneously is the simple product of the individual probabilities. P (x and y) + P (x), P(y). (ii)    The ‘second, limitation of the theory is that although the correlation coefficient is a measure of the co variability of variables,’ it does not necessarily imply any functional relationship, between the variables concerned. Correlation theory does not establish, and/or prove any causal relationship between the variables, It seeks to discover if a co variation exists, but it does not suggest that variation in, say, y are ‘caused’ by variation in x, or vice versa knowledge of the value or r, alone, will not enable us’ to predict die value of y from x ....... A high correlation between variables y and x may describe any one of fee following situations: (1)    Variation in x is the cause of variation in y. (2)    Variation in y is the cause of variation in x. (3)    y and x are jointly dependent. (4)    There is ‘another common factor (2); that affects x and y in such way as to show close relation between them. (5)    The correlation between x and y may be due to chance. (6)    Qualitative phenomena cannot be computed. (7)    When the number of observations is large, the calculation, of correlation coefficient becomes laborious and time consuming. Self-Check Exercise-4 Q1. Define a)     Regression Lines. b)     Regression Coefficients Q2. Discuss the properties of Regression Coefficient. Q3. Obtain the regression equation of Y on X by the least square method for the following data: | X | 1 | 2 | 3 | 4 | 5 | |---|---|---|---|---|---| | Y | 9 | 9 | 10 | 12 | 11 | 3.7    Summary In summary the linear correlation coefficient measures the degree to which the points, cluster around a straight line, but it does not give the equation for the line, that it does not assign numerical values to the parameters of the function which is represented by this lines. These parameters are elasticity (or components of elasticity’s) and the knowledge of their numerical value is of particular interest both to entrepreneurs and policy makers. 3.8    Glossary •    Correlation: Correlation is a statistical measure that describes the extent to which two variables are related. It indicates whether increases or decreases in one variable correspond to increases or decreases in another variable. •   Degree of Correlation: The degree of correlation measures the strength and direction of the relationship between two variables. It ranges from -1 (perfect negative correlation) to +1 (perfect positive correlation), with 0 indicating no correlation. •    Scatter Diagram: A scatter diagram is a graphical representation of the relationship between two variables. Each point on the graph represents an observation, with one variable on the x-axis and the other on the y-axis. It helps visualize the correlation between the variables. •   Karl Pearson’s Coefficient of Correlation: This is a method for calculating the correlation coefficient, denoted by ‘r’, which measures the strength and direction of the linear relationship between two variables. It ranges from -1 to +1. •    Regression Lines: Regression lines are lines drawn through a scatter plot of data points that best express the relationship between those points. There are two regression lines: one for predicting the dependent variable from the independent variable and the other for predicting the independent variable from the dependent variable. 3.9    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 3.3. Answer to Q2. Refer to Section 3.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 3.4. Self-Check Exercise-3 Answer to Q1. Refer to Section 3.5.1. Answer to Q2. Refer to Section 3.5.3.1. Self-Check Exercise-4 Answer to Q1. Refer to Section 3.6. Answer to Q2. Refer to Section 3.6.2. Answer to Q3: Y = 8.1 + 0.7 X 3.10    References/Suggested Readings 1.    Draper, N., & Smith, H. (1966). Applied regression analysis. John Wiley & Sons. 2.    Gupta, S. P. (2014). Statistical methods. Sultan Chand & Sons. 3.    Stevenson, W. J. (1978). Business statistics: Concepts and applications. Harper & Row. 4.    Jain, T. R., &Aggrawal, S. C. (2020). Business statistics. V K Publications Pvt. Ltd. 3.11    Terminal Questions Q1. Explain the concept of regression and correlation. Also distinguish between correlation and regression. Q2. What are regression coefficients? Explain the properties of regression coefficients. Q3. Find the regression of Y on X and X on Y by the least square method for the following data: | X | 1 | 2 | 3 | |---|---|---|---| | Y | 2 | 4 | 5 | ***** UNIT-4FITTING OF REGRESSION EQUATION AND STANDARD ERROR OF ESTIMATE STRUCTURE 4.1    Introduction 4.2    Learning Objectives 4.3    Regression Analysis 4.3.1    Descriptive Measures of Regression for Ungrouped Data Self-Check Exercise-1 4.4    Linear Regression 4.4.1    The Standard Error of Estimate Self-Check Exercise-2 4.5    Summary 4.6    Glossary 4.7    Answers to Self-Check Exercise 4.8    References/Suggested Readings 4.9    Terminal Questions 4.1    Introduction Dear students, In the preceding lesson, correlation coefficients and limitation of correlation coefficients were discussed. The present lesson deals with basic statistical methods for studying relationship between two or more variables. In economics we are seldom able to give exact prediction of the value of a variable from a knowledge of the value of other variable. Rather, if a relationship Between two variables X and Y be established, the relationship can tell us the value of y which on the average’ can be expected to’ be associated with a given value of X. Relationships of this kind are’ also referred to as stochastic relationship in contrast with exact relationships. If a stochastic relationship between two or more variables can be expressed by a mathematical equation, so that on the basis of this equation we can estimate the average value of y associated with, the given Xs, the method of analysis is known as regression analysis. In regression analysis we are concerned with statistical, not functional relationships among variables.’ Although regression analysis deals with the dependence of one variable on other variable, it does not necessarily imply causation. In the words of Kendatt and Stuart : “A statistical relationship, however strong and however suggestive, can never establish causal connection : Our ideas of causation must come from outside statistics, ultimately from some theory or other.” 4.2    Learning Objectives 4.3    Regression Analysis In regression analysis, the variable whose average value is being estimated’ is called the dependent variable. The variables on which the estimate is to be based are referred to as the independent or explanatory variables. When a regression relationship contains only one independent variable, it is referred to as a two-variable or simple regression. When more-variables are being used to explain, the behaviour of the dependent variable y, the analysis is known as multiple regression. Where the main objective of regression analysis is to investigate the nature of relationship between two variables x and y, the correlation coefficient measures the strength or degree of linear association between two variables. For example, we may be interested in finding the correlation coefficient between smoking and lung cancer. But in regression analysis, we are not interested in such a measure. Instead, we try to estimate or predict the average value of one variable on the basis of fixed value of other variables. There are some fundamental differences hi the two techniques of regression and correlation which are worth noting. In, regression analysis there is an asymmetry in the way “the dependent and explanatory variables are treated. The dependent or explained variable is assumed to be statistical, random or stochastic, that is, to have a probability distribution. The explanatory, variables on the other hand, are assumed to have fixed values. In correlation, both variables are, assumed to be random, whereas in regression analysis dependent variable is stochastic but the explanatory ‘variables are fixed. Descriptive Measures of Regression for ungrouped data. Suppose we wish to investigate the relation between the height and weight of adult males for some given population. If we plot the pair [x,y] = [height, weight], a diagram like figure 1 ‘will result. Such a diagram, is conventionally called a scatter diagram. x = g(y) Height in cms (x). Note mat for any given height there is a range of observed weights and vice-versa. This variation will be partially due to measurement error but primarily due to variation - between individuals. Thus no unique relationship between actual height and weight can be expected. But it can be observed that average observed weight for a given observed height increases as height increases. The locus of average observed weight for a given observed Height (as height varies) is called a regression curve of weight on height. Let us denote it by-y = f(x) There also exists a regression curve of height on ‘weight similarly defined which we can denote by x = g(y). A pair of random variables such as (height, weight) follows some sort of bivariate probability distribution when we are concerned with the dependence of a random variable y on quantity X, which is a variable but not a random variable an equation that relates y to r is usually called a regression equation. Self-Check Exercise-1 Q1. Define Regression Analysis. 4.4    Linear Regression The simplest, and’ most commonly used relationship between, two variables x and y is that of a straight line. We may write, the linear, first order model as. Yc = a + bx + ε                          .......... (4.1) That is, for a given x, a corresponding observation yc consist of the value a + bx plus an amount ε, the increment, by which an individual y may fall of the regression line. Equation (4.1) is the model of what we believe a, b are palled the parameters of the model a being the intercept of the straight line on the y axis and b its slope. The constant b is known as the regression coefficient of y on x. Clearly we want to determine the values of a and b in such a way that the fitted line is as close as possible to the N plot points, i.e. we want to minimize the overall discrepancy between the plot points and the line, by figure 4.2, consider the typical pair, of observations, denoted by (xi, yi). If the line fitted to the pointy has, intercept a and slope b, then the value of y computed from, the relationship wheh x is Yc = a + bxi and the deviation of the observed value yi from the computed value yc is measured by e1 = Yc - Ye. The deviation ei is called a residual. Eventually, the residuals can be positive or negative as the actual point lies above or below the fitted line. Since the positive negative residuals tend to cancel, out, the summation of these i.e. Σei cannot be used as a measure of overall discrepancy. However, if these are squared and summed, then it is possible to make them as small as possible. Σe12 - = f(a,b) The principle of least square is that a and b are to, be chosen so that Σe2c is a minimum i.e. the sum of squares of the vertical deviations of the observed points from the line is to be a minimum.” This procedure is known as fitting a curve by the method of least squares. An important advantage of the method is mat while being computationally simple, it also yields estimators with certain desirable statistical properties like unbasedness, efficiency and consistency. To determine the values of a and b that satisfy the above requirements, the partial derivatives of the sum with respect to a and b Should both be zero ei = Yi - Yc ei2 = (Yi - Yc) 2 Ee2 = (Yi — Yc) 2 = E (Yi - a - bxi) 2 Differentiating Eei2 with respect to (a Yi = a + -bx) a and b,” we have dEel           zx -7   = - 2 E (y - a - b>xi ) dai dEe2      .. db = - 2 E x (yi -a - bxi ) = - 2E (xy - ax - bx2) ,                 dEe2dEe For Ee2 to be minimum----and —— must both equal to zero, which they will do when dadb E(y- a - b)x ) = 0i and i.e. when E(xy - ax - bx2) = 0 } Ey = na + bEx Exy = ax + bEx2 These two equations are known as the normal equations for determining a and b. If we determine the numerical values of a and b such that these equations hold, the least square equation Yc = a + bx will satisfy two algebraic properties. First the deviation of observations about the regression line sum to zero, i.e. Ee = 0; secondly the sum of squared deviations from the line, i.e. Ee2 is minimum. Solving the normal equations for a and b, we have from the first” equation: a = y - bx and from the second Exy = asy + bex2 = (y - bx ) ex + bsx2 y Ex + bx ex2 + bsx2 Exy - nxy = n - nb + bsx2 vx Ex n vEx = nx Exy - nxy = b [sx2 -nx2 ] _ Exy - nxy b =       ^2" Ex - nx The expression for b can be further simplified by putting x = x - x and y = y - y i.e. x and y are the deviations from their respective means. Now X = x + x , and y = y + y Sxy = S(x + x ) (y + J ) = Sxy + x sx + y sy = N xy = Sxy + N xy            ( '•' Sx = 0 and Sy = 0) = Sxy = n Xx = Sxy and similarly Sx2 = S(n + x ) 2 = Sx2 + nx2 + 2 x sx = Sx2 + nx2 Sx2 = Sx2 - nx2 ,   Sxy Hence b = b -   ■ Sx The normal equations can also ‘be solved with the help of Cramer’s rule. na + bSx = Sy aSx + bSx2 = Sxy The solution to these equations is easily written as Follows : Sy Sx Sxy - Sx2 = SySx2 - Sx Sxy n Sx ~ nSx2 - (Sx)2 Sy Sx2 n Sy Sx Sxy nSxy -SxSy n Sx    nSx2 - (Sx)2 Regression equation is a measure of the average relationship between x and y such that for a given x, yc is the value of y which we would on the average expect to be associated with that x. The regression coefficient b measures the change in y which occurs on average per unit change in + Yc is of course, expressed in the same units as y. It will be noted that when x = x . yc = y so that the regression line passes through the means of Xs and YS . An Example: Data on the annual sales of a company in lakhs of rupees over the past eleven years is shown in the table below. Determine a suitable straight line regression model, y = a + bx + s for the data in the table: | Year | Annual sales in Lakhs of Rupees | |---|---| | 1988 | 1 | | 1989 | 5 | | 1990 | 4 | | 1991 | 7 | | 1992 | 10 | | 1993 | 8 | | 1994 | 9 | | 1995 | 13 | | 1996 | 14 | | 1997 | 13 | | 1998 | 18 | Solution : The independent variable in this problem is the year whereas the response variable is the annual sales. We see that to estimate the parameter b we require the four summations 2x1 Sy1, 2x12, and Sx. y.. Thus, calculations can be organized as shown below where the totals of four columns yield the four desired summations : | Xi | Y i | X12 | X1Y1 | Y12 | |---|---|---|---|---| | 1 | 1 | 1 | 1 | 1 | | 2 | 5 | 4 | 10 | 25 | | 3 | 4 | 9 | 12 | 16 | | 4 | 7 | 16 | 28 | 49 | | 5 | 10 | 26 | 50 | 100 | | 6 | 8 | 36 | 48 | 64 | | 7 | 9 | 49 | 63 | 81 | | 8 | 13 | 64 | 104 | 169 | | 9 | 14 | 81 | 126 | 196 | | 10 | 13 | 100 | 130 | 169 | | 11 | 18 | 121 | 198 | 324 | | Sxi = 66 | Syi =102 | 2x12 506 | Sxi yi = 770 | Syi2 = 1194 | we find feat n = 11 Sx1 = 66 x = 6 Sy1 = 102 y = 9.2727 Sx12 = 506 ; Sx1 y1 = 770 b= nΣxy - Σx Σy nΣx2 - (Σx)2 (11×770)-(66×102) (11×506)- (66)2 8470-6732 5566-4356 1738 1210 1.4363 b = 1.44 a = Y - bX = 9.27 - 1.44 × 6 = 9.27 - 8.64 = 0.63 The fitted equation is thus Regression equation of y on x y - y = σy r (x-x) σ x y = 0.63 + 1.44x Now that the model is completely specified we can obtain the predicted values yi and the errors or residuals yi - yˆ corresponding to the eleven observations. These are shown in table below. | X i | Y i | yˆ 1 | E = Y1 -yˆ | e12 | |---|---|---|---|---| | 1 | 1 | 2.07 | -1.07 | 1.15 | | 2 | 5 | 3.51 | 1.49 | 2.20 | | 3 | 4 | 4.95 | -0.95 | 0.90 | | 4 | 7 | 6.39 | 0.61 | 0.37 | | 5 | 10 | 7.83 | 2.17 | 4.71 | | 6 | 8 | 9.27 | -1.27 | 1.61 | | 7 | 9 | 10.71 | -1.71 | 2.92 | | 8 | 13 | 12.15 | 0.85 | 0.72 | | 9. | 14 | 13.59 | 0.41 | 0.16 | | 10 | 13 | 15.03 | -2.03 | 4.12 | | 11 | 18 | 16.47 | 1.53 | 2.34 | 4.4.1 The standard Error of Estimate. It has been shown above feat with the help of the regression equation it is possible for us to estimate the value of y for any given value of x. Thus when x is 8, the sales is estimated to be 12.15 lakhs rupees. Now misestimated value of y is less than its observed value (Rs. 13 lakhs). This means that the regression line is not a perfect fit and the entire points on the scatter diagram do not follow on this line. In other words, it may be said that regression, line will not enable 92 us to make estimates equal to the observed value, of the sales. It may thus be said that the estimates will be in error. This error is due to the fact that variations in Y may not be due exclusively to variations in X. There are other forces as well which influence the size of y. In order to know as to how far the regression equation has been able to explain the variation in y, it is necessary to measure the scatter of fee point around the regression line. If all the points on the scatter diagram fall, on the regression line, it means that regression equation anables us to make absolutely correct estimates of the values of y. In other words we can say that the variations in y are fully explained by variations in x and there is no error in the estimates. The scatter of the points from the regression line is called the standard error of estimating y. It is obtained commonly by the formula: Sy S( y - yc )2 Se2 \ n-2 n= =k Where    Sy standard error of estimate y = the observed values of y yc = the estimated values of y N - 2 = degrees of freedom (Since two parameters a and b have been estimated from n i.e. 11 observations) Sy 21.25 9 = 1.53 It will be observed that the, method of computing Sy is similar, to that of calculating o or with, the only difference that whereas in calculating o deviations are measures from the mean, in the case of Sy the y are measured from the regression, line, (the estimated y) for the purpose of computing the value of standard estimate of y, the formula can further be simplified so that we, avoid the need, for calculating individual ei’s the deviations of the observed values from the regression line. e = (y — a — bxi) Se2 = S (yi - <2 - bxi)2 •••x=y-bX By substituting for α we obtain = S (y, - (y - bx) - bxi) = S [( y/ - y) - b( xi- x )]2 = s[(y/- y)2 + b2(xi - x)2 - 2b (y - y)(xi - x)] y- y )2 + b2( Xi- x )2 - 2b (Xi- x )2 ] r _ (xi- x)(yi-y) Since b ,    -\2 (xi - x ) = S[( y/ - y )2 + b2( x- x )2 ] = Sy (S iN b2Sxi2 + b 2 (-X)2 N = S, •lSy' iN /2 S.x2 (SX) ˆ 2 N Sy Sy Sy 2(Sy)2 iN ˆ b2 Ei® ' N 2 n - 2 1194(102)2 11 - (1.4363)2 506 I 9 - nr J 21.25 = 1.53 9 which is much easier, to compute that taking value of ei by subtracting estimated; yS from observed yS and then squaring and summing. The various computations outlined in the case of straight line regression equation will now be illustrated to compute the standard error of fee intercept and slope. Most commonly the intercept and slope. Most, commonly the estimators of var ( aˆ ) and var (bˆ ), are denoted by S2aˆ and S2b respectively. The formulas for calculating var ( aˆ ) and var ( bˆ ) are : S 2a = S 2b = 1 — + n x2 Sxi2 - a2 (Sx)2 N a 2 =   -a .   a a S(x1 - x )2 a2 Sxi2 - (^ N 2 2(x1 — x) a2 Sx2 Where a2 Se2 n - 2 Standard Error of the intercept 506 S2 a = 11   a2 110 a . 46 =---a' 110 • 2 = .418 x 2.36 = 0.98 estimate of standard error ( aˆ ) = esv(aˆ) = 0.99 Then 95% confidence limit is for aˆ are a . t.(9, 0.975) s^2] 0        [n S(xi- x)2]12 = 0.63 ± (2.262) (0.98) = 0.63 ± 2.2167, that is -1.5867 and 2.8467 Standard error of the slops bˆ . S2(b) a Xx2 - (W A 1 2 a Xx-2 2.36 110 = 0.0215 estimate of standard error ( bˆ ) = est(S2bˆ) = 0.1465. a I Suppose a = 0.05, so that I n 2, 1 y I = (9, 0.975) = 2.262 from the table of the distribution. Then 95% confidence limits for bˆ are + (9, 0.975)S 1 Ex.^ — ^ 1 N = 1.44 ± (2.262) (0.1465) = 1.44 ± 0.3314 that is 1.7714 and 1.1086 Self-Check Exercise-2 Q1. Define Linear Regression. Q2. What do you mean by the standard error of estimate? Q3. Obtain the regression equation from the data given below: | X | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | |---|---|---|---|---|---|---|---|---|---| | Y | 9 | 8 | 10 | 12 | 11 | 13 | 14 | 16 | 15 | Plot the regression equation on a graph paper and determine X and Y. 4.5    Summary In this unit, we studied regression analysis, focusing on fitting regression equations to explore relationships between variables. We examined descriptive measures for ungrouped data, and delved into linear regression, emphasizing the calculation and significance of the standard error of estimate in predictive models. 4.6    Glossary •    Regression Analysis: A method to find relationships between a dependent variable and one or more independent variables. •    Linear Regression: A technique that models the relationship between variables using a straight-line equation. •    Standard Error of Estimate: A measure of how much the observed values deviate from the predicted values in a regression. •    Coefficient of Determination (R²): A measure indicating how well the independent variables explain the variance in the dependent variable, ranging from 0 to 1. 4.7    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 4.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 4.4. Answer to Q2. Refer to Section 4.4.1. Answer to Q3: Y = 0.95X + 7.25 ; X = 0.95Y + 7.25 4.8    References/Suggested Readings 1.    Draper, N., & Smith, H. (1966). Applied regression analysis. John Wiley & Sons. 2.    Jain, T. R., &Aggrawal, S. C. (2020). Business statistics. V K Publications Pvt. Ltd. 4.9    Terminal Questions Q1. What is regression line? Why are there, in general, two regression lines? Under what conditions can there be only one regression line? Q2. Explain the meaning of standard error of estimate. Q3. Suppose that you are interested in using past expenditure on research and development by a firm to predict current expenditure on R & D. You got the following data by taking a random sample of firms where X is the amount on R & D ( in lakhs of rupees) 5 years ago and Y is the amount spent on R & D(in lakhs of rupees) in the current year. | X | 30 | 50 | 20 | 80 | 10 | 20 | 20 | 40 | |---|---|---|---|---|---|---|---|---| | Y | 50 | 80 | 30 | 110 | 20 | 20 | 40 | 50 | (1) Find regression equation of Y on X. ***** UNIT-5THE GENERAL LINEAR REGRESSION MODEL :MATRIX FORMULATION & SOLUTION-1 STRUCTURE 5.1    Introduction 5.2    Learning Objectives 5.3    Relation Between Three Variables Self-Check Exercise-1 5.4    Elements of Matrix Algebra 5.4.1    Set of Rules and Definitions 5.4.2    Multiplication of Matrices Self-Check Exercise-2 5.5    Summary 5.6    Glossary 5.7    Answers to Self-Check Exercise 5.8    References/Suggested Readings 5.9    Terminal Questions 5.1    Introduction Dear Student, In the simple linear regression model, the objectives of the analysis were to determine the degree of relationship between two variables and to predict the behaviour of the dependent variable on the basis of an independent variable. It is generally seen that the dependent variable is ‘related ‘not only to one independent variable but to a number of independent variables all operating at the same time. For example the quantitative demanded for a given, commodity (y) depends on its price (x1) and on consumer’s income (X2) and price of related goods (X3). In the present lesson we shall extend the simple linear regression model to relationships with two explanatory variables and later on we shall develop some practical rules for the derivation of the normal equations for models including any number of variables. In section I we shall examine the model with two explanatory variables. 5.2    Learning Objectives After going through this unit, you will be able to •    Understand the relationship between variables and how it applies to regression models. •    Master essential matrix algebra operations, particularly multiplication, crucial for solving regression equations. •    Learn matrix techniques to formulate and solve general linear regression models effectively. 5.3 Relations between three variables: We shall illustrate the three variable model with an example from the theory of demand. It is well known that quantity demanded for a given commodity is a function of its price and consumers income x1 and x2 respectively. y = f(x1, x2) We assume that there is a linear relationship, between y, x1 and x2. Yi = β0 + β1 x1i + B2 × 2i, (= l,2,3   n) The, above relationship, is an exact relationship showing that the n variations, in the quantity demanded are fully explained by charges in price, is and income However, if the gathered information is plotted one diagram of it. It will be observed that, some will lie on it, but others will lie above or (0) below it and all of them will of it lie on a plane. This scatter is due to error (s) account by introducing a random variable u. in the function, which thus becomes stochastic. A sample on n households would give related observations. Yi = B0+ B1 xn + B2 x2i+ 4i systematic component random component On a priori grounds, the coefficient βˆ , have a negative sign, showing an inverse relationship between quantity, demanded and price, while βˆ is expected to have, a positive sigh as quantity demanded and income are positively correlated. Where x1i = 1 for all i; that is y can be regarded as a linear function of die X’s, the sample values of the first x variable always being a set of units. Our basic objective is to give the student a firm grasp and understanding, of the basic concepts required in the analysis of a relationship between three variables, so that extension to the general case is facitiated. Let the regression equation of y on x1 and x2 is of the form y = β0 β11 x1 + β21 x2 Yi = B0 + B1 x1+ B2 X2 AAA Where β0 , β1 , β2 are estimates of the true parameters β0, β1 and β2 of the demand relationship. As before, the estimates will be obtained by minimizing the sum of squared residuals. nn      n I e2=2 ( yi- y i)2=2 (yi- Po- Pixii- P2 x2i I i=1 i A necessary, condition for this expression to assume a minimum ‘value’ is that its partial derivatives with respect to β0 , β1 and β2 be equal to zero : ÎA     A       AV yi - P0 - P1 x1i - P2x2 )= 0 dPi ÎA     A yi - Po - Pi xi A dPi - ˆ β2x2 = 0 ÎA     A yi - Po - Pi xi A dP2 - ˆ β2x2 = 0 Performing the partial differentiations we get the following system of three normal equations, in the three unknown parameters β0 - β1 and β2 . ^yi = n Po+ Pi ^xu + P2 Sx2i 2xii yi = Po 2xii + P1 S x + P2 Sxu x2i .... (5.2) Sx2i yi = Po Sx2i + P1 Sx1i x2i + P2 Sx2i This system can be solved by using, determinants or by the use of normal equations. The above normal equations are expressed in terms of deviations from means. In this case the first equation disappears as x1i Sx2i and 2yi are not equal to zero (because the sum of deviation from the inspected mean is equal to zero): 2(x1— x ) = 2x1 = 0, the remaining two equations are reduced to the form: S x^yi = Pi^xii- + p2sxii x2i S x^yi = P, 2x..x,+ P,-2x,; 2i             1     1i 2i 2i      2i 5.3 The coefficient of correlation between y and x1 from mean is given by Syxi r= yx1   Nay ax or           Nay ox1 ryx1 = Syx1 Similarly     Nay ox2 ry x2 = Syx2 and           Nex1 ex2rx1 x2 = Sx1x2 The standard deviation (from mean) of x1 series ax1 Z x2 N or Nax12 = Sx12 and Nex22 = 2x22 Substituting the computed values Ney ex.r = R, N ex2 + R, Nex. ex.r. x, 5.4 1 yx1        1           1 2          1 2 x1 2 Ney ex.r = R, N ex. ex.r . x, + R, N ex 25.5 2 yx2 1          1     2 x1 2       2 Cancelling N ex1 from equation (5.4) and Nex2 from equation (5.5) we get. ey r ,= B, ex.+ R. ex.r. x.5.6 yx1 1      1      2     2 x1 2 zx ey r = R. ex.r. x+fk ex,                              5.7 yx2 1      1 x1 2      2 Dividing equation (5.7) by rx1 x2 and subtracting it from 5.6, we get f ayryX1 - ay ryx2 A ˆ i x rx1x 2 / = A ay2 rx1x2 x f 64 ryx1 x f ay ryx1 x - - ryx2 A ˆ i rx1x 2 / = £ 6x2 rx1x 2 x ryx2 A ˆ i rx1x 2 / = £ ax2 rx1 x 2 x ax 2 A rx1x 2 / ax2 A rx1x 2 / - 1 A rx1x 2 / ryxl rxlx2 rx2 ˆ A2 = — ox2 V rx1x2 2 rx1x2 r rx1x2 7 o ox2 °y ox2 ryxl rxlx2 ryx2 X rxlx 2 <     oxlx 2 2 rx1x2 17 r ryxï rr yx1 x1x2 A V 2 rx1x 2     7 o Ai = ox1 r ryxï rr ryx2 rx1x2 \ 2 V 1   rxlx 2     7 or by solving 1.4 and 1.5 algebraically we get ~ _ Lyxl Lx2 - Lyx2 Lxlx2 l     Lx2 Lx2 - ( Lv,x 2 )2 Which may be reduced to the following expression in forms of zero coefficient. A A = A Ai = ryxd ryx2 rxlx2 V   1 “ rx1x22 rx2 7 r , r, yx1  x1x2 A 2 V l rxl x2    7 oy ox1 oy ox2 Self-Check Exercise-1 Q1. Discuss the relationship between three variables model with an example from the theory of demand. 5.4 Elements of Matrix Algebra It is excessively tedious and complicated to build up a general case of K variables in a stepwise fashion. Fortunately, by the use of matrix algebra compact and powerful way of the treating, the general case be obtained by the use of matrix algebra. A matrix may be defined as on system of m n numbers arranged in the form of an ordered set of m row s, each row consisting of an ordered set of n numbers (or m, n numbers arranged in the form m rows and n columns). This matrix is called m×n matrix (first, letter of m × n always denotes rows and second denotes columns). or A matrix is defined as a rectangular array of elements arranged in rows or column as in A = | a11 | a12 ..... | .... a1 j.... | ..... a1n | |---|---|---|---| | a21 | a22 ..... | .... a2 j.... | ..... a1n | | ai j | a ....... i | .... ai j...... | ..... ain | | am1 | am2 .... | ... amj .... | ... amn | m × n If it has mn elements arranged in m rows, and n columns, it is said to be of order m by n, which is often written as m × n. The elements in the ith row and jth column is represented by aij. The matrix may be indicated more concisely by A = [aij] A matric of order I’n contains only a single row of elements and is commonly referred as a row vector, for example, b = [b1b2 ........bn] While a number of order mx1 is a column vector, C = c21 cm To economize space, comlumn vectors, may be written in horizontal position but enclosed in braces, c = [c1 c2 .........cm] 5.4.1    Set of Rules and Definitions (1)    Two matrices A and B are said to be equal when they are of the same order and aij = bij for all i, j: that is matrices are equal element by element. (2)    If A and B are of the same order, then the define A + B to be a new matrix c of the same order in which cij = aij + bij for all i.j. A 23 3 then - 4 5 and p = :4 1 1 1 C = A + B = 0 4 6 (3)    It r is a scalar, then we define scalar multiplication such that r, A = [r, aij] that is each element of A is multiplied by r. For example if A 2 -4 3 5 and r = - 2 then - 4 - 6 + 8 -10 If follows from the raters for addition and scalar multiplication that A-B = [aij -bij) 5.4.2    Multiplication of Matrices: If A = [aij] and B = [bij] are two matrices, then their product AB is defined to be a matrix of order m × p (if A is of the order m × n and B is of the order n × p). Whose ij the element is n Cu = Xaik bkj then k=1 that is, the ijth element in the product matrtrix is by multiplying the elements of the ith row of the first matrix by the corresponding elements of the jth column of the second matrix and summing over a U terms. Example A = a11 a21 a12 a22 a13 a23 J 2 X 3 B = b11 b21 b31 b12 b22 b32 J 3 X 2 Then AB is 2 × 2 matrix a11 b11 + a12 b21 AB = a21 b11 + a22 b21 + a13 b13 a11 b12 a12 a22 + a13 b32 + a23 b31 a21 b12 a22 a22 + a23 b32 While BA is a 3 × 3 matrix b11 all + b12 a21 b11 ai1 + b12 a21 b11 a13 + b12 a23 AB =  b21 a11 + b22 a21 b21 a12 + b22 a22 b21 a13 + b22 a23 b31 a11 + b32 a21 b31 a12 + b32 K12 b31 a13 + b32 a23 I.    It must be noticed that the commutative law of addition holds. A and B must, of course, be of the same order, and the result follows directly from die definition of the addition of matrices. 2 31 1 0 -1 - 5 0 + - 26 0 -12 3 + - 2   6- 5 0 22 - 7 6 II.    AB ^ BA except for rather special square Matrices : i.e., the commutative law of multiplication does not, in general hold. It the matrices are of order m × n and n × m, then both products will exist, but they will be of different order and hence cannot be equal. It both is square matrices of the same order then froth products will exist and will be of the same order but not necessarily equal as the following examples show. Examples A = 2 1 11           3 1        - 2 0 2 AB = 6 + 1 0 + 2 3 + 1 0 + 2 72 42 6 + 0 3 + 0IP 6 3' 2 + 2 1 + 2 = 4 3 21 1 - 1 Whereas if A = 11 B = -1 2 AB = 1 0 0 1 = BA. III.    (A+B) + C = A + (B + C) : that is the associative law of addition holds. Since addition of matrixes is simply achieved by the addition of corresponding elements and since it does not matter in which order elements are added together, the associative law holds. IV.    Matrix Multiplication is associative, i.e. AB (C) = A(B). We can first form Ab and then post-multiply by C. or multiply A and BC, of course BC will have to be found first e.g. if A is m × n and if C in p × q then conformability requires that. The distributive law holds for matrix multiplication i.e. A(B+C) = AB + AC (pre-multiply by A) (B+C) A = BA + CA (post-multiply by A) VI.    1 (A+B) = XA + IB and (1 + p) = XA + pA: that is the distributive law of scalar multiplication. These are the important results for the manipulation of matrices and the students should learn them by working out numerical examples. A unit of Identity matrix is “defined by | | 10 0 0 0 10 0 | |---|---| | In= | 0 0 10 0 0 0 1 | a scalar matrix has a common scalar element in the principal diagnol and p nos everywhere, else and is defined by A 0 ... 0 21 = 0 2 0 0 0 0 2 A diagnal matrix is defined in which scalar elements are not necessarily equal in the principal, diagnonal and zeros in off diagnonal positions, i.e., A = [ aij] i, j = 1,2 ....... n aij = 0, i * j. That is a,, A = 0 0 a22 0 0 a00 0 0 diag [a11 a22 a33] A scalar matrix is thus a special form of a diagnal matrix. Transposition, the transpose of A is defined to be the matrix obtained from A by inter changing, rows and columns that is, the first row of A becomes the first column of transpose, the second row of A becomes the second column of the transpose, and in general” the ijth element inthe transpose is the ijth element of the original matrix. A = [aij] m × n or A’ = [aij] n × m For example if | | | | | | an | |---|---|---|---|---|---| | A= | ■all | a12 | a13 ’ | Than A' = | a12 | | | _a 21 | a22 | a 23 _ | 2×3 | _ a13 | a21 a22 a23 3×2 A = a11 a21 a12 a22 a13 a23 Than A’ = 2×3 a11 a12 a13 a21 a22 a23 3×2 A’ A = and A’ A a11×a11 - a12 × a12 + a13 a13 a11× a21 - a12 × a22 + a13 a23 a21 a11 - a22 × a12 + a23 a13 a21 a21 + a22 × a22 + a23 a23 22 a11 a21 a a12 a11 + a^ a2i a 111 a12 + a21 a22 22 12   a22 a11 a13 + a21 a23 a12 a11 + a22 a21 a13 au + a23 a2i a13 a12 + a23 a22 a13   a23 3x3 If x is a column vector of n elements x£ is then a row vector of m elements and n x2 xx 1 i=1 and x 2 1 x2x1 x1x2 x22 1 XX'=   1 x1xn x2xn 1 1 1 1 1 xnx1 xnx2 1 2 n Some Important properties of transpose of matrix. (1)    (a')' = A. (2)    (Ak)' = (A') where k is a positive integer. (3)    (A + B)' = A’ + B’ (4)    (AB)' = B' A' (5)    (ABC)' = C' B' A' | | ’3 | 0 | 1’ | | ’1 | 2 | 0 | |---|---|---|---|---|---|---|---| | Example: If A = | 0 | 4 | 2 | and B = | 0 | 3 | 4 | | | L1 | 2 | °J | | 5 | 0 | 1 | Prove that    (i) (A+B)' = A' + B' (ii) (AB) = B' A' (2) | | ’3 | 0 | 1’ | | ’1 | 2 | 0’ | | ’4 | 2 | 1’ | |---|---|---|---|---|---|---|---|---|---|---|---| | (1) A + B = | 0 | 4 | 2 | + | 0 | 3 | 4 | = | 0 | 7 | 6 | | | _1 | 2 | 0_ | | _5 | 0 | 1_ | | _6 | 2 | 1_ | | | ’4 | 0 | 6’ | | |---|---|---|---|---| | (A + B)' = | 2 | 7 | 2 | ......(1) | | | L1 | 6 | 1J | | | | ’3 | 0 | 1’ | | ’1 | 0 | 5’ | |---|---|---|---|---|---|---|---| | Also A' = | 0 | 4 | 2 | and B' = | 2 | 3 | 0 | | | _1 | 2 | 0. | | .0 | 4 | 1. | A A' + B' = | 4 | 0 | 6 | |---|---|---| | 2 | 7 | 2 | | 1 | 6 | 1 | from (1) and (2), we get (A+B)' = A' + B' (AB) = 301 042 120 12 03 50 0 4 1 r 3x1+ 0 X 0 + 1x5 3x2 + 0 x 3 + 1x0    3x0 + 0 x 4 + 1x11 0x1+ 4 x 0 + 2x5 0x1 + 4 x 3 + 2x0    0x0 + 4 x 4 + 2x1 1x1+ 2 x 0 + 0x 5 1x2 + 2 x 3 + 0x0    1x0 + 2 x 4 + 0x1 861 10 12 18 188 8 10 /.(AB)' = B 'A' = 6 1 12 18 1 8 8 (I) 10 2 0 3 4 5113 0 1 0 0 4 2 112 0 8 10 1 6 12 8 (II) 1 18 1 From (1) and (2), we get (AB)' = B'A' Self-Check Exercise-2 Q1. Discuss the elements of Matrix Algebra. Q2. Discuss the set of rules of matrix algebra. Q3. Consider two matrices A and B where A is an m x n matrix and B is an n x p matrix. 1)    Explain the condition for the multiplication of two matrices A and B. 2)    If C is the resulting matrix from the multiplication AB, what are the dimensions of C? 5.5    Summary In this unit, we explored the general linear regression model using matrix algebra. We learned how to understand relationships among three variables and applied matrix operations to analyze these relationships. This unit aimed to provide a clear foundation for using matrices to solve and understand regression models. 5.6    Glossary •    Intercept and Slope: In a simple linear regression equation y=a+ bx. •    Intercept (a): The value of y when x is zero. For example, if the regression equation isy=5+ 2x, the intercept a is 5, indicating that whenx is 0, y is 5. •    Slope (b): The change in y for a one-unit change in x. In the equation y=5 + 2x, the slope b is 2, meaning that for each additional unit of x, y increases by 2 units. •   Matrix Algebra: A branch of mathematics involving the manipulation and use of matrices to solve linear equations and perform various calculations. •    Matrix Multiplication: A specific operation in matrix algebra where two matrices are combined to produce a new matrix, essential for solving regression equations. 5.6    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 5.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 5.4. Answer to Q2. Refer to Section 5.4.1. Answer to Q3. Refer to Section 5.4.2. 5.7    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Freund, J. E., & Williams, F. J. (1959). Modern business statistics. London: Pitman and Sons. 5.8    Terminal Questions Q1. Discuss the elements of matrix algebra. What are the set of rules of matrix algebra? Q2. Let A be a 2×3 matrix and B be a 3×2 matrix. 1.   Show that the product AB is defined and find its dimensions. 2.    Is the product BA defined? If so, what are its dimensions? 3.    Discuss whether the matrix multiplication AB is commutative, i.e., whether AB =BA holds in general for matrices. ***** UNIT-6THE GENERAL LINEAR REGRESSION MODEL : MATRIX FORMULATION & SOLUTION-II STRUCTURE 6.1    Introduction 6.2    Learning Objectives 6.3    The General Linear Regression Model 6.3.1    Model with One Explanatory Variable 6.3.2    Model with Two Explanatory Variables 6.3.3    Model with Three Explanatory Variables Self-Check Exercise-1 6.4    Goodness of Fit (R2) Self-Check Exercise-2 6.5    Summary 6.6    Glossary 6.7    Answers to Self-Check Exercise 6.8    References/Suggested Readings 6.9    Terminal Questions 6.1    Introduction In this unit, we will delve into the general linear regression model using matrix formulation and solutions. You’ll learn how to build and understand regression models with one, two, and three explanatory variables. We’ll also cover how to assess the goodness of fit using R2. By the end of this unit, you will be equipped with the knowledge to apply these concepts to real-world data, enhancing your analytical skills and understanding of statistical modelling. 6.2    Learning Objectives After completing this unit, you will be able to •    Understand and construct general linear regression models with one, two, and three explanatory variables. •    Apply matrix algebra to formulate and solve regression models effectively. •    Evaluate the goodness of fit for regression models using R2R^2R2 and interpret the results. 6.3    The General Linear Regression Model The discussion so far was limited to the regression models containing one or two independent/ explanatory variables. Now we shall generalize the model assuming that it contains K explanatory variables. The general linear equation will be of the form. Y1 = β0 + β1 X1i + β2 X2i +   βik Xki + Ui(i = 1   n)               ....... (2.1) We need some assumption about the error term u. Assumption I : µi is a random real variable and has ~ normal distribution. Assumption II : The mean value of g for each xi is zero or E (^i) = 0. Assumption III : The variance of the distribance term is constant in each period : E(gi2) = eg2 (eg2 is a constant) Assumption on IV : The convariance of gi and g. is equal to zero. E(gi gj) = 0 for i / j. Assumption V: Every disturbance term, is independent of the expiatory variables: E(x2i ui) = X2 E (ui) = 0 E(x3i ui) = X3 E (ui) = 0 There are k parameters to be estimated (k = k + 1) clearly the system of normal equations will consist of k equation s, in which B0, B1, B2   and Bk are the unknown parameters, and the known term s will be the sum of squares and the sum of the products of all the variables in the structural equation. 6.3.1.    Model with one explanatory variable Structural form Y= P0 + 01 X1 + g A      A estimated from Y= β0 + β1 X1 + e '                 A        A                     ' Normal equations SY = n ßo + bi LX SX1Y = ß0 SX1 + ß1 SX2 6.3.2.    Model with two explanatory variables structural form y = β0 + β1 X1 + β2 X2 + u A      A          A estimated form y = β0 + β1X1 + β2X2 + e Normal equations Sy = b p0 + 01 SX1 + 02 SX2 SYX2 +00 SX2 + 01 SX1 X2 + 02 SX22 X2 = 00 SX2 + 01 SX1X2 + 02 SX22 6.3.3.    Model with k explanatory variable structural form Y 2= 00 + 01 X1i+ 02 X2i+     bk X1+ gi (i = 1.2 ............ n) Since Subscript represents the ith observation, we shall have n number of equations with n number of observations on each variable : Y 1 = 00 + 01 + X1+ 02 + X1 + 03 + X31 +Bk Xk1 + g1 Y 2= 00 + 01 + X21+ 02 + x22+ 0X32 +...... 0k Xk2 + g2 Y n= 00 + 01 + X1n+ 02 + X2n+ 03X3n +...... 0k Xkn + gn These equations are put in matric form y = Sb + u where: Y1 Y2 X11 X21 X12  X22 Xk1 Xk2 y= x = β = Yn & 01 X X2n X1n Xkn nx (k +1) and U = u1 u2 L 0k J U n Taking (1.1) and assumptions together, we now apply the least square principle to estimate the parameters of 1.l. p = K pt.......Pk} denote a colum vector of estimate of b. Then wewrite y = Xβˆ + e ........ 2.2 Where denotes the column vector of n residuals (y - Xp). You should carefully notice the basic difference between (2.1) and (2.2). In the former the unknown coefficients β and the unknown disturbances u appear, white in the latter, we have some set of estimates βˆ and the corresponding set of residual e. From 1.2 the sum of squared residual, is n Ë e2 = e' e i=1 We have to minimize : ^ n e 2 1 i=1 2222  2 Se, = Se1 = e1 + e2 + e3     en = [e1 e2 e3  en] IX e 1 e2 ..........2.3 enxl Le2 = e' e = (y - Xpy )(y - Xp) = (y - X'p' )(y - Xp) = Y' Y - p'X’ Y' - p' Y'X' + X'Xp' A             A          A = Y' Y - 2p' X’ Y' - p' X'X' p                        .........2.4 Which follows from neting that p' X'Y is a scalar and thus equal to its transpose Y'X' p. To find the value of β which minimizes the sum of squared residuals we differentiate (2-4) . (e'e) = 2X'Y + 2XX'p dβ equating to zero gives XX'p = 2X'Y p = (X'X)-1 X'Y                         ........2.5 This is the fundamental result for me least squares estimators. As an illustration of this result consider the two variable case. Here n  LX          LX X' X = LX LX2 and X' Y LXY So that writing (2.5) in the attentive form (X' X)p = X' Y and substituting gives. _          A A _ LY = a p0p1 LX LXY = p0 LX p1 LX2 which are two normal equations already derived. For the three variable case (2.5) gives LY = b po + pi LX2 + p2 LX3                        .... 2.6 LX2Y = p0 LX2 + p1 L X2 + p2 LX2X3               .... 2.7 LX3Y = p0 LX3 + p1 LX2X3+ p2 L X32              .....2.8 and it is clear from the symmetry how these equations can be built up for higher order cases. Solving (2.7) and (2.8) for βˆ1 and βˆ2 . ZK ß1= Sxi y Sx2 - Sxix2 Sx2 y Sx2 Sx 2 2 - (Sxix2)2 2.9 ZK ß2 = Sx2 y Sx32 - Sx2 x3 Sx3 y Sx2 Sx 32 - (Sx2 x3 ) 2.10 Self-Check Exercise-1 Q1. Define General Linear Regression Model. Q2. Discuss the assumptions of error term. Q3. Explain the following terms a)    Regression model with one explanatory variable b)    Regression model with two explanatory variables c)    Regression model with k explanatory variables 6.4 Goodness of Fir (R2) In the previous section, we discussed the general linear regression model with one explanatory variable, two explanatory variables and k explanatory variables. In this section we now discuss the Goodness of Fit of the fitted regression line, means how well the sample regression line fits the data. Sei2 (2.11) R2 = 1 - -yi2 (In the case of one explanatory variable) In the present model of two explanatory variables Se,2 _ S(y - ß,X,m - ZK β2X2 = Sei (y- ßiXii - ß2X2i) = Yeiyi - ßi Seixii- ß2 Seix2i = Seiyi [ v Sexii = Sex2i = 0] = Sy (y,- ßi xii -ß2 x2i) Sei2 = Sy2 - ßiSxiiyi - ß2Sx2iyi Sy = ßi SxH y i + ß2Sx 2i y i + Sei -y2    = Total sum of squares or total Dividing both sides by Syc2 ˆ ˆ Pi Ex^y i + p2Sx 2iy, Explain sum of squares (explained Variations) + Sei2 Residual sum of square (unexplained Variation) ˆ 2 . _ ßi Sxiiyi + ß2Sx2iyi+ sei Sy2 + Sei Sy2 ' _ Sum of Squares explained by X1 and X2 ^ V Total sum of squares 7 For estimation of the standard errors of ß1 and ß2 we need an estimate of an2. Sei2 R2 = 1 - syi2 or Sei2 = Syi2 (1-R2) Var (ßi ) = on2 2x2 Sx^Sx^.- (Sx1i Sx 2i)2 (2.14) Var ^2 )= y 2y 2°“2 S'2 V  ,2 . Sxu Sx2i - (Sx1i Sx 2i) .... (2.15) These variations can be expressed in terms, of simple correlation coefficients. Var (ßi )= on2 2 (Sxii SX2i) Sxii    Yy2 Sx2i 0112 | Sxi2i | ’ (Sx1i Sx 2i)2 ’ | |---|---| | l-Sx’-Sx*] | ˆ i.e var β1 on2 Sxii (1 - ri2) ˆ Similarly var β2 2 on ou2 Sxn (1- r2) Se2 1 (k being the n - k .... (2.16) ..... (2.17) ....... (2.18) total number of parameters to be estimated. In the two explanatory variables model. k = 3 and = on2 Se2 n - 3 Example: 3.1 Applications using matrix algebra with one dependent and two independent variables. | Y | :      49    40    41    46    52    59    53    61    55    60 | |---|---| | X1 X2 | :      35    35    38    40    40    42    44    46    50    50 :      53    35    50    64    70    68    59    73    59    71 | Worsheet for the model : Y = α + β1x1+ β2x2+ 4 | y       x1 | x2      y1       x       x       x2      y12      x12      x22      x1y1    x2y2    x1x2 (y - y)          (x - x)         (x -x2 ) | |---|---| n | 1      49 | 35    53    -3     -7     -9     9      49    81     21     27    63 | |---|---| | 2     40 | 35    53    -12   -7    -9     144   49    81    84    108   63 113 | | 3 | 41 | 38 | 50 | -11 | -4 | -12 | 121 | 16 | 144 | 44 | 132 | 48 | |---|---|---|---|---|---|---|---|---|---|---|---|---| | 4 | 46 | 40 | 64 | -6 | -2 | 2 | 36 | 4 | 4 | 12 | -12 | -4 | | 5 | 52 | 40 | 70 | 0 | -2 | 8 | 0 | 4 | 68 | 0 | 0 | -16 | | 6 | 59 | 42 | 59 | 7 | 0 | 6 | 49 | 0 | 36 | 0 | 42 | 0 | | 7 | 53 | 44 | 68 | 1 | 2 | -3 | 1 | 4 | 9 | 2 | -3 | -6 | | 8 | 61 | 46 | 73 | 9 | 4 | 11 | 81 | 16 | 121 | 36 | 99 | 44 | | 9 | 55 | 50 | 59 | 3 | 8 | -3 | 9 | 64 | 9 | 24 | -9 | -24 | | 10 | 60 | 50 | 71 | 12 | 8 | 9 | 144 | 64 | 81 | 96 | 108 | 72 | | n = | 10 Sy | Sx1 | S x2 | Sy | =0 Sxi = | 0 Sx2 = | 0 Sy? | Sx12 | S22 | = 594 | = 270 | = 630 | | | 520 | = 420 = | 620 | | | | | | | O' | CM O o ^r | | K K M M y = 52, x = 42, x = 62 12 Y = α + β1 X1 + β2 X2 + u .......... 3.1 In the matric notations P = (X' X)-1 = X’ Y Where (when we use die quantities in deviation form). | .A P = | ’ Pi ’ _ P2 _ | x = | xii . Xi2 | x21 x22 _ | so that | |---|---|---|---|---|---| | | | r | Sx2 | x1 | x2 | | | Sxi y | | x1 x | = | Sx1x 2 Sx2 | J | | and X' Y = | SX2 y | Substituting the relevant quantities, we have 270 240 240 630 and X' Y = 319 492 X' X = 270 x 630)-(240 x240) = 112500. | | 1 1 | 630 - | - 240 | |---|---|---|---| | (X' X)-i = | 112500 | - 240 | 270 | | | | | | | 0056 | - 002i | | 3i9 | | = - 002i | 0024 | X' Y = | 492 | βˆ1 = 0.7532 βˆ2 = 0.5109 The value of a is determined from the relation: ˆˆ (3.2) (3.3) a = y -Pi Xi - P2 X2 = 52 - (.7532 × 42) - (.5109 × 62) = 52 × 31.63 - 31.6578 = - 11.29 y = 11.09 + 7532 X1 + 5109 X2 Goodness of Fit R2 A                     /V R Pi sxiy i + 02^x2i y i sy? = 7532(319) + 5109 x 492 594 = 240•2708+251•3628 = 491•6336 =        594         =   594 = 82                                         ..... (3.4) The variables X1 and X2 explain 82 percent of the total variations in Y. Se12 = (1 - R2) Syi2) (1-82) 594 106.92 Self-Check Exercise-2 Q1. Define Goodness of Fit (R2). Q2. Derive the following terms a)    Explained Sum of Square b)    Residual Sum of Square c)     Total Sum of Square Q3. The Following data show the experience of machine operation and their performance rating s as given the number of good parts turned out per 100 pieces. | Operator | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | |---|---|---|---|---|---|---|---|---| | Experience (X) | 16 | 12 | 18 | 4 | 3 | 10 | 6 | 12 | | Performance Rating (Y) | 87 | 88 | 89 | 68 | 78 | 80 | 75 | 83 | Calculate the regression line of performance ratings on experience and estimate the probable performance if an operator has 10 years’ experience. 6.5    Summary In this unit, we explored how to create and solve general linear regression models using matrix methods. We learned to handle models with one, two, and three explanatory variables and assessed their effectiveness using R2. This unit helps you understand and apply regression analysis to real-world data. 6.6    Glossary •   Dependent Variable: The outcome variable in a regression model that the explanatory variables aim to predict or explain. •    Explanatory Variable: An independent variable in a regression model that explains or predicts changes in the dependent variable. •    General Linear Regression Model: A statistical method used to predict the value of a dependent variable based on one or more independent variables using linear equations. •   Matrix Formulation: A way to represent and solve regression models using matrices, which simplifies calculations and provides a clear framework for handling multiple variables. •    Goodness of Fit (R²): A measure that indicates how well the regression model’s predicted values match the observed data. A higher R2R^2R2 value means a better fit. 6.7    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 6.3. Answer to Q2. Refer to Section 6.3. Answer to Q3. Refer to Sections 6.3.1, 6.3.2 and 6.3.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 6.4. Answer to Q2. Refer to Section 6.4. Answer to Q3: Y = Y = 2 + 1.6 X 6.8    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Freund, J. E., & Williams, F. J. (1959). Modern business statistics. London: Pitman and Sons. 6.9    Terminal Questions Q1. Define the General Linear Regression Model. What is the purpose of the coefficients β0, β 1 and β2 in a General Linear Regression model? Q2. Explain the concepts of Total Sum of Squares, Residual Sum of Squares, and Explained Sum of Squares in the context of a General Linear Regression model. How are these concepts used to assess the model’s goodness of fit? ***** UNIT-7MULTIPLE AND PARTIAL CORRELATION-I STRUCTURE 7.1    Introduction 7.2    Learning Objectives 7.3    Multiple Correlation 7.3.1    Distribution of Three Variables Self-Check Exercise-1 7.4    Variance of a Residual Self-Check Exercise-2 7.5    Multiple Correlation Coefficient Self-Check Exercise-3 7.6    Summary 7.7    Glossary 7.8    Answers to Self-Check Exercise 7.9    References/Suggested Readings 7.10    Terminal Questions 7.1    Introduction Dear Student, In correlation and regression, the main purpose of the analysis was’ to determine the degree of linear relationship between two variables and to predict the behaviour of the dependent variable on the basis of n independent variable. But often, it is necessary to find correlation between three or more variates. For example, the stature of men is influenced by those of all their ancestors, and the yield of grain is affected by the amount of irrigation and fertilizers used. Whenever we are, interested in the combined influence of a group variates upon a variate not included in the group, our study is that of multiple regression and multiple correlation. Therefore there arises a need for the study of multiple correlation. 7.2    Learning Objectives After the completion of this unit, you will be able to •    Understand and explain the concept of multiple correlation and how it involves the relationship between three variables. •    Calculate and interpret the variance of a residual in the context of multiple correlation analysis. •    Determine and analyze the multiple correlation coefficient to assess the strength of relationships among multiple variables. 7.3    Multiple Correlation Multiple correlation may be defined as a statistical tool designed to measure the ‘degree of relationship “existing among three or more variables. For example; we may be asked to find the relationship, between the yield of wheat and amount of irrigation, fertilizers, seeds, insecticides and the spacing of the plants. In a multivanate population, the different variates, may be mutually correlated and the correlation will, in general, be influenced, by the other variates of the population. To study the relationship between any two variables, there are two methods, firstly, we may consider only those members of the observed data in which, the other members have ‘specified values. Secondly, we may eliminate mathematically the effect of other variables on the two variables under study. The first method has the disadvantage that it limits the size of the data and also the result of this, will be applicable to only those data in which the other variables have assigned values. In the second method it may not be possible to eliminate the entire effect of other variables but we can easily eliminate the linear effects. The correlation between two variables when the linear effect of the other variables in them has been eliminated from both is called partial correlation. 7.3.1    Distribution of three variables: The theory of multiple and partial correlation was developed by Kari Pearson (1896) for three variables and then generalized by G. Udny Yule (1897). For the sake of simplicity, we shall study the distribution of three variables only though the arguments will apply to the case of n variables also. Let these variables, be measured from their respected means and the quantities so obtained fee denoted by x1, x2 and x3. In multiple regression with two independent variables, there are three constants involved in the equation. Three normal equations are required to compute the values of three constants. The regression equation of X1, on X2 and X3 is of the form X1 = a + b12.3 X2 + b13.2 X3 .......... 1. Where the constants a and b’ s are such as to give on the average the ‘best’, estimate of x1 corresponding to any assigned values of x2 and x3.Thus we are to find a and b′ such that µ = Σ(x1 - x1) 12 = Σ(x1 - a-b12.3 x2 - b13.2 x3) 2 Σx2 123 where x1.23 = x1 -a-b12.3 x2 - b13.2 x3) 2 (6.1) is a minimum said the summation is taken place over all sets of values of x2 and x3. The three normal equations for determining a & b′s are Σ(x1-a-b12.3 x2 –b13.2 x3) = 0 Σ x2(x1-a-b12.3 x2 –b13.2 x3) = 0 Σ x3(x1-a-b12.3 x2 –b13.2 x3) = 0 | Σ x1.23= 0 | | Σ x2x1.23= 0 | | Σ x3x1.23 = 0 |                         …… (6.2) The first of these equations gives a = 0 and the last two equation’s may be written in the form. Σx1 x2 -b12.3 Σx12 -b13.3 Σx2 x3 = 0                      …… (6.3) Σx1 x3 -b12.3 Σx2 x3 -b13.3 Σx32 = 0                     …… (6.4) The coefficient of correlation between X1 and X2 (from mean) is given by: Σx1x2 r = 12   Nσ1σ2 or n o1 02 r12 = 2x1 x2 2x1x3 r = 13    Ng\g3 N o1 03 r13 = 2x1 x3 2x2 x3 r = 23  Na2a3 Similarly N O2 3 r23 = 2x2 x3 The standard deviation (from mean) of X1 series o, = ^xL or No,2 = 2x,2 111 N Similarly No12 = 2x12 and No32 = 2x32 Substituting and above values in the equations (6.3) and (6.4) we have Nr12 o1 o2 = Nb12 3 o22 + Nb13 2 r23 o2 o3                      (6.5) Nr13 o1 o3 = Nb12.3 o23 + r23 + b13.2 o3 Since No2 is common in equation 5 and No3 in equation           (6.6) r12 o1 = b12.3 o2 + b13.2 r23 o3                                         r13 o1 = b12.3 o2 + r23 b13.2 o3                                          Dividing the equation (6.8) by r23 and subtracting the result from equation (6.7) we get: r12 o1 = b12.3 o2 + b13.2 o3 r23 r13 r23 ^3 o1 = b12.3 o2 + b13.2 r r23 o1 r12 r13 - o, 1 r23 = b12.3 o3 r23 — b ^3 13.2 r23 or o1 or o1 or o1 r12 r13 \ = b k r23 7 r12 r23 r13 \ 12.3 o3 = b r23 k 1 \ r23 7 k r23 7 13.2 o 3 r23 1 \ k r23 7 ^1 r12 r23 r13 \ r23 \ ^3 k r23 7 2 k r23 a 1 7 or b13.3 51 ( ■        1 53 V  r23   1 J Substituting the value of b13.2in equation (7) we get: or b12.3 51 ( iliJi1 52  V  r23   1 J or b13.3 can be write as = 51 [       ' 1 53 V 1 r23 J and 5 ( rl2 r— r3 b“ = 51 V 1 b 12.3 r12   51 r12   51 r13 5 5 _ 51 r12 r23 + 5 r23 r23 52 53 53 5 r13 1 1 r23 (6.10) Where Aij is the co-factor of the element in the ith row and jth column in the determinant. | 1          r12         r13| A =     । r21      1         r23 | r31       r32           1| 1 and r33 = 1 and r13 = r31 ; r12 = r21 and r23 = r32 1 b12.3 = 53 53 1 r23 1 r13 r23 1 r12 r13 51 A13 52 A11 Hence on substituting the values of b12.3 and b13.2, the equation to the regression plane of x1 and x2 and x3 is 0-1 A12 c2 A11 X2 + o1 A13 03 A11 x3 (6.11) Example 6.1: Find the least square regression equation of X3 on X1 and X2 from the following data : X = 6.8 X2= 7.0 X3 = 74 O1 = 1.0       02 = 0.8      o3 - 9.0 r12 = 0.6       r13 = 0.7       r23 – 0.65 The regression equation, of X3 on X1 and X2 is of the form: X3 and a3.12 + b3.12 X1 + b32.1 X2 Computation of regression on coefficients: = °3 f r13  r23 rl2 b31.2    O1  [ 1 - r12   J Substituting the values, 9.0 f 0.7 - 0.65 x 0.6A b31.2 = 1.0 [    1 - (0.6)2 b31.2 = 9 0.7 - 0.390A 1 - 0.36 J = 4.36 A d b31.2 03f rv^ ' O2 < 1 - r2 J 9.0 = 0.8 = 4.04 f 0.65- 0.7 x 0.6A < 1 - (0.6)2 y The value of a  may be computed from the following equation: . a3.12 + X 3 - b31.2 X 1 - b32.1 X 2 = 74 – 4.36 × 6.8 – 4.08 × 7.0 = 16.7m Thus the required equation is X3 – 16.07 + 4.36X1 + 4.04 X2 Ans. Example 2 : Given the following: Determine the regression equation of X1 on X2 and X3. X = 6.8      01 = 1.0       r12 =0.6 X2 = 7.0     02 = 0.8      r13 -0.7 X3 = 7.4     03 = 0.9      r13 -0.8 Sol. Regression equation of X1 on X2 and X3 in determinant form is A11 A12     A13 o1  x1 + °2 x + x = 0 2    03    3 The values of cofactors Δ11, Δ12 and Δ13 are determined from the following relationship Δ = r11 r21 r31 r12 r22 r32 r13 r23 r33 Since r11 = 1, r22 = 1; and r33 = 1 and r12 = r21 and substituting the values r13 = r31 and r23 = r32. 1   0.60.7 0.6    10.8 0.7  0.81 1 0.8 Δ11 0.8 1 = (1-(0.8)2)= 1- 0.64 = 0.36. 0.6 0.8 Δ 12 Δ 13 0.7 1 = [0.6-(0.8 × 0.7)] - 0.6 - 0.56 = 0.04 0.6     1 0.7 0.8 = [0.48 - 0.70] = - 22 Substituting the values in above regression equation. 0.36       0.04      -22 1   x1 + 0.8 x2 + 0.9 x3 = 0 0.36 x1 = - 0.05 x2 + 24x3 0.05      .24 xi = .36 x2 + .36 x3 x1 = 0.138x2 + 66x3 The value of ‘a’ is computed from the following equation : a1.23 + X 1 – b12.3 X 2 – b13.2 X3 Putting the value of X 1, X 2, X 3, b12.3 and b13.2 a1.23 = 6.8 + 0.138 17 - .66 × 7.4 = 6.8 + 0.966 - 4.88 = 2.88 Thus the required regression equation of X1 on X2 and X3 becomes X2 = 2.88-0.138 X2 + 0.66 X3. Self-Check Exercise-1 Q1. Define multiple correlation. Q2. The table shows the corresponding values of three variables X1, X2, and X3 | X1 | 3 | 5 | 6 | 8 | 13 | 14 | |---|---|---|---|---|---|---| | X2 | 16 | 10 | 7 | 4 | 3 | 2 | | X3 | 90 | 72 | 54 | 42 | 30 | 12 | Find (i) Least Square Regression equation of X3on X1 and X2, and estimate X3when X1= 10 and X2= 6. Variance of a Residual : We shall now obtain formula for the variance α21.23 of the residual x1.23 (the deviation of the observed Values x1 from its computed value of the regression plane), in terms of 12 and the correlation coefficients. We have NO21.23 = Sx2i2.3 = Sxi xi.23 = Sxi(xi - bi2.3 x2 – b13.2 x3) = N< -Nbi2.3 G1 O ri2 - Nbi3.2 G1 O3 ri3 Or N< NO . = Nbi2.3 ^1 O ri2 + Nbn.2 ^ O r^ Dividing both sides by o1 or °i 2 °1.23 i = b12.3 ^ r12 + b13.2 ^ r13 °1 2 1—ir \     °12 7 b12.3 °2 r12 + b13.2 °3 r13 eliminating b12.3 and b13.2 between this equation and equations. r12 °1 = b12.3 ^ + b13.2 r13 ^ r13 °1 = b12.3 ^ + b13.2 ^ We have 2 1-i1.23 r r r12 r13 i12 | r12      1         r3 | = 10 | r13       r23         1 | 2 ^1 A — i V = 0 <23 A °1 A11 Thus the variance of residual of order 2 is expressible in variance of zero order and correlation of zero order. Self-Check Exercise-2 Q1. Derive the formula for variance of the residual. 7.5. Multiple Correlation Coefficient: Consider the regression equation for x1 on x2 and x3 viz., Next the correlation between x1 (observed value of the variable) and X1 (expected value of the variable) is given by 2x1X1 R- = We have 2xX1 = 2(x1 (x1 - x1 23)) = Sx12 -dx1 x1 23 = Sx12 = 2x21 23 = W -^1.23 = W12 — <23) Also 2x12 = ^(x1 x12.3) 1x1   2 Sxi xi.23 + Sx 1.23 = 2x12 = 2 2x21.23 + 2x 2 1.23 = Sxi2 = 2x21.23 = N^12 —N<23 1. R1.23 22 —1   ^1.23 2 °1.23 22 01    °1.23 \ —1 — ^1.23 1 — 2    1/2 ° 1 J 2 1 0123 - Where c1 stands for the standard deviation of a dependent variable x1c1 23 stand for the standard error of estimated of x1 on x2 and xy. It is computed by the following formula : = 2(X1 - X1est )2 °123         n Where X1 est stands for the estimated value of X1 as compared with the aid of regression equation of X1 on X2 and X3. It may also be computed by the following formulae °1.23 = °. 7(1^ ^ Or S21.23 = ° ■ (1-r212) (1-r213.2) 2.    For the two independent variables, the coefficient of multiple correction may also be determined with the help of zero order coefficient of correlation. _    r12 + r13    2r12 r13 r23 R1.23 = V 1 - r23 Where r12, r13 and r23 stand for zero-order coefficient of correlation r12 + r14   2r12 r14 r24 R1.24 R1.23 SX 21.23 2x12 Where x2c1.23 stands for explained variations in the values of variable X1 which has been explained by two (independent variables X2 and X3. It may be completed by the following formula: x2c1.23 = b12.3 = Sx1x2 + bi3.2 Sx^ Example : Calculate the multiple correlation coefficient of X1 and X2 and X3 from the following data: | X1 X2 X3 Solution: | :             5 :         10 :        21 | 7 18 21 | 8 6 15 | 10 5 17 | 12 4 20 | 16 3 12 | |---|---|---|---|---|---|---| | X | 1 X1 X2 X2 | X2 - | X3 | | | | | X1   X2 | X3     = x1 | x2 | x3 | x12 | x22 | x32     x1x2      x2x3    x1x3 | | 5    10 | 21      -5 | +4 | +3 | 25 | 16 | 9      -20     12     -15 | | 78 | 21      -3 | +2 | +3 | 9 | 4 | 9      -6       6      -9 | | 86 | 15     -2 | 0 | -3 | 4 | 0 | 9     0      0     +6 | | 10   5 | 17     0 | -1 | v1 | 0 | 1 | 1      0       +1     0 | | 12   4 | 20    +2 | -2 | +2 | 4 | 4 | 4      -4      -4     4 | | 18   3 | 14     +8 | -3 | -4 | 64 | 9 | 16    -24     +12   -32 | | SX1 SX2 | SX3   S x1 | S x2 | Sx3 | S x12 | Sx2 | 2   Sx32 =48 S X1 X2 | | =60 = 36 = | 108     =0 | =0 | =106 =34 | | | | X=10 X 2 | =6 X3= 18 | | | | | | | | | | | | | Sx2 x3=27 | | | | | | | | SX1 X3 = - 46 | For computing the value of explained variations. We need the value of b12.3 and b13.2 and these may be computed, from regression equation of X1 on X2 and X3which may be written in the deviation form as given below: x1 = b12.3 x2 + b13.2 x3 Normal equations: Sx1X2 + b12.3 S^ + b13.2 Sx2x3 Sx1X3 + b12.3 Sx2 x3 + b13.2 Sx3 Substituting the values in above equations: -54 = 34 b12.3 + 27 b13.2 -46 = 27 b12.3 + 48 b13.2 Multiplying equation 3rd by 16 and 4th by 9 we get: . (1) . (2) (3) (4) -864 = 544 b12.3 + 432 b13.2 -414 = 243 b12.3 + 432 b13.2 -450 = 301 b . 450 b13.2 = - 301 = - 1.4950. b12.3 = 0.4010. Calculation of Explained variation: Σx2c 1.23 = b12.3 Σx1 x2 + b13.2 Σx1x3 = - 1.4950 (-54) + 0.4010 (-46) = 80.73 + 18.4473 2   ΣXc21.23      62.2826 2c R Σx12          106 = 0.7665 Ans. Self-Check Exercise-3 Q1. Find the multiple correlation coefficient R1.234 when R12 = 0.9, r14 = 0.4193, r13 = 0.75, r23 = 0.7 7.6    Summary In this unit, we explored multiple correlations, focusing on how three variables interact and their distribution. We learned to calculate the variance of residuals and understand its significance. The unit also covered the multiple correlation coefficient, providing insights into the strength of relationships among multiple variables. 7.7    Glossary •    Multiple Correlation: Examines the relationship between three variables simultaneously, assessing how they interact and influence each other. •    Multiple Correlation Coefficient: A statistic that quantifies the strength and direction of the relationship among three or more variables, providing a comprehensive view of their associations. •    Variance of a Residual: Measures the variability or difference between observed and predicted values in a regression model, indicating how well the model fits the data. •   Residual: The difference between the observed value of the dependent variable and the value predicted by the regression model. Residuals are used to assess the model’s accuracy and fit. 7.    8 Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 7.3. Answer to Q2: X3 = 61.4 – 3.64X1 + 2.54 X2 Self-Check Exercise-2 Answer to Q1. Refer to Section 7.4. Self-Check Exercise-3 Answer to Q1. Refer to Section 7.5. 7.9    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Mill, F. C. (1955). Statistical methods. Pitman and Sons: London. 7.10    Terminal Questions Q1. Define multiple correlation with the help of an example. Calculate R1.23 from the following zero order correlations: r12 = 0.9, r23 = 0.4 and r13 = 0.5 Also obtain the coefficient of multiple determination and interpret it. If n = 30, how would you test the significance of r .. Q2. If r12 = 0.60, r13 = 0.50 and r23 = 0.45, then calculate r12.3, r13.2 and r 23.1. ***** UNIT-8MULTIPLE AND PARTIAL CORRELATION-II STRUCTURE 8.1    Introduction 8.2    Learning Objectives 8.3    Partial Correlation Self-Check Exercise-1 8.4    Partial Correlation Coefficient Self-Check Exercise-2 8.5    Standard Error of Estimate Self-Check Exercise-3 8.6    Summary 8.7    Glossary 8.8    Answers to Self-Check Exercise 8.    9 References/Suggested Readings 8.10 Terminal Questions 8.1    Introduction In the previous unit, we learned about multiple correlation. In this unit, we will discuss partial correlation, which measures the relationship between two variables while controlling for the influence of one or more additional variables. After completion of this unit, you will learn to calculate and interpret the partial correlation coefficient and understand its significance. Additionally, we will discuss the standard error of estimate, which helps in assessing the accuracy of predictions in regression analysis. 8.2    Learning Objectives After going through this unit, you will be able to •    Understand and explain the concept of partial correlation. •    Calculate the partial correlation coefficient to assess the relationship between two variables while controlling for other variables. •    Evaluate the accuracy of regression predictions using the standard error of estimate in the context of partial correlation analysis. 8.3    Partial Correlation Partial correlation is a statistical technique used to understand the relationship between two variables while removing the influence of one or more additional variables. Imagine you want to study the relationship between exercise and weight loss, but you also know that diet plays a crucial role. By using partial correlation, you can measure how much exercise alone contributes to weight loss while keeping the effect of diet constant. This helps you get a clearer picture of how two specific variables interact without the noise from other influencing factors. Partial correlation is valuable in research and data analysis, providing deeper insights into complex relationships. Self-Check Exericse-1 Q1. Define Partial correlation. 8.4    Partial Correlation Coefficient A partial correlation coefficient measures the relationship between any two variables, when all other variables connected with those two are kept constant. For example, let us assume we want to measure the correlation between the number of cold drinks (X1) consumed during summers in Shimla and the number of tourists (X2) coming to Shimla. It is obvious feat both these variables are strongly, influenced by weather conditions, which, we may designate by X3. On a priori, grounds we expect X1, and X2 to be positively correlated when a large number of tourist arrive in the summer resort (Shimla), one should expect a high consumption of cold drinks and vice-versa. The computation of the simple correlation coefficient between X1 and X2 may not reveal the true relationship connecting; these two variables, because of the influence of third variable X3. In other words the above positive, relationship Between number of tourists and number; of cold drinks consumed is expected to hold if weather conditions can be assumed constant If weather changes, the relationship between the consumption of cold drinks and number of tourists arrived may be distorted to such an extent as to appear negative. Thus if the weather is cold (due to Unexpected rain), the number of tourists will be less, but because of the chilly weather, they will prefer to consume hot drinks (tea or coffee) rather than cold drinks. If we overlook weather and look only at X1 and X2. We will observe a negative correlation between, these two variables which is explained by the fact cold drinks and as well number of visitors is affected by unexpected weather. In order to measure the true relationship between X1 and X2 we must find some way of accounting for changes in X3. This is achieved with the partial correlation coefficient between X1 and X2 when X3 is kept constant. The partial correlation coefficient is determined in terms of the simple correlation coefficient among the various variables involved in multiple relationship. In the above mentioned example there are three simple Correlation Coefficients. r12 = Correlation Coefficient between X1 and X2. R13 = Correlation Coefficient between X1 and X3. X23 = Correlation Coefficient between X2 and X3. There are two partial correlation coefficients. r12.3 = partial correlation coefficient between X1 and X2 when X3 is kept constant r12 - (r13) (r23) r12.3 = a/ 1 - r123 (1 - r223) and r13.2 = partial correlation coefficient between X1 and X3 when X2 is kept constant. r12 -r12 r23 r13.2 = a/ 1 - r122 (1 - r223) We may postulate the following functional relationship between the number of tourist in Shimla (X1) and She consumption of cold drinks (Y) : Y = b0 + b1 X1 + b2 X2 + 4 Y = consumption of cold drinks X1 = number of tourists, in Shimla (summer) X2 = weather conditions measured by an index of rainfall or temperature The partial correlation coefficient measures the correlation between any two variables, when all the other variables are held constant, that is when we assume, that there is no other, factor influencing, the relationship. Partial correlation coefficient for the model including two explanatory variables is ryx1, x2 ryx1   ryx2 rx1x2 1   ryx2 (1   rxlx2 ) The partial correlation, on coefficient between y and x2 when x1 is kept constant is obtained from this expression by interchanging the position of the subscripts 1 and 2. ryx2, x1 ryxl   ryx2 rxlx2 V1-rÿx1”(1—"rxlx^^ Proof: The following relationship between regression coefficients and simple correlation coefficients has been established in lesson. r2yx = b or ˆ = r -         Sy2        i yx In the above mentioned case aˆ1 = ryx2 and c=r c1    rx1x2 From the correlation coefficient of the two regressions, we get Se22        Sx22 r2 yx2 = 1 ■ Sy2 = 1 - Sy and Se2 r2 x1x2 = 1 - sx2 = 1 — Sx*2 Sx12 Therefore Sy*2 = Sy2 (1- r2 yx2) and Sx1 *2 = Sx12 (- r2 x1x2) Substitute the terms with asterisks in the formula of partial correlation coefficient S (y - a1 x2 )(x1 -c1x2) ry" x2 = JSy’RS 4 Sx?(1 - r1 x, >2) S (yx1 — â1x1x2 —I 2 ■ c1 yx2 + a1 c1 x2 ) 7 Sy2 Sx2 JÔ-îÏKi-rïxx) The rationalization of these formulae may be propounded as follows. In order to measure the pure correlation between y(cold drinks) and x1 (numbers of tourists), the influence of the third variable X2 (weather conditions) has to be eliminated from both Y and X1. This can be done by regressing Y on X2 and X1 on X2. Y = a0 + a1 X2 + u2 X1 = C0 + C1 X2 + u2 Where u1 and u2 and error term satisfying the usual assumption of zero mean and constant variance. From the application of least squares, fee following, estimates are obtained. ~    Yyx2            Se2 ai = ex2 r2yx2 = 1 - Sy2 1x1x2              Se2 c1 = Sx22 r^x2 = 1 — Sx2 The unexplained variance in each regression is e1 = Y1 – yˆ = Y1 aˆ1x1 = Y* and    e2 = X1 - X1 = X1 cˆ1x2 = X1* These are variations’ in Y and in X1, respectively, left unexplained after removing the influence of X2 (weather conditions). The partial correlation coefficient between Y and X1 which is defined as the simple correlation between the above mentioned unexplained parts of the two variables can be proved. ryx1x2 ry* x1* ΣY*X1* ΣY*2 Σx1*2 2 1yx1 - a1x1x2 -c1 1yx2 + a1 c11x2 ïryX' —ryx2 1X1X2 rx1x2 A -yX 'r. Srx1x2 J 1x22 Tsy7 T1XF 7(1—ryx 2)(1 - r2 x1x2) Multiplying each term of the numerator by appropriate unitary term so that these are transformed into simple correlation coefficients. Multiply the first term ^/Sy2 7ÏX2 -JSy2 71x12 = 1, the second term by ryx1 x2 = Syx1 -ryx2 ^xXl rx1 x2 + ryx2 rxi x2 TiTiXF \yy' a/2? F(1 - ryx 2)(1 -r 2 xix2) ^y^^x^ (ryx1 — ryx 2 rxlx 2 - ryx 2 - rxlx 2 Sy2 2x12 ^Q1 - r2xïx2~) + ryx 2 rx1x2 ) r — r r ryx1 ryx2 rx1x2 V1 — ryx2 F1 - rxlx2 Problem: If r12 = + 0.80, r13 = - 0.40, r23= - 0.56, find the values of r12.3, r13.2, and r23.1. we have r12 — r13 r23 r23 = V1—ryx 2 Vï-Fl 0.80 — (—0.40)(—0.56) = J (0.40, } {1 — (0.56)2} = 0.759 r13 - r12 r32 Similarly ri3.2 = J (1 — r.2 )(1 — r322 ) = 0.097 r23 — r21 r31 and      r- = 4 (1—'24(1—r2J = 0.436 Example: On the basis of observations made on 50 households the linear correlation of coefficients, between X1 (quantity demanded of tea), X2 price of tea) X3 (households income) one as follows : r12= 0.75, r13= 0.80, r23 = 0.55. Where rij is the correlation coefficient between X1 and Xj Calculate partial correlation coefficient’s of : (a)    quantity demanded of tea with price of; tea. (b)    quantity demanded of tea with household’s income. Solution : (i) The coefficient of partial correlation between quantity demanded and price when the effect of income is kept constant is given by : r12 r13 r23 ■ = »-a Putting the values 0.75 - 0.80 0.55 = 7 {1 - (0.80)2 }{1 - (0.55)2} 0.75 - 0.44   _ 0.31 7036 x 7039 " 0.6 .8 3 0.31 0.498 = 0.62 (ii)    The coefficient of partial correlation between quantity demanded and the income of households, when the effect of price iskept constant, isgiven by r13 — r12 r23 r“ = Í-ÍO Substituting the given values for r12, r13 and r23 we obtain 0.80 - 0.75 0.55 = 71 — (0.75)21- — (0.55)2 0.80 - 0.41     0.39 = 41 - .56 V1 - 30 66 x 83. Example : The variable X1 is thought to be a linear function of X2 and X3. A sample of 12 pairs of readings (X2, X3) produced the values of X1 shown in Table 6.1 below : (a)    Find the least square regression equation of X1 on X2 and X3. (b)    Determine the estimated values of X1 from the given values of X2 and X3. (c)    Estimate X1 when X2 = 54, and X3 = 9 Table 6.1 | X1 | 64 | 71 | 53 | 67 | 55 | 58 | 77 | 57 | 56 | 51 | 76 | 68 | |---|---|---|---|---|---|---|---|---|---|---|---|---| | X2 | 57 | 59 | 49 | 62 | 51 | 50 | 55 | 48 | 52 | 42 | 61 | 57 | | X3 | 8 | 10 | 6 | 11 | 8 | 7 | 10 | 9 | 10 | 6 | 12 | 9 | Solution: The linear regression equation of X1 on X2 and X3 can be written X1 = b12.3 + b12.3 X2 + b13.2 X3 The normal equations of least square regression equation are ΣX1 = b1.23 N + b12.3 ΣX2 + b13.2 ΣX3 ΣX1 X2 = b1.23 ΣX2 + b12.3 ΣX22+ b13.2 ΣX2 X3 ΣX1 X3 = b1.23 ΣX3 + b12.3 ΣX2 + X3 b13.2 X32 | ΣX1 = 753 | ΣX12 = 48139 | ΣX1 X2 = 40830 | |---|---|---| | ΣX2 = 643 | ΣX22 = 34843 | ΣX1 X2 = 5779 | | ΣX3 = 106 | ΣX32 = 976 | ΣX1 X3 = 6796 | Using the values, the normal equation (i) becomes. 12 b1.23 + 643 b12.3 + 106 b13.3 = 753. (2)    643 b1.23 + 34843 b12.3 + 5779 b13.2 = 40830. 106 b1.23 + 5779 b12.3 + 976 b13.2 = 6796. Solving       b12.3 = 3.6512 b12.3 = 0.8546 b13.2 = 1.5063, and the required regression equation is (3)    X1 = 3.6512 + 0.8546 X2 + 1.5063 X3 or X1 = 3.65 + 0.85 X2 + 1.50 X3 (b) Using the regression equation (3) we obtain the estimated value of X1, denoted by X1est., by substituting the corresponding values of X2 and X3. For example, substituting X2 = 53 and X3 = 8 in (3) we find X1est.= 64.414. Similarly the other estimated values of X1 are obtained and given in the Table 6.1 together with the sample values of X1. Table 6.2 | X1est. | 64.41 | 69.13 | 54.56 | 73.20 | 59.29 | 56.92 | 65.71 | 58.22 | 63.15 | 48.58 | 73.85 | 65.92 | |---|---|---|---|---|---|---|---|---|---|---|---|---| | X1 | 64 | 71 | 53 | 67 | 55 | 58 | 77 | 57 | 56 | 51 | 76 | 68 | (c) Putting X2 = 54 and X3 = 9 (in) (3) the estimate is X1est. = 63.556 or about 63 Ans. Self-Check Exercise-2 Q1. Define partial correlation coefficient. Q2. If r12 = +0.80, r13= - 0.40, and r23 = - 0.56, find the values of r12.3 ,r13.2 r23.1. 8.5    Standard Error of Estimate Compute the standard, error of estimate of X1 on X2 and X3 for the data of problem 6.1. Solution: From Table 6.1 of problem 6.1 (b), we have (1) S1.23 Σ(X12-X1est)2 N (64-64.41) +(71-69.13)2+ (68-65.92)2 = 4.6447 = 4.6. The population standard estimate is estimated by Sn,=,/N/N- 3 S 9,= 5.3 Ans. . or Alternatively c          1 - (0.8196)2 - (0.7698)2 -(0.7984)2 + 2(0.8196) (.7689)(.7984) S n, = 8603„ -------------------------------------z-------------------------- 123                                    1-(0.7984)2 = 4.6 The standard error of estimate can be found without use of regression equation in the above mentioned method. | Example : | If R123 = 1, Prove that (a)    R2.13 = 1 (b)    R3.12 = 1 | |---|---| Solution: _    r12 + r23 - 2r12 r13 r23 (1)     R1.23 =            1-r223 and r12 + r23   2r12 r13 r23 2 1 - r23 (a) In (1) setting R = 1 and scaring both sides, r212+ r213 - 2r12 r23= 1 r223. Then r212 + r213 - 2r12 r23= 1 r213 or r12 + r23 - 2r12 r13 r23 2 1 - r23 = 1. i.e. R22.13 = 1 or R2.13= 1. Since the coefficient of multiple correlation is considered non-negative. (b) R3.12 = 1 follows from part (a) by interchanging subscripts 2 and 3 in the result R2.31 = 1 Example: If R1.23 = 0, does it necessarily follow that R2.13 = 0 ? Solution: R = R2.13 r12 + r23 - 2r12 r13 r23 1- r223 or R1.23 = 0 if and only if r122+ r213 - 2r12 r13 r23 = 0 r122 + r213 - 2r12 r13 r23 R2.13 r12 + r23 - 2r12 r13 r23 2 1 - r23 Since 2r12 r13 r23 = r122+ r213 By putting these values 22 22 r12 + r23   (r12 + r13) 2 1 - r13 22 r23 + r13 1-    r123 which is not necessarily zero. Partial Correlation: Compute the coefficient of linear partial correlation (a) r12.3 (b) r13.2 and (c) r23.1 for the data in problem 6.1. (a) The quantity r12 is the linear correlation coefficient between the variables X1 and X2, ignoring the variable X3. r12 = NSX1X2 - (XXJ^XJ 7 [n sx 2 - (lx2 )] [n sx 2 - ( sx 2 )] =________(12)(40830)- (753) x (643)________ 7[(12)(48131)- (753)2] [(12)(34843)- (643)2] = 0.8196 or 0.82. Using the above formula, we obtain r13 = 0.7698 or 0.77 r23 = 0.7984 or 0.80 r12 — r13 r23 Now ri2.3 = (1 — r¿ )(1 — r-22 ) r13 r12 r23 r13'2 = (1 — ri22)(1 — ri2 ) r23 -r12 r13 r23'1 = (1 -r-n)(1 — ri! ) Putting the values, we find r12.3 = 0.5334 r13.2 = 0.3346 r12.3 = 0.4580 It follows that the constant X3 the correlation coefficient between X1 and X2 is 0.53. For constant X2, the correlation coefficient between X1 and X3 is only 0.33. Since these results are based-on a small sample of only 12 observations, they are of course not that reliable as those which would be obtained from a larger samples. Self-Check Exercise-3 Q1. Compute the standard error of or X1 on X2 and X from the table 6.2. 8.6    Summary In this unit, we focused on partial correlation, which examines the relationship between two variables while controlling for additional variables. We learned to calculate the partial correlation coefficient and explored its significance. The unit also covered the standard error of estimate to evaluate the accuracy of regression predictions. 8.7    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 8.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 8.4. Answer to Q2. Refer to Section 8.5. Self-Check Exercise-3 Answer to Q1. Refer to Section 8.5. 8.8    Glossary •   Multiple Correlation: Examines the relationship between a dependent variable and two or more independent variables simultaneously. •    Partial Correlation: Measures the relationship between two variables while controlling for the influence of one or more additional variables. •   Control Variable: A variable that is held constant or whose impact is removed to better understand the relationship between other variables. •   Partial Correlation Coefficient: A statistic that quantifies the strength and direction of the relationship between two variables, excluding the effect of other variables. •    Residuals: The differences between observed values and the values predicted by a regression model, used to assess model accuracy. •    Standard Error of Estimate: Indicates the accuracy of predictions made by a regression model, showing the average distance between observed and predicted values. •    Coefficient of Determination (R²): A measure of how well the independent variables explain the variance in the dependent variable. 8.9    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Mill, F. C. (1955). Statistical methods. Pitman and Sons: London. 8.10    Terminal Questions Q1. From a triradiate distribution, the following correlation coefficient are obtained r12 = 0.7, r13 = 0.6 and r23 =0.4 Verify that 1 – R21.23= (1 – r212) (1 – r213.2) Where r1.23 and r13.2 are multiple and partial correlation coefficient respectively. ***** UNIT-9PROBABILITY AND PROBABILITY DISTRIBUTION STRUCTURE 9.1    Introduction 9.2    Learning Objectives 9.3    Probability Self-Check Exercise-1 9.4    Approaches of Probability 9.4.1    Classical Approach 9.4.1.1    Difficulties in Classical Approach 9.4.2    Relative Frequency Approach 9.4.3    Subjective Approach Self-Check Exercise-2 9.5    Permutations and Combinations Self-Check Exercise-3 9.6    Probability Theorems 9.6.1    Addition Theorem 9.6.2    Multiplication Theorem Self-Check Exercise-4 9.7    The Binominal Expansion Self-Check Exercise-5 9.8    Summary 9.9    Glossary 9.10    Answers to Self-Check Exercise 9.11    References/Suggested Readings 9.12    Terminal Questions 9.1    Introduction Dear Student, According to the syllabus, you are required to have an elementary knowledge of probability. In order to understand the normal, probability curve, some understanding of probability is needed. Knowledge of normal curve itself is important to understand the theory of sampling. 9.2    Learning Objectives After going through this unit, you will be able to •    Understand the basic concepts and different approaches to probability. •   Learn the principles of permutations, combinations, and key probability theorems. •    Explore and apply the binomial distribution in various contexts. 9.3    Probability The idea of probability has great importance in many decision making problems. Most of the people like an HMT watch because the probability that it gives correct time is very high. A person will assign very low probability to the event that an Usha Machine selected from a lot will be defective. Such probabilities help persons in making decisions about the purchase of articles. Manufacturers also benefit from the idea of probability. Hence probability theory help in various decision making situations. Self-Check Exercise-1 Q1. Define probability and explain the importance of this concept in statistics. 9.4    Approaches of Probability (i)    Classical Approach (ii)    Relative frequency Approach (iii)    Subjective Approach 9.4.1    Classical Approach Definition of an experiment can result in N different, equally likely outcomes and NA of these outcomes correspond to event A, then the probability of. event A is NA P(A) = N Let A be event not A. Then probability of event not A is P(A) = N-NA N NA = 1- N 1 – P(A) Thus P(A) + P(A) = 1 The word experiment is used to describe, any act that can fee repeated under given conditions. The experiment may be tossing of one or more coins rolling die or drawing a card from a pack of 52 cards. The coin may land either head or tail. These are the outcomes of the experiment of tossing a coin. Each outcome is called a simple event. A simple event is the outcome of an experiment which cannot be decomposed into a combination of other events. Suppose an experiment consists of casting a pair of dice, in this case, sample space consists of 36 different elementary events. This is shown in this table. No. of spots in a die I | Y/X | 1 | 2 | 3 | 4 | 5 | 6 | |---|---|---|---|---|---|---| | 1 | (1, 1) | (1, 2) | (1, 3) | (1, 4) | (1, 5) | (1, 6) | | 2 | (2, 1) | (2, 2) | (2, 3) | (2, 4) | (2, 5) | (2, 6) | | 3 | (3, 1) | (3, 2) | (3, 3) | (3, 4) | (3, 5) | (3, 6) | | 4 | (4, 1) | (4, 2) | (4, 3) | (4, 4) | (4, 5) | (4, 6) | | 5 | (5, 1) | (5, 2) | (5, 3) | (5, 4) | (5, 5) | (5, 6) | | 6 | (6, 1) | (6, 2) | (6, 3) | (6, 4) | (6, 5) | (6, 6) | Suppose event A is defined as throwing a total of 7 point. In this case, event A is a composite event. The event A consists of 6 simple events, (6,1) (5,2) (4,3) (3, 4), (2,5), (1, 6). If dice are fair, then all the 36 events are. equally, likely. Now what is the probability of occurrence of event A. Hence NA = 6 and N = 36. NA ∴ P (A) = N 61 36 = 6. Note that probability of Not—A is P(A) = 1 — P(A) = 1 — 1/6 = 5/6 Note that if a single die is rolled, then there will be six points in the sample space. The probability of each will be 1/6. Since there are there even numbers (2,4 and 6) and three odd numbers (1,3 and 5), hence P (even number) = (odd number) = 1/2. 9.4.1.1    Difficulties in Classical Approach : (i)    One difficulty in regard the phrase “equally likely” We know that as long as a deck of playing cards is well shuffled. One is equally likely to draw any of the 52 cards. Similarly as long as a die is fair each side is equally likely to turn up. But it is very difficult to Verify that these conditions apply in the given instance. In reality the assumption that different outcomes are equally likely is not justified. The assumption is based on abstract reasoning. It is not based on experience. (ii)    One more difficulty arises when number of cases in a trial is infinite. No one can observe infinite number of trials. (ii)    If we use the classical approach, men it will not be possible to assign probabilities to certain events. For example what is the Probability that it will rain during the 48 hours or what is the probability that Ram will become a millionaire within two years? Classical definition breaks down in such cases. Here the events are not the result of an experiment that can be repeated under same conditions. 9.4.2    Relative Frequency Approach : In certain cases relative frequency is taken as an estimate of true, probability. If an experiment is repeated n times under uniform conditions and the outcome A is observed m times. In this case m m relative frequency is n and is taken as probability of event A, i.e., P (A) = n But according to relative frequency approach it is assumed that relative frequency n approach true probability as n becomes large. Thus probability of an event. A is Defined, as a number which the relative frequency tends to approach as n tends to infinite. The main difficulty that arises with this definition is that it is very costly and time consuming. 9.4.3    Subjective Approach According to this approach, probability measures the confidence that a particular individual has in the truth of a particular proposition. On the basis of this definition different persons may 140 assign different probabilities to an event by using the same evidence. The probabilities assigned are not necessarily based on abstract reasoning. The subjective approach has become important in recent years. A person using this approach can talk about probability of railway strike the next month or the probability that it will rain in the next 48 hours. Thus a person using subjective approach can assign probabilities to many, event that, from die classical view point, do not have probabilities associated with them. Consider another case where a person states that the probability his firm, will fail is 5 percent. Such a statement has no meaning from the point of view of relative frequency approach. But it has meaning and utility according to subjective approach. As pointed out earlier, the main defect of this approach is that different persons may assign different probabilities to the same event. This is due to personal prejudices. It may be pointed out that no single definition of probability is completely satisfactory. If one approach is suitable in some cases, the other approach may be more suitable in other cases. Suppose you toss a coin. What is the chance of the coin falls with the tail upwards? I think you know the answer to this question. The chance is 2 . Suppose you cast a die. (dice) A die has six faces. What is die chance that the die will fall with 6 upwards. Again you know the answers. The 1 chance is 6 Well this chance is probability. We can also see that odds are 1:5 in favour of getting 6 or 5 :1 against our getting 6. Let us put this common sense knowledge in more concrete and mathematical language. If an event can happen in m ways and fails to happen in n ways and each of these ways is m equally likely, the probability of the chance of i s happening is P = m + n and that of its not m m happening for failing to happen is q = 1- p = 1- m + n = m + n . In other words the chances are m to n that even will happenor n to m that the event will not happen. Let us take the case of the casting of a die, with faces having 1, 2, 3, 4, 5 and 6 dots inscribed on them. Let us call the occurrence of in a throw as the event. The probability of any of the faces turning up on a throw, is . 6 Hence probability of 6 occurring i.e. p is and the probability of its not occurring 6 11 i .e q is 1- 6 - 9. This is because the sum of the probability of the event happening (p) and of its not happening (q) must be unity. Hence p = q = 1or p = 1 — q = 1 — p. The probability that an event will happen is obtained by dividing the number of ways favourable to the happening of the event by the total number of ways in which the event can happen and cannot happen. You draw a card from a pack of cards. What is the probability, it is a queen of hearts ? ...... i-r4_ 5 = 2 I. What is the probability that it is a queen ? .... I 5 = 2 I What is the probability that it is a (2_6] black card ? I 5 = 2 I. What is the probability that you will get first Prize in a state lottery in which 1 there are 10,000,000 tickets ?.... 10,000,000 or almost 0. (Assuming you purchased only one ticket). In order to find the probability of an event happening, you have just to know the number of ways in which the event can happen (m) and the total number of ways in which the event can happen and not happen (m+n). Self-Check Exercise-2 Q1. Explain different approaches of probability a)    Classical Approach b)    Relative Frequency Approach c)    Subjective Approach 9.5    Permutations and Combinations In order to find the numbers of ways in which an event can happen or cannot happen, a knowledge of the algebraic concepts of permutations and combinations is required. Take the letters ABC. The number of ways in which these letters can be arranged by taking all three at one time is called the number of permutations. The permutations are 6 in the case and they are ABC, ACB, BAC, BCA, CAB, CBA. Although there aredifferent arrangement or permutations, they are one combination of selection.. ABC is not a separate combination or group, from ACB although they are different arrangements. If we want to know the number of ways in which n things can be arranged and permuted by taking r things at a time, the number can be found with the help of a formula. n P = n! r   (n - r)! npr stands for permutations of n things taking r at a time, n ! is called n factorial (or ! is the factorial sign.) n ! = n. (n-1) (n-2) .... x 2 x 1 6 ! = 6 x 5 x 4 x 3 x 2 x 1 2 ! = 2 x 1; 1! = 1; 3 = 1 If four letters ABCD are to be arranged :3 at a time: 4 Pa (4 - 3)! 4!      4 x 3 x 2 x 1 1 = 24 If four letters ABCD are to be arranged four at a time; 4! 4 P4 0 ! 4 x 3 x 2 x 1  _ --------= 24 (.-.0 ! = 1) Combinations or selections are different, from permutation or arrangements. The number of combination of n things taken r at a time is given by the formula. n n Cr =    Pr r! n! (n - r)! r If three letters are to be selected out of the four letters. ABCD. What shall be the number of these selections? 4      4P3         4 !         (4.3.2.1) Co 3 r!     (4-3)! 3!     1.3.2.1 Example 1. In how many ways can 4 persons be selected put of 7 ? (It is a question on. combinations. The arrangement of the 4 selected persons is not important here). The required number of ways or combinations is 7 7C _   1 4 _         7 ! 4’i7_ (7 -4)! 4 ! 7 x 6 x 5 x 4 x 3 x 2 x 1 (3 x 2 x 1) x (4 x 3 x 2 x 1) = 35 Example 2. In how many ways can 4 persons sit on 7 chairs ? Here the order or arrangement of sitting is also important. Itis a question or permutations. 7P - --— 4 (7 - 4)! 7 x 6 x 5 x 4 x 3 x 2 x1 3 x 2 x1 = 840 After this brief discussion on permutation and combination let us now come back to the consideration and calculation of probability. Example 3. From a bag containing 5 red and 6 black balls three balls are drawn at random. What is the probability they all three balls drawn Would be red ? m We know that p = m + n We have to find values of m and m + n, m standing for the number of ways favourable to the event i.e. the number of ways in which 3 red balls can be drawn and m + n for the total number of ways in which three balls can be drawn whether they be red or not, now 3 red can be drawn out of 5 red balls in 5C3 ways. a m = 5C3 3 balls can be drawn out of 11 ball’s in 11C3 ways. a m + n = nC3 . _ m _ 5C3 P m + n   11C3 3! 3! 11 3 3! 5 V 3 ■ 7 x V  r3 7 5 !        (11 - 3) ! =------x ------- (5 - 3)!     (11)! 5 x 4 x 3 x 2 x 1          8 x 7 x 6 x 5 x 4 x 3 x 2 x 1 ------------- x ---------------------------------- 2 x1         11 x 10 x 9 x 8 x 7 x 6 x 5 x 4 x 3 x 2 x 1 2 = 33 Simple and Compound Events When we discuss the probability of the happening or not happening of a single event, the event is called a simple event. All the examples, taken so far are examples of simple events. When two or more events, happen together they are called compound events. If we want to find, the probability of a card drawn from a pack of cards being a queen, the event is simple. If we want to find the probability of the first card being a queen and the second card a king the event is a compound event. Events can either be independent or dependent If we draw a card from the pack and put it back before drawing the next, the second sub-event is not’ affected by the first draw.” The events are independent. If we do replace the card the probability of the second draw is influenced by the first draw such events are dependent events, the tossing of queen and throwing of die shall always be independent events. Drawing of balls or cards may or may not be independent events, depending on whether replacement were or were not made after the first draws. Self-Check Exercise-3 Q1. In how many ways can 4 persons be selected out of 7? Q2. In how many ways can 4 persons sit on 7 chairs? 9.6    Probability Theorems There are two very important theorems of probability which are applicable in the case of compound events. They are (i) Addition Theorem, (ii) Multiplication Theorem. 9.6.1    Addition Theorem: If an event can happen in different ways which are mutually exclusive, the probability that it will happen, is the sum of the probabilities of its happening in these different ways. The theorem is simple and self evident. If you throw a die what is the probability “that in the first throw either 6 comes upwards, or 5 comes upwards ? The two events are mutually exclusive. It is a question of either this, or that the probability is the sum of the separate probabilities and is 112 --+ — — — 666 The addition theorem will hold good only if (1)    events are mutually exclusive (2)    mutually exclusive events belong to the same set. The condition (1) is satisfied if events belong to Either r’ category. For condition (2) just 4 read this statement. The probability that a card is a queen is 52 . The probability that 6 comes upwards, in a throw of a dice is 1 6. What is the probability that either a queen is drawn on 5 comes up? Not + . The addition theorem cannot be applied. The items do not belong to the same set. 52   6 9.6.2 Multiplication Theorem If a compound event is madeof a number of separate but not mutually exclusive, sub-event the probability of the occurrence of the compound event is the product of the probabilities of each of the sub-event happening. If a dice is thrown 2 times what is the probability that 6 comes up in the first throw and 5 in the second throw? The probability is 1 — X 6 11 — = — in multiplication theorem we do 6 36 not consider the probability of either this or that event, but of this as well as that event. Suppose there are two events 1 or 2, p1 and p2 are the probabilities of their happening and q1 and q2 the probabilities of their not happening respectively. Let m be the number of ways in which, the first each can happen and n in which it cannot happen. Let m1 be the number of way in which the second event can happen n1 in which it cannot happen. mn pi - T- qi - "v m + n m + n m1 n1 P2 = T-r q2 = T-r m + n m + n The number of ways in which the first as well as the second event can happen: mm1 or     pp2 = 7     77 i n (m + n) (m + n ) The number of ways in which the first event can happen and the second event can be happen. mn1 or pp - 7 : \ /i , n (m + n)(m + n ) nm Similarly, qq - 7    77 n ■’              (m + n)\m + n ) nn qq - (m + n )(m + n ) The following illustration of drawing a card three times from a pack of cards may Help you, decided when the addition theorem is to be used and when the multiplication theorem. The card is 4 replaced after each draw. What is the probability that the card is a queen in the first draw...     = 4 3 . What is the probability that it is a king in the second draw ? Again 52 = 3 . That it is an ace in 41 the third draw ? Again 52 = 13 . What is the probability that the card is either 5 queen in the first draw or a king in the second draw. The composite probability should be greater than the separate 11 probabilities. It is 13 + 13 . What is the probability that the card is queen in the first draw or a king in the second draw or an ace in the third draw? It is greater still and is — + — + — + ~ and so on. n e secon raworanace n e r raw sgreaers an s 13 13 13 13 an so on. Probability of ship A arriving safely = Probability of ship B arriving safely = Probability of ship C arriving safely = | 2 | = 2 | |---|---| | 2 + 5 | = 7 | | 3 | = 3 | | 3 + 7 | = 10 | | 6 | 6 | 6 +11 17 The chance that all the three (A as well as B as well as C) arrived safely 23 — x- 7 10 6    18 x — =-- 17   595 . What would be the probability that at least one of these ships arrives safely? Find the solution yourself and then compare with the following: Probability of A not arriving safely = 5 7 Probability of B not arriving safely = 7 10 Probability of C not arriving safely = 11 17 :. Probability that none of them arrives safely 5 7 x 7 — x 10 11 17 11 34 :. Probability that at least one ship of three arrives safely 1123 = 1= — 3434 2 4 ? (A, B, C) all reach : A, B Can you imagine another way of arriving at the answer 3 reach; C does, not; A, C reach, B does not; and so on. Add all such mutually exclusive probabilities. The method is however, lengthy and time consuming). Example 6. The probability that A can solve a problem is 2 3 the probability that B can solve it is. If both try, what is the probability that the problem is solved. The probability that A cannot solve the problem :— 3 The probability that B cannot solve the problem: — 1 - 4 5 (H.P. University Feb. 1972) 1 3. 1 5. The probability that both A as well as B cannot solve, the problem (i.e. if A and B both try, the problem, will remain unsolved) = 11  1 - x — = — 3 5    15. :. the probability that if both A and B try the problem the problem will be solved (i.e. if A and B both try, he problem will not remain unsolved) (i) 1    14 1 -15 - 15 Ans- Alternatively the problem could be solved like this— 2 The probability that both A and B solve the problem 3 x 4 = 8 5 = 15’ (ii) 2 The probability that A solves and B does not solve the Problem 3 12 x — = — 5    15. (iii) 14 Probability that A does not solve and B solve the problem 3 x 5 4 15. The problem will be solved if either (i) or (ii) or (iii) above happens (mutually exclusive and events). 8    2    4   14 : the probability the problem will be solved - — x — - — - — Ans. Example 7. Two balls are drawn from a bag containing 9 red and 14 white balls. Find the chance that they are — (i)    both of the same colour. (ii)    each of a different colour. Solution: (i)    The requirement is met if both the balls are either red or white. Two red balls can be drawn out of 9 red balls in 9C2 ways. The total number of balls in the bag is 9 + 14 = 23, and two balls can be drawn out of 23 balls in 23C2 w ways. : the probability of drawing of 2 red balls is 9 x 8 9    C2 _ 2 x1 _ 9 x 8 _ 36 23    C2 = 23x22 = 23 x 22 = 253 2 x1 the probability that both the balls are of the same colour (either both red or both white) 36    91    127 --- + --- — --- 253   253   253 (ii)    The requirement is met if one ball is red and the other is white. One red ball can be drawn out of red balls in 6C1 or 9 ways. One white ball can be drawn out of the white balls in 14C1 or 14 days. :. total number of ways favourable to the drawing or one red and one white ball. = 9 x 14 = 126. The total number of ways in which two balls can be drawn out of a total of 23. 23 C2 23 x 22 2 x1 — 253 : the chance that one ball is red and the other white 126 253 . Self-Check Exercise-4 Q1. Explain the additive theorem of probability. Q2. Explain the multiplication theorem of probability. 9.7 The Binomial Expansion The following are some binomial expansion. (p + q)1 p + q (p + q)2 = p2 + 2pq + q2 (p + q)3 = p3 + 3p2q + 3pq2 + q3 You may be remembering these expansions from your high school days. But suppose you were to write down the value of (p + q)12. You may not be able to expand this binomial so easily. We shall make use of a general formula writing such expansions. (p+q)n = pn+nc1 p n+1 q1+nc2 pn-2 q2+nc3 pn-3 q3 + qn n(n -1)          n(n - 1)(n - 2) pn + npn-4 q1+         pn-2 q2 +                pn-3 q3 ………. qn 1 x2                1 x 2 x 3 (Compare these terms with the terms, of Newton’s forward formula for interpolation). Now we can write the terms of say (p+q)4 more easily. (p+q)4 = p4+ 4c1 p 4-1 q1+4c2 p1-2 q2 4c3 p4-3 + 4c4 p4-4 q4. = p4+ 4p3 q = 6p2 q2+4 pq3 + q1. 1 —--+ 16 1 1 1 + I 2 2 J 4 + + 4 6 16 16 41 + 16   16 1 x —+ 6 x 2 1 + 4 x — x 2 Such binomial expansions are of immense use in finding certain probabilities. Suppose P stands for the probability of the happening of an event in a single trial, q for the probability of the event not happening in a single trial. If two trials are made, the probability that the event will happen both the times, one time and zero time (i.e. is not happening at all) are the term are the teams of the expansion (p+q)2,i.e p2, 2pq and q2 respectively. If 4 trials are made, the probability that the event will happen all the 4 times, 3 times, 2 times, 1 time and 0 times are the respective terms of the expansion (p+q)4, i.e, p4, 4P3q, 6p2q2, 6pq2, and q4, respectively. Example 8: Let us take specific cases now. If tossing of a coin is called a trial and coming tip of a head in a toss of a coin is called success or an event then the probability of event happening of the probability of success in a single trial or p — 21, q then is automatically 1 . If the coin is 2 tossed 4 times, (i.e. 4 trials are made) what is the probability that we shall get heads in all the 4 tosses, in 3 tosses, in 2 tosses, in l toss, and in no toss at all? These are given simply by the successive terms in the expansion of the binomial (p+q)4 or i i 4 + I 2 2 J The terms are 14641 16,16,16,16,16. Example 9 : Suppose a die is cast and coming upwards of 1 or 2 is called success What would be, the value of p? 21        12 It will be 6 or 3 q, therefore will be 1 - 3 - 3. Suppose the die is cast 4 time. What is the probability that we shall get success all the 4 times. What is the probability that we shall get success all the 4 times, 3 times, 2 times, 1 time and no success at all ? (Remember that success is coming up (I* of l or 2 whosep= 3). The probabilities are the respective terms of (p-q)4 or I 3 + 3 J .They are:- 2 x —+ 6 x 3 3 3 3 1 Y 2 .    1 - I x —+ 4 x - x ( -11 181 J 3   34   32 + — + — + — + 81   31   81 16 81. Probability of getting 4 successes = Probability of getting 3 successes = 1 16 8 81 22 Probability of getting 2 successes = 32 Probability of getting 1 successes = 16 Probability of getting 0 successes = From specific we move back to the general p is the probability of an “event” happening or of “success” in a single trial q has the opposite meaning. The trial is repeated n times, what are the probabilities of the event happening n times, n-1 times n-2 times.......2 times, 1 time, 0 times ? They are terms of the expansion of the binomial (p+q)n. Let us write these terms. (Plus signs here, are not important) (p+q)n pn + nC1 pn-1 + nC2 pn-2n2 +……. nCn-2 p2 qn-2 + nCn-1 p3 qn-1 + qn From here, we can write down. In n trials. The probability of getting n successes = pn The probability of getting n-1 successes = nC1 pn-1 q1 The probability of getting n-2 successes = nC2 pn-2 q2 The probability of getting r success = nCp-r pr qn-r The probability of getting successes = nCr pr qn-r The probability of getting 2 successes = nCn-2 p2 qn-2n C2 p2 qn-2 The probability of getting 1 success = nCn-1 p1 qn-1 npqn-1 The probability of getting 0 success = nCn pn-n qn = qn Do a little brain twisting to understand how the probability of getting r success which is nCn-r pr qn-r or nCr pr qn-r fits into the above scheme. If you find it is difficult to start from the top, start from the bottom. Probability of getting I success is nCn-1 p1 qn-1; of 2 successes is nCn-r p2 qn-2; of r successes it should be nCn-r pr qn-r. Since nCn-r is the same thing as nCr pr qn-r. Let us now write a general rule. The probability of the happening of an event in one dial being known, the probability that the event will happen exactly r times in n trials is = nCr pr qn-1 whether p stands for the probability of its happening and q for the probability of its not happening in a single trial. The probability that an event will happen at least r times in n trials is pn+ nC1 pn-1 q1 + nCn-r pr qn-r. Because the probability of an event happening exactly n times is pn, exactly n-l times is nC1 pn-1 q1exactly n-2 times is nC2 pn-2 q2 …... and exactly r times is pn+ nCn-r pr qn-rand because all of them are mutually exclusive and in any of them the event happens at least r times, therefore the probability that the event happens at least r times issimply the sum of these different probabilities. Example 10. Find the chance of getting exactly 5 heads in 6 throws of an unbiased coin. (H.P. University Feb. 1972) Solution: 1 2 Probability of a head or p = 1 q = 2 Number of trials or n = 6 Number of successes desired or r = 5. the chance of getting exactly 5 heads in 6 throws of the coin = nCr pr qn-r. 6 C5 1Y r i Y 2 J I 2 11 = 6 x — x — = 32 2 63 64 = 32 Ans’ Example 11. Find the chance of getting (i) at least 5 heads, (ii) at least 4 heads, in six throws of unbiased coin. Solution (i)    If p is the probability of success in a single trial, the probability of getting at least r successes in n trials = pn + nC1 pn-1 q1 + nC2 pn-2 q2 + …. nCn-r pr qn-r probability of getting a head in one throw of a coin or 21. 1 . q = 2 Probability of getting at least 5 heads in 6 throws of a coin 6    51 = 1 -1 + 6CI -1 I -1 16 J 112 J ( 2 J 167 — -- +--— -- 64 64 64 . (ii)    Probability of getting at least 4 heads in 6 throw of a coin 6   51    12 — I 11 + 6C 111 111 + c 111 111 12 J 11 2 J I 2 J 21 2 J 12 J 1    6    15 22   11 — -- +--+--— -- — -- 64   64 64 64 32 The problem could have been solved as follows also : (i Y 1 (i)    Portability of getting 6 heads: = I 2 I - —. i11* 64 12 J Probability of getting 5 heads = 6C5 p5 q1 . Probability of getting 5 heads. 167 — -- + -- — -- 64   64   64 . Similarly, for the (ii) part. Let us now use the knowledge of probability to study Binomial and Normal Distributions. As we shall see, these distributions are indispensable for analysis and interpretation of data and are the foundation, of sampling methods. In previous lessons we have dealt with many frequency distributions. Those distributions were based on observations or experiments. In binomial and normal distributions we start on certain assumptions and then try to calculate different frequencies. Such distributions which are not based on actual observations or experiments, but are calculated mathematically on the basis of certain assumptions, are called “Theoretical Frequency Distributions.” We are studying only two binomial and normal distributions. A third important distribution. Poisson distribution is not being discussed here. The Binomial Distribution Suppose two coins, a,b are tossed simultaneously. There are 4 possible ways is which they could fall: Possible outcomes on coins ab (1)   HH (2)    TH (3)    HT (4)    TT Probability _ *2 = P2   = P2   — 24 112 = pq3 = 2pq  - 2 x 3x 2 = pq 1 the first outcome is one out of four, or 4 . Second and third outcomes represent 1 head and 1 tail, the 2 probability of one head and one tail is, therefore two out of 4 or 4 . Fourth outcome is both tails; the 41. probability of both tails is one out of four or Here the probabilities of 2 heads, 1 head and 1 tail, 2 tails have been written, as p2, 2pq + q2. These terms are simply the expansion of (q + q)2 In this illustration, p = q = 21. That is why the probabilities of various outcomes are 1 1 ^ + I 2 2 J 121 = - + - + — 444 Similarly if we toss three coins simultaneously, there shall be 8 possible outcomes. The probabilities of 3 heads, 2 heads, 1 head, 0 head (i.e. 3 tails) can be found simply by binomial expansion. But let us write again all the 8 possible outcomes to satisfy ourselves that the binomial, expansion really gives correct results. | Possible Outcomes on coins | Probability | |---|---| | | a | b | c | | | (1) | H | H | H | = p2   = p3 | | (2) | T | H | T | = p2q | | (3) | H | T | H | = p2q =3p2q | | (4) | T | T | H | = pq2 | | (5) | T | T | H | = pq2 | | (6) | T | T | H | = pq2 =3pq2 | | (7) | H | T | T | = pq2 | | (8) | T | T | T | = q3   = q3 | Probabilities of 3 heads, 2 heads; 1 head and 0 head are given by the successive terms of the 1 binomial expansion (p+q)3 – p2 + 3p2q + 3pq2 + q2. Here p = q = 2 . Therefore, probabilities of 3, 2, 1 and 0 heads are respectively 1 8 3 8 We can see easily that these results are correct. Out of the 8 possible outcomes, there is only one way in which three heads fall upwards. The probability is only one out of eight or . There are 8 three outcomes where coins fall with, two heads upwards. Therefore, the chance of two heads is three out of eight or . And so on. 8 By the way expansions of type (p+q)nare called binomial because there are two terms p and q which are to be expanded and bi means two bi-cycle, bilateral, binoculars. If several terms were involved the expansion would have been called multinomial. Suppose the experiment of tossing three coins simultaneously is repeated 200 times. How many out of the 200 experiments or tosses or trials can we expect 3 heads, 2 heads, 1 head and, 0 1     1 ^3 + I 2 21 head ? This shall be given by the successive terms of 200 13 3 1A = 200 + + l| 8 8 8 8 J • = 25 + 75 + 75 + 25. The probable frequencies of heads, 2 heads, 1 head and 0 head are 25, 75, 75 and 25 respectively. One thing should be clear now. If we want to find the probable frequencies of various outcomes in a given number of experiments or trials, we can use the expression. N(p+q)n where capital N stands for me number of times the experiment was repeated and n for the number of independent events. Example 1: Three dice are thrown 27 rimes. If coming upwards of 3 or 4 is considered to be a success, find the expected frequencies of 3 successes, 2 successes and 0 success. 3 or 4 is considered to be success. 21 :. p or probability of success = -7-7 63 q = 1 - 2 3 The required expected frequencies are given by N(p+q)n or 27 6    12    8 A 1+ 27 27 27 J - 27 33 21 — + 3 x — x + = 1 + 6 + 12 + 8 Therefore, once out of the 27 trials we can expect all the three successes, 6 times 2 successes, 12 times 1 success, 8 times no-success. Self-Check Exercise-5 Q1. Explain the general formula for binomial expansion. 9.8    Summary This unit covers fundamental concepts of probability and probability distributions. It introduces probability, explores various approaches, and discusses permutations, combinations, and key theorems. The unit also explains the binomial distribution, providing a comprehensive understanding of these essential statistical topics. 9.9    Glossary •    Binomial Distribution: A probability distribution that summarizes the likelihood of a value occurring in a specific number of trials. •    Classical Approach: A method of determining probability based on equally likely outcomes. •    Probability: The measure of the likelihood that an event will occur. •    Relative Frequency Approach: Probability is determined by the ratio of the number of times an event occurs to the total number of trials. •    Subjective Approach: Probability based on personal judgment or experience rather than objective calculations. 9.10    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 9.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 9.4. Answer to Q2. Refer to Section 9.4.1, 9.4.2 and 9.4.3. Self-Check Exercise-3 Answer to Q1. Refer to Section 9.1. Self-Check Exercise-4 Answer to Q1. Refer to Section 9.6.1. Answer to Q2. Refer to Section 9.6.2. Self-Check Exercise-5 Answer to Q1. Refer to Section 9.7. 9.11    References/Suggested Readings 1.    Emory, C. W. (1976). Business research methods. Richard D. Irwin Inc. 2.    Plane, D. R., & Oppermann, E. B. (1986). Business and economic statistics. Business Publications Inc. 9.12    Terminal Questions Q1. Explain the term probability. State and prove the addition and multiplication theorems of probability. Q2. Two balls are drawn from a bag containing 8 red and 7 white balls. Find the chance of (i)    they are both red, (ii)    they are both white. (iii)    one is red and other white. ***** UNIT-10MATHEMATICAL EXPECTATION STRUCTURE 10.1    Introduction 10.2    Learning Objectives 10.3    Univariate Probability Distribution Self-Check Exercise-1 10.4    Mathematical Expectation Self-Check Exercise-2 10.5    Parameter and Statistic 10.5.1    Estimation of Parameters 10.5.1.1    Problem of Estimation 10.5.1.2    Point Estimation and Interval Estimation 10.5.2    Properties of Good Estimation Self-Check Exercise-3 10.6    Interval Estimation for Population Mean Self-Check Exercise-4 10.7    Confidence Interval for Proportions Self-Check Exercise-5 10.8    Summary 10.9    Glossary 10.10    Answers to Sel-Check Exercise 10.11    Reference/Suggested Readings 10.12    Terminal Questions 1 0.1 Introduction Dear Student, You have already studied Probability; and Theoretical distributions (Normal and Binomial distributions). Elementary knowledge of these two topics particularly, that of the Normal distribution will be frequently needed to understand this topic. Besides we should also be conversant with the concept of Mathematical expectation about which you have not been told so far. So before we switch over to the main topic of Estimation we will talk a little about Expectation first. 10 . 2 Learning Objectives After going through this unit, you will be able to •    Understand the concept of mathematical expectation and its application in probability. •  Learn the methods for estimating parameters and constructing confidence intervals. •   Learn point estimation and interval estimation. 10.3.    Univariate Probability Distribution If the variables X assumes X1 X2 ...... Xn (denoted by X1, for i = 1, n   N) valaes with probabilities pr pn   P1 (denoted by p1 for i = 1, 2   N) respectively, corresponding to the 156 N exhaustive and mutually exclusive cases the X is called a Chance or Random variable and the set of values Xi together with their probability p constitutes what is called. Univariate Probability, distribution of the variable. For example if a fair die is cast and X.....denotes the number of the die, then Probability Distribution for the variable X can be given as: X: 1 23456 111111 : 666666 | Similarly, if in a throw of a pair of fair, dice and | denotes the sum of the number on two dice | |---|---| | the Probability. X: | Distributions will be: 234 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | | P: | 123 | 4 | 5 | 6 | 5 | 4 | 3 | 2 | 1 | | 36   36    36 | 36 | 36 | 36 | 36 | 36 | 36 | 36 | 36 | (ii) Self-Check Exercise-1 Q1. Explain Univariate Probability Distribution. 10.4.    Mathematical Expectation: Let φ xy be a function of X such that it takes values X1 [i = 1, 2   N) i.e. φ (X)1 φ(X) φ (Xn) when X takes values Xi (i = 1, 2 .......... N) i.e., X1, X2 ......... Xn with probabilities P1 = 1: 2 ....... N) i.e. P1P2....... Pn, the EXPECIED OR PROBABIE value of φ (x) denoted as E (φ {x} is defined as {φ(x)} = ΣP1 + (φxi) = Pi φ(X1) + Pφ (X2) + ....... + Pn φ (Xn) Where ΣP1 P1 P2 ×   n = 1 (Since X1. i = 1, 2 .....N) are mutually exclusive and exhaustive cases. The expected values of X in cases (i) and (ii) above, 67 + 6 × 1 = 2 and E (X) = ΣpiXi = 2 can be given as E (X) = ΣpinXi1 × + 2 × + 3 × 666 123 36 + 3 × 36 + 4 ×36+ 12,     = 7 resp. 36 Self-Check Exercise-2 Q1. Explain the concept of mathematical expectation. 10.5.    Parameter and Statistic: Any function of observation in the population is defined as parameter and that in the sample as statistic. 10.5.1 . Estimation of Parameters: 10.5.1.1    Problem of Estimation: Let the population under investigation have the density function (probability distribution f(x) :θ1; θ2; θ3; ......φm) where x is the variable and θ1, θ2 ..... θm are m 1 21 (x- µ)2 σ parameters of the distribution. For example the density function of Normal dist. σ∠(π) -e can be written as f(x; µ, σ). Suppose that a sample of observation x1  x2  xn is available. The estimation problem is to define estimator’s for one are more of the parameter θ1; θ2; θ3  θm as function’s of the sample observations x1  x2  x3  xn. It is customary to represent the as θ1; θ2;   θm for the parameter θ1; θ2;  θm respectively. For one parameters there can be number of estimators. The main part of the probe of estimation lies in selecting one estimators out of many of the parameter and under consideration in such a manner that its distribution is concentrated around the true value of the parameter θ as closely as possible. Let us try to understand this concept clearly. Suppose, the population consists of N observation X1 X2 ........Xn. If we want to draw a sample of size n from this population to estimate the value; of x which is unknown, it can be done in Nen different ways. In other words there will be Nen different sample of size n each. Suppose x = n ∑x1 i=1 is defined as an estimator of Parameter x. Each one of the Nen samples will give one estimate for the estimator x. Let these estimates by represented by x1, x2, x3  xNcn for sample Nos. 1, 2 Ncn, respectively. This forms a distribution of x values. These distributions should have concentration round about the true value, of x. 10.5.1.2    Point Estimation and Interval Estimation (i)    Point Estimation:— As the name indicates, estimates is point or a single number which is calculated as function of the observation in the sample. For example sample mean is a point estimates of the population mean. In case of point estimates it is not possible to indicate, the amount of confidence that may be placed on it. (ii)    Interval Estimation:— A single value ‘d’ or point estimate is quite unlikely, to concede with the true value of the parameter. It is, therefore, considered more appropriate to obtain a range or values or an interval in which the true value of the parameter may be expected to lie with some definite probability or degree, of confidence. This interval is called “Interval Estimate” or “Confidence Intervals” and the probability or degree of confidence is called “Confidence coefficient”. Obviously corresponding to different confidence coefficients, there will be different “Confidence intervals” Higher the value of probability or Confidence Coefficient wider will be the confidence interval there will be one confidence coefficient associated with it. 10.5.2    Properties of Good Estimators We have already, mentioned in 4.1 that out of a large number of possible estimators for a parameters. We have to select one those distribution is most closely concentrated round about the true value of the parameter. For presence of the characteristic, an estimator is tested by examining the presence of the following properties in it. (i)    Unbiasedness (ii)    Efficiency (iii)    Consistency (iv)    Sufficiency Here, we will discuss (1) and (ii) only. (i) Unbiasedness An estimator θ for θ a parameter θ is said to be unbiased if its expected value is equal to θ i.e. if E(θ) = θ 0. then θ is an unbiased estimator of θ. In other words, if the value of the estimator is calculated for all possible samples and if on the average estimator assumes, the true value of the parameter, then the estimator is said to be unbiased estimator. For example; η Let X = C∑=1 xi/η be the estimator for the population mean where x1, x2 ,   xn are the Ncn estimator for Ncn possible samples of size n Ncn of the estimator x . It can be shows that x1 +x2 ..+ x . Ncn = Nc Ncn x Hence, we say x i.e. sample mean is an unbiased estimate of population mean x . This can be proved mathematically also. A When E = θ = θ: 0is said to be biased estimator, η ∑(x1-x2) i=1 η For example σˆ = ∑ (x1-x2) is a biase i=1 η ∑(x1-x2) estimator of σ= i=1         as E (σˆ )3= σ2. n η ∑(x1-y)2 but = 1=1         is an unbaised n-1 estimator of σ2 as E (s2) = σ2. A                                                                              A If E(θ) > θ estimator is said to be positively biased and if E(θ) < θ, estimator is said to be negatively biased. Efficiency— If there are two unbiased estimators θˆ1 and θˆ2 of the same parameter i.e. E (θˆ1 )= 0 and E (81 j = θ, the comparison between the two is made on the basis of their variations. The estimator with smaller’ variance is said to be more efficient that the other with higher, variance. Thus if Var (θ1 ) Var θ2 then is said to be more efficient than θ2 and vice versa, Efficiency of θˆ1 respect to θˆ2 is defined as E= Var(θˆ2) × 100 Var (θ1 ) If there aremore than one unbiased estimators of a parameter then the one with the smallest variance, is sailed “most efficient estimator”. If the criteria of goodness of an estimator be efficiency only then the defined most efficient estimator can be named as the best unbiased estimator. Self-Check Exercise-3 Q1. Explain the problems in estimation of parameters. Q2. Explain point estimation and interval estimation. Q3. Discuss the properties of good estimators. 10.6    Interval Estimate for Population mean A point estimate of a parameter is not very meaningful without some measure of the possible error in the estimate, An estimate 0 of parameter should be accompanied by some interval about possibly of the form (0-d) and (0+d) and together with some measure of assurance that the true parameter 0 lies within the interval. “A cost accountant for a publishing company may estimate the cost to be 80 ± 5 Rs. per volume with the implication, the correct cost very probably lies between 75 and 85 Rs. per volume. Suppose, a samples of n observation (x1, x2........... xa) is drawn from a normal population Lx with unknown mean, and known standard deviation cx - — is a point estimate of population mean. n We wish to determine the upper and lower limits which are rather certain to contain the true parameter value between them. We know the y - x - * < cr / Z(n ) will normally distributed withmean 0 and S.D. = 1. Thus Prob. (—1.96 < y < 2.96) = 0.95 or P -1.96 < k x * < 1.96] 0.95 c / Z(u)      J or P -1.96 c ^ kAn) J < (x -1) < 1.96 c kz(n) J J - 0.95 or P - x -1.96 c k A(n) J < * < 1.96 c k A(n) J - x - 0.95 f or P ^ x -1.96 c ^ Zw J< * - 1.96 c kz(n) J ^ - 0.95 or P x -1.96 r c ^ kAn) J < * < x -1.96 r c ^ kAn) J - 0.96 Thus the two limits (x-1.96)a.W(n) and x + 1.96 cN(n) have been obtained which we may say with 65% certainly to contain the true parameter value, between them, (iii) has to be clearly understood, We mean that if samples of size n were repeatedly from the population and if the interval (x-1.96) [σ/√(n) to x=1.96 [σ/√(n) were computed for each sample, then 95% of those intervals, would be expected to contain the true mean. We therefore, have considerable confidence mat the interval x-1.96 σ/√(n) to x-1.96 σ/√(n) contains the true mean. The measure of confidence is 0.96. The interval x—1.96 σ/√(n) to x + 1.96 σ/√(n) is called a 95% confidence interval, the probability 0.95 in this case called the confidence coefficient. We can obtain interval with any desired degree of confidence less than one confidence interval with confidence coefficient 0.93 and 0.99 can be obtained by replacing 1.96 by 2.33 and 2.50 respectively. Example Find 95% 98% and 99% confidence interval, for µ a when a example of four observations (1, 2, 3, 4, 016 and 5,6) has been drawn from a normal population, with unknown mean µ and known S.D. —3. 12+3.4+6.6+5.5 x= = 2.7; α=3 4 σσ and x + 1.96       . ∠(n)              ∠(n) This limits of 95% confidence interval are given by x — 1.96 33 = 2.7 - 1.96 ∠(4) and 2.7 + 10.96 ∠(4) . = - 0.24 and 5.64 and hence 95% confidence interval can be written (-0.24, 5.64). Similarly, find out 98%’ and 99% confidence intervals by replacing 1.96 by 2.33 and 2.58, yourself. The method described above cannot ordinarily be used to find interval, estimate of the mean of a normal population because σ2 is not ordinarily known in such a case σ2 is estimated by Σ(x - x) σ2 = n -1 which is an unbiased estimate of σ2 when the sample is small. Then Σx - µ Sn follows the ‘t’ distribution with (n—1) degree of freedom. This is evident from the property of normal distribution that the area covered between the points= 1.96 and +1.96 and (in terms of Standard Normal Variate) is 0.96 when the total area under normal curve is considered = 1. In this case, the 95% confidence interval is given by [z-t00.5n-1 (s/√n) x + t00.5 n-1 (s/√n)] Similarly to find 98 % and 99% confidence intervals we replace t0.05 n = 1, by t0.02 n-1 and t0 01, n - 1 values respectively. t0.05 is the table value of t at 5% level of significance of (n-1) of you should show to consult at table: Example : Deduce that for a random sample of 16 values with mean 41.5 inches and the sum of squares of deviations from the mean 135 (inches drawn from a normal population, 95%) confidence limits for the mean of the population are 39.9 and 43.1 inches. n = 16, x = 41.5 S (x1 - x)16= 135 ' ' ll ■ ' ' k n 1 J 135 = VI----I = V 9 = 3. k 5 J t-value at 5% level of significance for 15 d.f. i.e. t0.05 15 = 2.13 (from t-table). Confidence limit will be f s ì f s Ì X-t0.05 15 k Zn J and x + t0'05 15 k Zn J f 3 I                        i 3 I or 41.5-(2.13) | —I and 41-5 + (2-13) I 216 I or 39.9 and 43.1 Note. When n > 30, t values can be replaced by 1.96, 2.33 and 2.58 for 95% 98% and 95% conf coefficients respectively. This is so because for large samples, the sampling distribution of mean resembles the normal distribution. Self-Check Exercise-4 Q1. Explain the interval estimate of the population mean. Q2. Find 95 %, 98% and 99% confidence intervals for µ when an example of four observations (1, 2, 3, 4, 0.6 and 5, 6) has been drawn from a normal population with unknown mean µ and S.D-3. 10.7    Confidence intervals for Proportions: Let p' be the proportion in the population. It’s point estimate in the sample can be given by the x sample proportion P =   where x is the number of items processing the attribute under consideration n P - P' and n it the total number of observation in the simple A p—q I is approximately normally distributed I n I with mean 0 and S.D. = 1 for sufficiently large n (q' = 1 - p'). For sufficiently large n the variance p-q n can be estimated pq hence n p -q z{ pq)/Zn} can also be taken as standard normal variate i.e. Similarly 91% and confidence intervals can be obtained by replacing 1.96 by 2.33 and 2.50, respectively. Example: A sample poll of 100 voters chosen at random from all voters in a given district indicated that 55% of them were in favour of a particular candidate. Find 95% and 99% confidence limits for the proportion of voters in favour of that particular candidate if an unlimited number of voters are allowed to cast their votes. x 55 p = — =---= 0.55 n 100 q = 1 - p = 1 -0.55 - 0.45 / (0.55)(0.45) ^ I 100 J So 95% confidence limits will be < 0.55 -1.96Z (0-55)(0-45)^0.55 +1.96 A(Q-55)(0-45)^ . I I 100 J k 100 JJ i.e. (0.45 and 0.65) Similarly you can find out 99% confidence limits by replacing 1.96 by 2.58. Self-Check Exercise-5 Q1. A sample poll of 100 voters chosen at random from all voters in a given district indicated that 55 % of them were in favour of a particular candidate. Find 95% and 99% confidence limits for the proportion of voters in favour of that particular candidate if an unlimited number of voters are allowed to cast their votes. 10.    8 Summary In this unit, we covered univariate probability distributions, mathematical expectations, and the difference between parameters and statistics. We explored methods for estimating parameters, including point and interval estimation and constructing confidence intervals. 11.9    Glossary •   Univariate Probability Distribution: A probability distribution of a single random variable. •    Mathematical Expectation: The weighted average of all possible values that a random variable can take on, with each value weighted according to its probability. •    Parameter: A numerical characteristic of a population, such as a mean or a standard deviation. •    Statistic: A numerical characteristic of a sample, used to estimate a parameter. •   Confidence Interval: A range of values, derived from a sample, that is likely to contain the value of an unknown population parameter. 10.10    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 10.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 10.4. Self-Check Exercise-3 Answer to Q1. Refer to Section 10.5.1.1. Answer to Q2. Refer to Section 10.5.1.2. Answer to Q3. Refer to Section 10.5.2. Self-Check Exercise-4 Answer to Q1. Refer to Section 10.6. Answer to Q2. Refer to Section 10.6. Self-Check Exercise-5 Answer to Q1. Refer to Section 10.7. 10.11    References/Suggested Readings 1    Mason, R. D. (1986). Statistical techniques in business and economics. Homewood, IL: Richard D. Irwin, Inc. 2    Plano, D. R., & Oppermann, E. B. (1987). Business and economic statistics. Plano, TX: Business Publication, Inc. 10.12    Terminal Questions Q1. Suppose 250 randomly selected people are surveyed to determine if they own a tablet. Of the 250 surveyed, 98 reported owning a tablet. Using a 95% confidence level, compute a confidence interval estimate for the true proportion of people who own tablets. Q2. Let X be a random variable defining number of students getting A grade. Find the expected value of X from the given table | X = x | 0 | 1 | 2 | 3 | |---|---|---|---|---| | P(X=x) | 0.2 | 0.1 | 0.4 | 0.3 | ***** p - q' Z{ pq)/Zn} can be given by UNIT- 11 STATISTICAL HYPOTHESIS STRUCTURE 11.1    Introduction 11.2    Learning Objectives 11.3    Statistical Hypothesis Self-Check Exercise-1 11.4    Testing a Statistical Hypothesis 11.4.1    Null Hypothesis and Alternative Hypothesis 11.4.2    Critical Region Self-Check Exercise-2 11.5    Type I and Type II Errors and their Sizes Self-Check Exercise-3 11.6    Level of Significance Self-Check Exercise-4 11.7    Properties of the Normal Curve Self-Check Exercise-5 11.8    Equation of the Normal Curve Self-Check Exercise-6 11.9    Summary 11.10    Glossary 11.11    Answers to Self-Check Exercise 11.12    References/Suggested Readings 11.13    Terminal Questions 11.1    Introduction Dear Student You have leant about the problem of estimation in an earlier section. The text problem under the tide of statistical inference is that of the Testing of Hypothesis’, let us try to understand what is meant by Hypothesis or more precisely Statistical hypothesis. 11.2    Learning Objectives After completing this unit, you will be able to •    Understand the concept and importance of statistical hypotheses. •    Learn how to test a statistical hypothesis using null and alternative hypotheses. •    Understand the implications of Type I and Type II errors in hypothesis testing. 11.3    Statistical Hypothesis: Statistical hypothesis, in the general form is an assertion about the density function or probability distribution of a random variable. Thus the assertion an example of Statistical hypothesis. For example, let the variable be height, measurement of students in a class of Students. Then the statement that the height measurements of student in the said class follow normal distribution is an example of statistical hypothesis. There is another form also of statistic, hypothesis Let it for example, be presumed or known that the random variable follows, normal distributions. Then the statement that mean p the normal random variable is12 is also an example of statistical, hypothesis. In practice, it is latter type of hypothesis which have to tackled most of the times. Thus, therefore, amounts to saying, that density function will be assumed to be known and hypothesis will consist of an assertion about values of all or a few of the parameter of the density function. Let the density function of x1 for example be of the form 0(fx) = e0x then the assertion that value of 0 is 2, is a statistical hypothesis. By nature, a statistical hypothesis can be either simple or composite. If a hypothesis specifies the values of all the parameters of a density function completely. It is called “Simple hypothesis”. Otherwise it is called. Composite hypothesis”, i.e. statistical hypothesis which is not simple is called a “Composite hypothesis.” Let the density function be 1       1 ( x - u A =---------e--I--------I oZ2n    2 ^ ex ) 1 fx*0) = xzne If the hypothesis m = 10 and o = 2 then this is a simple hypothesis. If, however, the hypothesis is p =10 and o = 2 or p = 10 and o > 2 or p = 10 o = 2 then the hypothesis will be composite hypothesis, because the value of α has not been completely specified. Self-Check Exercise-1 Q1. Define statistical hypothesis. Q2. Distinguish between simple and composite hypothesis. 11.4 . Testing a statistic hypothesis, Null hypothesis, Alternative Hypothesis and Critical region The procedure for deciding whether to accept or reject the hypothesis is called p. Testing a Statistical hypothesis”. As per definition, the statistician has unlimited freedom in developing this procedure i.e. designing the test, naturally, he will be guided by its properties. For example let us take (i) as the density function of the variable X, i.e. let the density function be f(x) = 0e. Let it be further supported that 0 can take only two values, either 2 or 1. The statistician who is working on the problem, either with his intuition, on the basis his experience or on account of some information available on the related factors favours the value 0 = 2. He therefore assumes that 0 has the value 2. The assumption is the statistical hypothesis to be tested. Let this hypothesis be denoted as H0. This is known as, “Null hypothesis”. 11.4.1    Null Hypothesis and Alternative Hypothesis According to Prof. R.A. Fisher, null hypothesis is the hypothesis which the tested for possible rejection under the assumption that is true. In other worlds, testing the null hypothesis involves testing whether the observed values are significantly different from the expected values. The null hypothesis may be rejected or accepted with some specified probability depending upon the outcome of the test applied on it. When testing the null hypothesis, the position of a statistician is on the null, inclined neither to accept it nor to reject it, but allowing full say of observational facts and of the testing procedure adopted, in taking decision either way decision about it. A possible or acceptable hypothesis alternative to the null in many instance, we formulate a statistical hypothesis for the sole purpose of rejecting or mortifying it. Similarly, if we want to 166 decided whether one procedure is better than another, we formulate the hypothesis that there is no difference between the procedure (i.e. any observed differences are merely due to fluctuation in sampling from the sample population). Such hypothesis are often called null hypothesis and are denoted by H0. Any hypothesis which differ from a given hypothesis, is called “Alternative hypothesis” and is denoted by H1. So in the problem under consideration θ = 1 is our Alternative hypothesis i.e. H1 : θ = 1. Thus our problem is not that of testing H0, against H1 where. H0= θ = 2 and         H1 = θ = 2 Now, to test H0, sample of observations on the random variable X, will be taken. To avoid the complications, and to make it easier to understand the phenomenon involved in testing, let the sample consist of one observation only. In real life problems, one usually takes several observations but here we are considering only one observation for the reason mentioned above, the philosophy of the testing procedure remains the same. Now, the decision whether to accept the hypothesis or reject it will be made on the basis of the value of the random variable X obtained. This is quite obvious that rejection of H0 implies acceptance of H1 and vice-versa. So the problem then is to determine those values of X which will correspond to rejection of H0 and those which will correspond to acceptance of H0. It is obvious that if the choice has been made of the values of X that will correspond to rejection H0, then the remaining values will necessarily correspond to acceptance of H0. 11.4.2    Critical Region Aggregate of all such values which correspond to the rejection of H0 is called the Critical region of the test. All other values of X which correspond to acceptance of H0, constitute what is known as “Acceptance Region’, or Non-critical Region’. In the problem under consideration let us suppose that the variable X can take values on the positive half of X-axis only i.e. (0 ≥ x ≥ 06). Every positive out some of X can be represented by a point on the positive half of the X axis with its x coordinate giving the value of the associated random Variable X. The problem of constructing a test for H0, is therefore, the problem of choosing a critical reason on the positive half of the X-axis. Suppose the statistician arbitrarily chooses, the part of the X-axis to the right of X = 1 as the critical region. To decide whether this was a wise choice, its consequences will have to be considered fully well. Self-Check Exercise-2 Q1. Define a)    Null Hypothesis b)    Alternative Hypothesis Q2. Explain the concept of critical region in hypothesis testing. 11.5 . Type I and Type II Errors and their sizes : Since we are basing our decision of accepting or rejecting the hypothesis on the sample observations only arid the critical region has-been chosen arbitrarily, there is always the likelihood, howsoever, small it may be of committing an error in decision making. There, existing only two possibilities with regard to H0; either it is true or it is wrong (i.e. H1 is true) Similarly, there exist 167 two possibilities with regard to the Decision: either accept H0 or reject. H0 as a result of testing procedure. If H0 is really true and the observed value of X exceeds 1, H0, will be rejected because it has been agreed to reject H0 when sample observation falls in the Critical region. Obviously, this is an incorrect decision. If this kind of decision is taken, it will amount to committing an error. This kind of error is called “Type I Error”. On the other hand if H0 is really wrong (i.e. H is true) and the observed value of X does not exceed 1, H0, will be accepted. This is also an incorrect decision because we are accepting H0 while it is wrong. This kind of error is known as Type II Error. In fact, in all, there are four possibilities, two mentioned above lead to incorrect decision and the remaining two viz. accepting H0 when it is true and rejecting H0 when it is not true lead to the correct decision. These four possibilities have been displayed in the table given below: Table I | | H0 is true | H0 is wrong (H1 is true) | |---|---|---| | X > 1 Reject H0 | Type I Error (incorrect decision) | Correct decision | | X < 1 | Correct | Type II Error | | Accept H0 | decision | (Incorrect decision) | It is just not enough to know the kind of error that may be committing in decision making but it is also necessary to measure them in some way before one can judge, whether or not the choice of critical region was wise. This can be accomplished by using what is known as the size of an error as the measure of its seriousness. The “size of type I error” is the probability of making a type I error which in turn in the probability that the sample point will be fall in the critical region H0, is true. The size of type II error, similarly is the probability of making a type II error which in turn again is the probability that the sample point, will fall in the non-critical or acceptance region when H0 is wrong i.e. H1 is true. Size of type I and type II errors are denoted by α and b respectively. ∞ 2                     θ-2 if H0 is true a = 2 2e - xdx + 135 hence we putθ= 2 here 1 θ-2 if H0 is true β = e - xdx + 632 hence we put θ= 1 here 0 For the problem under consideration, these will be as follows: Now, in terms of the two types of errors, it is possible to introduce a simple principle to be followed in’ the process of determination of a good test of hypothesis from amongst many that may exist. Many principles can be, suggested. For example, one may think of minimizing the sum of two types of errors or and product of the two errors or and other desirable and function of two errors. However, among all possible alternatives, the principle : “Among all test-possessing the same size type I error chose one for which the size of the type II error is as small as possible” has, been found to be the best Size of the type II error usually Increases if the size of type I error is decreased and hence one cannot think of making type I error as small desired without paying for an increasingly large type II error. In real life experiments, it is often necessary to adjust the type I error until a satisfactory balance has been reached between the size of two errors. Self-Check Exercise-3 Q1. What is a Type I error in hypothesis testing and what are the potential consequences of committing this error? Q2. Define a Type II error in hypothesis testing. How does it differ from Type I error? 11.6    Level of Significance: In testing a given hypothesis, the maximum probability with which, one is willing to risk the type I error is called the level of significance of the test, Thus the level of significance is the size, of the critical region : i.e., the probability assigned to the critical set. The commonly assigned probability are 0./05 and 0.01. An event E is said to be (i)    significant if under H0 : 01 - P (E) < 0.5 This normal curve of distribution is the most important theoretical distribution in statistical theory. The probability distributions (we shall see fee meaning of probability distributions in the next lesson) of roost sample statistics closely resemble the normal distribution. The fundamental importance of the normal distribution in statistics arises from the fact that the measures computed from samples usually tend to be normally distributed, whether or not the original data conforms to a normal distribution. “The normal curve of error stands out in the experience of mankind as one of the broadest generalizations of natural philosophy. It serves as the guiding instrument in researches in the physical and social sciences and in medicine, agriculture and engineering. It is an indispensable tool for the analysis and interpretation of fee basic data obtained by observation and experiment. Self-Check Exercise-4 Q1. Explain the following term a) Level of Significance 11.7    Properties of the Normal Curve Drawing the graph of a normal curve. Any symmetrical curve is not necessarily a normal curve. Although every normal curve is a symmetrical curve. (i)    The arithmetic average, median and mode in a normal curve coincide. This holds in any bell shaped symmetrical distribution. They fie at the point where the curve has the maximum height Draw a normal curve and mark the point where Md and Mo lie. (ii)    The first and third quartiles are equidistant from the median. Similarly, third and seventh deciles, twentieth and, eightieth percentiles, etc. are equidistant from the median. This relationship also holds good, in all symmetrical curves. You know this from your knowledge of skswness. (iii)    Mean deviation is 7979 or about 5 th of standard deviation. This relationship is not needed in our, subsequent studies, but it should be given if a question is asked on the characteristics of a normal curve. (iv)    Semi inter-quartile range or quartile deviation is equal to the probable error and probable error is 6747 or approximately 3 rd of the standard deviation. (v)    The normal curve is asymptotic to the x-axis. Understand clearly what is meant by “asymptotic” here. This means that the normal curve continues to approach the base line but never reaches or touches it. It can go to any distance on either side of the mean point, but it will not touch the x-axis. (vi)    The points of inflection of the curve lie at a distance of one standard ‘deviation on either side of the mean. Ordinate points of inflection are the points where the curvature of the curve changes its direction. (vii)    This characteristic refers to the relationship of the ordinates of height of the curve to the height of the mean ordinate. The ordinate or the vertical length of the curve at the mean (or Md or Mσ) called the mean ordinate is the highest, ordinate. The height of fee ordinate at a distance of one standard deviation (σ) from the mean is 60.653% of the height of the mean ordinate. Heights of other ordinates at various distance from the mean are also in fixed relationship with the height of the mean ordinate. In your graph of the normal curve measure a distance of one standard deviation from the mean. Draw a vertical line from here to meet the curve. This vertical distance is 60.653% of the ordinate at the mean. (viii)    Area relationship is the most important relationship in a normal curve. Most of the sampling theory is founded on this area relationship. The area of the curve-covered between the mean ordinate and an ordinate at σ (standard deviation) distance from the mean always has a fixed relationship with the total areal of the curve. Thus, the area enclosed between mean ordinate and the ordinate at a distance of one s from the mean is always 34.136% of the total area of the curve. Thus, if you draw two coordinates each at a distance of one σ on either side of the mean the area enclosed between them would be 34.134% approx = 34,134% approx = 68.267%. This will always be true whether we are drawing one normal curve or another. The following figure shows the area relationship in a normal curve. Fig.1 SS^Si. Area relationship in a normal The above figure shows the area enclosed by ordinates at 0; o and 3 o distances from the mean ordinate. The following table shows the area relationship in a normal curve in more details. Similarly, the area enclosed between two ordinates at 2o distances on both, sides of mean is 95.45% of the total of the curve between two ordinates at 3 o distances on both sides of mean is 99.73% of the total area, of the curve. These relationships are available in the form of printed tables. Five area relationships, however are of special importance in sampling and should be remembered by all. | Distance on both sides of Mean ordinate | Percentage of total area enclosed | |---|---| | ± | 1.00a | 68/27 | | ± | 1.96a | 95.00 | | ± | 2.00a | 95.45 | | ± | 2.58a | 99.00 | | ± | 3.00a | 99.73 | Thus, if we draw ordinates at 2.58o tf distances on both sides of mean, they are a covered between these ordinates will be 99% of the total area, i.e. only 1% of the area shall be left out of these ordinates. Various tests of significance are constructed on the basis of taking 95% 99%, or 99.73 of the area into account. Self-Check Exercise-5 Q1. Explain the properties of the Normal Curve. 11.8    Equation of the normal curve : The density function of a normal curve is given as f ( x) = x = X - ^ = 1 oZ(2n ) 2 -1/2 e o Height of ordinate on any given point on x-axis for a normal curve can be written as follows. Ni y =------- - x 2/2o2 oZ(2n)e Where y is the ordinate e is a mathematical constant having a value of 2.71828, o is the standard deviation, and x is a given value of the independent variable expressed as deviation from the mean. N stands for total number of cases. The maximum ordinate or mean ordinate can be derived from the equation. y0 Ni oZ(2n ) where N is the total number of items in the simple, i is the class interval, n (called pie) is a constant 22 having a value 7 = 3.1416. ∠ (2π) = 2.5066. Ni 2.5066σ The equation of the normal curve can thus be written as: y= Ni 2.5066σ × 2.71828-x2/2σ The theoretical frequencies can thus be found with the help of this curve. But calculation of theoretical frequencies like this is neither necessary nor advisable. Example 1. Fit a normal curve of the following frequency distribution relating to the height of certain children, | Heights (Inches) | No. of Children | |---|---| | 40.5—42.2 | 1 | | 42.2—44.5 | 4 | | 44.5—46.5 | 2 | | 46.5—48.5 | 18 | | 48.5—50.5 | 14 | | 50.5—52.5 | 23 | | 54.5—56.5 | 10 | | 56.5—58.5 | 7 | | 58.5—60.5 | 10 | | 60.5—62.5 | 3 | | 62.5—64.5 | 0 | | 64.4—55.5 | 1 | | Total | 106 | in the above frequency distribution, number of frequencies is 106 and class interval 2. The standard deviation is 4.7 (you can calculate S.D. yourself). Height of the mean ordinate, Ni 2.5066σ N = 106 106 × 2        212 × 2.5066×4.7   11.78102 σ i = 2 = 17.993 or approximately 18. Thus the height of the mean ordinate is 18. The height of the ordinate at one σ distance from the mean would be Ni 2.5066σ 2.71828 - (4.7)2 2(4.7)2 (4.7)2 (4.7)2 106x2 2.5066 x 4.7 -1 2.71828 2 cancels out 106x2           1 2.5066x4.7 (2.71828-1/2) as e -1/2 can be written as = 18 x          = 18 x 271823 1 1.6489 = 10.9175. Thus the height of the ordinate at one o distance from the mean on either side would be 10.9175. Similarly, the heights of otherordinates can be calculated and these points can be plotted to obtain, a normal curve. Alternatively, the height of ordinate at o distance canbe obtained by multiplying the height of mean ordinate by .60653. In this case. 18 × 0.60653 = 10.91754. Self-Check Exercise-6 Q1. Explain the equation of the normal curve. 11.9    Summary In this unit, we discussed statistical hypotheses and how to test them using null and alternative hypotheses. We explored the concepts of Type I and Type II errors, the critical region, and the level of significance. Additionally, we examined the properties and equation of the normal curve, supported by practical exercises. 11.10    Glossary •    Statistical Hypothesis: A statement that can be tested statistically to determine whether it is likely to be true. •    Null Hypothesis (H0): A default hypothesis that there is no effect or no difference, which is tested for possible rejection. •    Alternative Hypothesis (H1): A hypothesis that opposes the null hypothesis, indicating there is an effect or a difference. •    Critical Region: The set of values for the test statistic that leads to the rejection of the null hypothesis. •    Level of Significance: The probability threshold below which the null hypothesis is rej ected, commonly denoted by alpha (á). 11.11    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 11.3. Answer to Q2. Refer to Section 11.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 11.4.1. Answer to Q2. Refer to Section 11.4.2. Self-Check Exercise-3 Answer to Q1. Refer to Section 11.5. Answer to Q2. Refer to Section 11.5. Self-Check Exercise-4 Answer to Q1. Refer to Section 11.6. Self-Check Exercise-5 Answer to Q1. Refer to Section 11.6. Self-Check Exercise-6 Answer to Q1. Refer to Section 11.8. 11.12    References/Suggested Readings 1    Mason, R. D. (1986). Statistical techniques in business and economics. Homewood, IL: Richard D. Irwin, Inc. 2    Plano, D. R., & Oppermann, E. B. (1987). Business and economic statistics. Plano, TX: Business Publication, Inc. 11.13 Terminal Questions Q1. Answer the following questions (1)    What is statistical hypothesis? (2)    What is Null Hypothesis? (3)    What is type I error? (4)    What is type II error? Q2. What is normal curve? Discuss the properties of the normal curve. ***** UNIT-12NON-PARAMETRIC TEST STRUCTURE 12.1    Introduction 12.2    Learning Objectives 12.3    Tests of Difference with Correlated Data 12.3.1    The Sign Test 12.3.2    The Sign Rank Test of Difference 12.3.3    The Run Test 12.3.4    The Median Test 12.3.5    Wilcoxon-Mann Whitney U Test Self-Check Exercise-1 12.4    Advantages and Limitations of Non-Parametric Tests Self-Check Exercise-2 12.5    Summary 12.6    Glossary 12.7    Answers to Self-Check Exercise 12.8    References/Suggested Readings 12.9    Terminal Questions 12.1    Introduction Dear Student, Most of the tests which we have considered in earlier lessons have been based on the assumption that parent population is normal. Even if the parent population is not normal, we can often find transformations to reduce it to the normal form. In practice, however, our knowledge, about the parent population may not be sufficient to enable us to find such a transformation and in such case we need tests which do not depend on any assumption about the form of the population. In the last few decades, many new statistical procedures have been developed especially, to take care of the experimental situation in which samples are smalls and the form of population distribution is net not normal. We shall consider some of these non parametric or distribution free tests in this lesson. It is obvious, however that the collection of these results cannot be so comprehensive as in the normal case, but these tests are not only capable of wider application, but also are simpler to apply and do not require complicated sampling theory Distribution free methods are based on ordered statistics or ordered samples i.e. we suppose that the sample x1 x2   xn is ordered so that the observations in it are in ascending order of magnitude. Also while in the parametric case the measure of legation and depression which are most commonly, used are the measured standard deviation, respectively which do not depend or order in thesample in the non-parametric case, we prefer to use the median, quartiles, inter-quartile range, etc. for which an ordered sample is desirable. Point Estimation and Confidence Intervals. The population median v is estimated by the sample median x which however is not an unbiased estimate, thoughthebias is not serious and tends to zero as n tends to infinity. Similarly to estimate the population quartiles or deciles we use the corresponding sample quartiles or deciles as estimators. To construct a confidence interval for v, we use the fact that the probability of an observation falling to the left or right of y is one half in either case. The probability that xr, rth order statistic, exceeds v is given by r-1 P( xr > v) = ^ (xc.)(l/2)n                   (1) i=0 n Similarly (P xs < v) = ^(xc)(1/2)(2) s=0 s-1 and   (Pxr < v < xs)= £(xcr )(1/2)n i=r (xr, xs) gives the confidence interval for the confidence coefficient given by right hand side, of the third equation. A confidence interval for every confidence, coefficient cannot be constructed as R.H.S. of (3) can take only a certain set of values for different values of r and s. Though linear interpolation can be done but generally we restrict ourselves to the confidence levels available with simple order statistic. Oneof the drawbacks of distribution free method is the pancily of confidence levels for small samples. For moderate samples sizes, we can compute the RHS of (3) either directly or by use of tables of incomplete data functions. For large samples we can use the results that the number of successes will be asymptotically distributed with mean (x/2) and variance (x/4) 12.2    Learning Objectives After going through this unit, you will be able to •   Understand the purpose and applications of non-parametric tests. •    Learn various non-parametric tests for analyzing data with correlated samples. •   Identify the advantages and limitations of non-parametric tests. 12.3    Test of differences with Correlated Data 12.3.1    The Sign Test One of the simplest test of significance in the non parametric category is the sign test Let us .say that we have parallel set of measurements that are paired off in some way or let (x1 x2 ........xn) and (y1, y2 .....yn be two random samples of same size x from two populations and let the sample value be paired by pairing the ith member of one wife the ith member of the other and consider the signs of the differences (xi-yi), (i = 1.2,   n). If the two populations are continuous and identical the probability of a difference being positive or negative is one half, so we can test the null hypothesis by treating the number of positive signs as a binomial variable with mean (n/2) and variance n/i. In case of ties (xi — yi), can either ignore them or decide to allot them positive or negative signs by tossing a coin or assign half of them positive sign or other half negative signs. Usually the. first choice of ignoring them is preferred and in that case the conditional probability of the sign being positive, given that the difference is non-zero, is one half. Some of the distribution free or non-parametric, methods have lower power to detect a teal difference as significant. When there is arty choice, a t test is generally more efficient them a sign test and we should prefer a parametric test. But the sign test is much easier to apply and is applicable even when t test is not applicable., viz, when the parent population is not necessarily normal. Example In Table 12.1, we find, two set of knee-jerk measurements, both from the same men but obtained under, two conditions. In the first case (x), the subjects were squeezing a hand dynamometer just before the stimulus struck the knee and in the sex and case (y) the relaxed kneejerk was obtained in a released sitting posture. The hypothesis to be tested is that they arose from random sampling from the same population. If this hypothesis is true, half the changes from x to y should be positive, and half should be negative or we can state null hypothesis that the median change is zero. Application of the sign test to 1- pairs of knee-jerk data from table 12. Table 12.1 | x | y | Sign of x-y | |---|---|---| | 19 | 14 | + | | 19 | 19 | 0 | | 26 | 30 | - | | 15 | 7 | + | | 18 | 13 | + | | 30 | 20 | + | | 18 | 17 | + | | 30 | 29 | + | | 26 | 18 | + | | 28 | 21 | + | | *X = Knee-jerk measurement under tension | | R = measurement under relaxation | | There are 10 pairs of observations; therefore, 10 changes are involved. Since one change is zero and hence cannot be included as either positive, or negative. The hypothesis now calls for 4.5 positive differences, whereas we obtained eight. Is this a significant deviation ? The null hypothesis is H0: P - l/2 H1 : P > l/2 The obvious test to make is based on the binomial distribution for P = 5 and N = 9. On this bias, 8 or more plus signs could occur by chance 10 times in 512 trials (1 chance in 512 for exactly 9 plus 9 chances for exactly 8). For a one tail test this deviation is significant with p equal to approximately 102. For a two tailed test we double the probability, which gives a departure, significant at 04 level. We would make one tailed test, if the alternate hypothesis at the start were to expect x values to be higher than y values. The assumption involved in making the sign test include mutual independence of the differences. The two parallel set of values may or may not be related. Nothing is assumed regarding the equality of variances. The differences need not even he measured accurately but the direction of each difference should be experimentally established. The Sign Rank test of Differences Let us use an illustration of the sign rank test of differences the same data to which the sign test was applied in Table 12.1. The ten pairs of knee-jerk measurements under tensed arid relaxed conditions are repeated for convenience in Table 12.2. Here the numerical differences with algebraic, signs, are also listed. As in the sign test, however, we cannot use zero differences, since the differences must be classified according to algebraic sign. Rank the differences according to size irrespective of their algebraic signs, giving the smallest difference a rank of 1. Table 12.2 | X* | Y | X*-Y | Rank of absolute difference | Rank with minority signs | |---|---|---|---|---| | 19 | 14 | +5 | 4.5 | | | 19 | 19 | 0 | - | | | 26 | 30 | -4 | 3 | | | 15 | 7 | +8 | 7.5 | | | 18 | 13 | +5 | 4.5 | -3 | | 30 | 20 | +10 | 9 | | | 18 | 17 | +1 | 1.5 | | | 30 | 29 | +1 | 1.5 | | | 26 | 18 | +8 | 7.5 | | | 28 | 21 | +7 | 6 | | | *X = knee jerk score under tension X = -3 | | Y = score under relaxation | Two difference of rank 1 are given an average rank of 1.5. The next smallest difference is 4, which is given a rank 3, and so on until all non zero differences ranked. Now single out all differences whose signs are in minority. If there are fewer negative than positive signs, we select all ranks corresponding to the difference having that sign. There is only one negative difference in Table 12.2. We put this rank with negative sign in the last column. We sum this column to give an statistic T. The hypothesis test is that the difference are symmetrically distributed about a mean difference of zero. If this is true T would coincide with the mean of much sums of randomly selected ranks T which is also the sum ofN successive ranks and is given by the formula - N(N +1) T  ---n--- (Mean of sum of ranks) The obtained T of (-3) (the algebraic sign does not matter in the use of table) is significant at the 0.2 level (a two tail test) when we have a nine differences involved. For samples larger than 25, a standard deviation and a z ratio can be computed and S can be interpreted in terms of the normal distribution. For a sample of size N, ^1 7 N(N + 1)(2N +1) 24 and S is equal to (T- T )/oi 12.3.3    The Run Test We have to test the null hypothesis that two ordered random samples x1 x2 …….xn, and y1,y2 ...... yn come from the same population. The two sets of observations are combined and arranged in order of magnitude as say. x1 y1, y2 x2 x3 x4 y3 and in this new arrangement we find the total number of funs, where a run is defined as sequence of letters the same kind bounded by letters of the same kind. Thus starts with a run of one x, followed by a run of two y’s which is succeeded by a run of three x’s and so on. It is obvious that the two samples are from the same population, x’s and y’s will be well mixed and r, the number of runs would be large. If the two populations areso widely separated that they do not overlap the number of runs would be only two. There will be a long run of x’s at the end if the two populations have the same mean/median, but the first population has a larger variance. Any difference in mean or variance tends to reduce r. The run test for testing the null hypothesis that the two populations are identical accordingly consists In counting the number of runs in the combined ordered samples and rejecting the null hypothesis if s < r0 where r0 is number to be determined from the distribution of runs and depends on n1 and n2 and the level of significance but not on the form of the distribution of parent population. The portability of exactly r runs is given by | | 2 n1 - 1c k-1 X n2 ~ 1 ck-1 r = 2k | |---|---| | P (r) = ] | n1 + n2cn1 n1 - 1c k X n2-1 k-1+ n1-1ck - 1 X n2-1ck r = 2k n1 + n2cn1 | r = 2k+1 To test the null hypothesis at probability level P, we find r0 from. r0 £ p(r) = P r=0 as nearly as possible and reject H0, if r < r0. For large n1, n2 the distribution of r is asymptotically normal with E(r) = r = 2nn —1 2- +1 n1 + n2 2n1n2(2n1n2 - n1 - n2) Var (r) = o. = 7 i i u    r2   (n1 + n2) (n1 + n2 +1) and we reject the null hypothesis if r < r1, -1.045 on. The above approximation can be used if both n1,n2 exceed 10. The run test can also be used for testing the randomness of a give a sample i.e., to test whether the sample, observations are independent and identically distributed or in other words to test whether the phenomenon yielding the data is in statistical control. 12.3.4    The Median Test The run test discussed above is sensitive to both differences in shape and location of two distributions. If we interested in differences in location only, the median test is to be preferred as it is very little sensitive to differences in shapes. Let x1, x2 ......... xn1 and y1, y2 …… yn2 be two ordered random samples from the two populations f(x) and f(y) respectively and z1,z2 zn1+n2 be the combined order, sample of za as its median. Under this null hypothesis, the portability that m1 or rs and m2 of the y’s are less than za is given by n 1cm2 n2cm2 / ni + n2ca (m1-m2 + 1 = a) where we take a to be (n1 + n2) /2if n1, + n2 is even and (x1 + x2 + l)/2 if x1 + x2 is odd. If the null hypothesis is true, m1 has a hypergemetric distribution with mean n1/2and variance. Example : The Median Test Table 12.3 Application of the median to two samples under conditions A and B. | Samples | |---| | A | B | Contingency table sample | | 14 | 5 | | | | | 13 | 7 | | A B Both | | | 10 | 6 | 10+ | 5     2    7 | | | 12 | 5 | 9- | 2     5    7 | | | 15 | 11 | | 7     7 | | | 9 | 8 | | | | | 9 | 10 | | | | Mdn = 9.5 A common median, has to be calculated of the two samples to carry out median test. The number of cases above and below the common median are to be counted in each sample resulting in a fourfold contingency table. The observations are not paired or correlated and the N may differ in the two-samples. Equal N’s would make the test easier to apply. Then chi-square & test can be used with the test statistics s follows, with equal number of observations in each sampler we can conveniently use Table M for a test of significance without computing chi-square. The median of the 14 observations is 9.5 Values of 10 and above are easily segregated from those of 9 and below, as show in the four-fold table. With such small frequencies, chi-square is not computed. P = 7c5 7 c5 14c7 21 x 21   441 3432   3432 = 1.12 As this probability is greater than 0.05 the level of significance, hence we can accept the null hypothesis, that median is the same for both the populations Median Test with more than two samples. Suppose that we have three samples, each from its own set of conditions. We want to test the homogeneity of their central values. For example consider the samples in Table 12.4. Table 12.4 | Sample | | | | | | |---|---|---|---|---|---| | N | P | K | | | | | | | 2 | 10 | 12 | Contingency table | | | | 7 | 7 | 15 | | N | P | K | all | | | 5 | 12 | 9 | 10+ | 0 | 4 | 4 | 8 | | | 6 | 14 | 16 | 9- | 6 | 3 | 1 | 10 | | | 8 | 9 | 14 | | 6 | 7 | 5 | 18 | | | 3 | 8 | | | | | | | | 10 | | | | | | | | | Ni 6 | 7 | 5 | Mdn = 9.0 | The median of all 18 observations is 9.0 we cannot make the point of dichotomy at exactly 9. In such a situation it is made clear that it should be near the median. Let it be point 9.5. We then set up a contingency table as in above table. From these data chi-square is 7.82 with 2df this chi-square is significant near the 0.2 point, We reject the null hypothesis and say that the three medians are not homogeneous. 12.3.5 Wilcoxon-Mann Whitney U Test Let two ordered samples combined and ordered in order of magnitude as follows, say x1, y1, x2, x3, y2, y3 …….. For each y we count the number of inversions i.e., the number of x’s that precede that y in the sequence. Thus for y1, y2 y3 in the sequence there are 1, 2, 3 inversions. Alternatively for each pair of observations xi, yi, let Zij = < 1 if x < yj 0 if X: < y, ij Then the total number of inversions would be nz ni U = ZZ zij j-=1 i=1 Under the null hypothesis Zij is a Berroullian variable with p =1/2 and nn       nn E(u) = 2". Var ^ =      (n1n2 +1) also n is asymptotically normal with, these parameters. The hypothesis being tested by the Mann-Whitney’ u test takes care of samples of unequal size and the operation through to the finding the sums of the ranks are the same as in the composite rank method When Na and Nb are both as large as 8, a z test can be used and z can be computed by the formula. 2Ri - Ni(N +1) \nanb (N = V nan 3 (z value for an obtained sum of ranks of 4 test). where R1 = one of the sum of ranks Na and Nb = replication in samples A and B respectively. N = total number of cases = Na + Nb. Ni = number of cases corresponding to Ri. The hypothesis being tested is that one set of measures, as a group, is equal to another. Statistically me H0 being tested is that the obtained u, minus the u to be expected for a particular combination of Na and Nb is zero. The expected u is equal, to na nb/2, u1 is given by the formula. U = NaNb + N(Ni + 1) - R i ab       2       i (The Mann-Whitney U statistic). Deducting the expected u from the obtained ui and multiplying through by z, the numerator of formula is obtained. The denominator is also 2 times the standard error of u. Let as apply formula to a problem with data in Table 12.5. Table 12.5 | Measurements | Ranks | |---|---| | A | B | A | B | | 14 | 5 | 13 | 1.5 | | 13 | 7 | 12 | 4 | | 10 | 6 | 8.5 | 3 | | 1.2 | 5 | 11 | 1.5 | | 15 | 11 | 14 | 10 | | 9 | 8 | 6.5 | 5 | | 9 | 10 | 6.5 | 8.5 | | | | S 71.5 | S 33.5 | | | | Ra | Rb | In this problem, Na from sample A is 10 and Nb from samples B is 8, All 18 measurements were ranked together. The sum for sample A(Ra) was 123, and that for sample B (Rb) was 48. The sum of those two values is 171, which gives us one check, since the total sum of ranks, which is given by N(N+l)/2, also 171. Using the Small Ri and applying formula. _ _ 2(48) - 8(19) z (10)(8)(19)/3 96-152 a/1520/3 -2.49 Assuming a normal distribution for this z the difference between the two set of ranks - and thus between die two set of measurements appears to be significant beyond 0.5 level in a two tailed test. The algebraic sign of z here does not matter. Self-Check Exercise-1 Q1. Explain how the sign test is used to compare two populations. Q2. To test whether students perform equally well on the verbal and non-verbal parts of SAT test, the scores from a random sample of 7 students are recorded (in pairs) as follow: | SAT (Verbal) | 560 | 510 | 620 | 480 | 590 | 610 | 420 | |---|---|---|---|---|---|---|---| | SATNon-verbal) | 620 | 650 | 700 | 620 | 450 | 460 | 400 | Use the sign test at 5% level of significance to test η1 = η2. Q3. A coin is tossed 20 times and the following sequence of heads (H) and tails (T) is obtained H T T H H T T H T H HH T T H H T H HH Use run test to determine at 5% significance if the coin is unbiased. Q4. Describe median test. Q5. Describe Mann-Whitney U Test. Q6. A researcher wants to determine whether or not there is a difference in the median weight loss under two different weight reducing diets. The weight loss (in pounds) in five individuals under each diet for a month are recorded as follow: | Diet A | 5 | 0 | 1 | 3 | 4 | |---|---|---|---|---|---| | Diet B | 8 | 2 | 6 | 6 | 1 | Use α = 0.05 12.4    Advantages and limitations of Non-parametric tests When certain assumption of normality are not satisfied, men the desirable properties of the estimators are no longer valid. In such cases the statistical tools to obtain the estimators satisfying the desirable properties are quite complicated and’at times inadequate too. This objective can be achieved pretty easily and quickly too using appropriate non-parametric tests. Non parametric tests are applicable to qualitative data as well as quantitative data. These tests can be applied to small samples and large samples. But the main drawback of these, statistic is that the sensitivity of results gets affected by the choice of initial and terminal observations. One weakness of the Sign test is that it does not use all the available information. If the measurement are on a scale of equal units, on which differences. may be computed for size as well as for direction, the sign test ignores the information provided by the size. Barring small samples, the sign test is only about 60 percent as powerful as at test would be for the same data, where both apply. Self-Check Exercise-2 Q1. Discuss the advantage and limitations of non-parametric tests. 12.5    Summary This unit discussed non-parametric tests and their applications for analyzing data. We covered several tests, including the Sign, Sign Rank, Run, Median, and Wilcoxon-Mann Whitney U tests. Additionally, we explored the advantages and limitations of using non-parametric tests in statistical analysis. 12.6    Glossary •    Non-Parametric Test: Statistical tests that do not assume a specific distribution for the data. •    Sign Test: A simple non-parametric test used to determine if there is a difference between paired observations. •    Wilcoxon-Mann Whitney U Test: A non-parametric test used to compare differences between two independent samples. •    Median Test: A test used to determine if two or more groups differ in their median values. •    Run Test: A non-parametric test that checks for randomness in a sequence of data points. 12.7 Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to 12.3.1. Answer to Q2. H0: n 1 = n2; H1 n1 * n2 (Two- tailed); X = Min (2,5) = 2 C. V. (for n =7, a = 0.05 for two-tailed test) = 0; C. R. X < 0 i.e., X = 0 Answer to Q3. We fail to reject H0 and conclude that the coin may be regarded as unbiased. Answer to Q4. Refer to Section 12.3.4. Answer to Q5. Refer to Section 12.3.5. Answer to Q6. Ans. H0: nA = nB, H1 = nA * nB ; U = Min (6.5, 18.5) = 6.5 ; C. R : U <2. Self-Check Exercise-2 Answer to Q1. Refer to Section 12.4. 12.8    References/Suggested Readings 1    Siegel, S., & Castellan, N. J., Jr. (1988). Non-parametric statistics for the behavioral sciences. McGraw-Hill. 2    Kapur, J. N., & Saxena, H. C. (1969). Mathematical statistics. S. Chand & Co. 12.9    Terminal Questions Q1. What is a non-parametric test? Describe the non-parametric test methods. Q2. Explain how the sign test is used to compare two populations. ***** UNIT- 13 CHI-SQUARE TEST STRUCTURE 13.1    Introduction 13.2    Learning Objectives 13.3    Chi-square Distribution Self-Check Exercise-1 13.4    Tests Based on Chi-square Distribution (X2) 13.4.1    Testing Goodness of Fit: Observed Frequencies 13.4.2    Independence of Attributes Self-Check Exercise-2 13.5    Summary 13.6    Glossary 13.7    Answers to Self-Check Exercise 13.8    References/Suggested Readings 13.9    Terminal Questions 13.1    Introduction In previous units, we discussed the testing of hypotheses and explored various tests based on different distributions. These tests are crucial for making statistical inferences and validating assumptions about population parameters. In this unit, we will focus on the Chi-square test, a versatile statistical tool used to test hypotheses concerning categorical data. For example, tests based on X2 = distribution have been found suitable and appropriate for solving problems related with goodness of fit and independence of attributes and tests based on normal distribution for testing if a given sample belongs to a particular population or not and to test the significance of difference between two sample means. 13.2    Learning Objectives After completion of this unit, you will be able to •    Understand the Chi-square distribution and its properties. •    Apply the Chi-square test to assess the goodness of fit for observed frequencies. •    Use the Chi-square test to examine the independence of attributes in categorical data. 13.3    Chi-square Distribution You have already studied Testing of Hypothesis. Hypothesis ate set in accordance With the objectives of the problems. For testing, various kinds of hypothesis different tests based on different distributions have been evolved. For example tests based on x2 = distribution have been found suitable and appropriate for solving the problems related with Goodness of fit and Independence of attributes and tests based on Normal distribution for testing if a given sample belongs to a particular population or not and to test the significance of difference between two sample means. There are a number of other tests based on other distributions but we will study only these two here. Self-Check Exercise-1 Q1. Define Chi-Square distribution. 13.4 . Test based on Chi-square Distribution (x2) As has been mentioned above that the tests based on x distribution are used mainly to solve the following kinds of problem; (i)    Goodness of fit (ii)    Independence of Attributes First of all we will discuss (1). 13.4.1 Testing Goodness of fit Observed Frequencies: Many a times situation arises when we want to see it an observed frequency distribution follows a particular theoretical distribution or not This is done by examining how well one fits into the other. It is on account of this fact that the test has been named as a test of the goodness of fit. For example, 200 digits may be chosen at random, from a table of random numbers frequencies of the. digits may be noted and it may be desired to example 1, here after). Similarly, 12 dice may be considered to have been thrown 4096 times. [Considering a throw of 4, 5 or 6 as a success it may be desired to test if the dice were unbiased], As a result, frequency distribution of number of successes obtained is as follows : Success 0    1    2      3     4     5     6     7     8     9    10    11   12 Frequency 7    60    198 430 731 948 847 536 257 11    11 - (To be referred as example 1, here after) Now if the dice were unbiased then prob. of 1 success = p = 2 (For each dice) and the frequencies of the distribution can be obtained by the terms in the Binomial expansion. 1      1 I 2 4 J 4096 Now the problem of testing if the dice were unbiased, reduce to the problem of examining how well do the observed frequencies fit into the expected (hypothetical) frequencies (ii). The laterones are called expected or hypothetical or frequencies because these have been obtained with the expectation or under the hypothesis that the observed frequencies follows this sort of distribution viz, binomial distribution in the example (ii) given’ above. Now before we proceed for defining x2 test for the problems of the kind mentioned above, it should be clearly understood that this test cannot be applied if (i) the total number of frequencies is not large enough i.e. at least 50 and (ii) the expected frequency of an individual class is very small i.e. less than 5. If these two conditions are fulfilled then x2 with n—k degrees of freedom is obtained with the help of the formula. n X2 -Z i=I (Ot- E) = Z(O - E )2 Ei          E ^ Simple way of ^ ^ writing ? where n is the effective number of classes K is the number of constraints O is the observed frequency of the 7 class. h is the expected frequency of the 7 class. In the problems of the kind given above, generally, k = 1 hence the test has (n—1) degrees of freedom. The calculated value of x2 is compared with the table valise* of x2 for the number of degrees of freedom of the test on specified level of significance. The most commonly used level of significance by applied statisticians is 05. If the calculated value of x2 is greater than the table value of x2, null hypothesis is rejected, if the calculated value of x2 less than the table value of x2, null hypothesis is accepted. Nail hypothesis is, Fit is good i.e. observed frequencies fit well into the expected frequencies. It is easy to follow that when all observed frequencies are identically equal to the expected frequencies, the calculated value x2 is 0, implying that the fit is the best. Obviously, larger the value of x2 wider the gap between observed and expected frequencies and proper the fit. Let us now try to understand the whole procedure with the help of the examples given above. Take example first, say the observe frequencies of the digits were : Digits :        0      1       2      3      4      5      6      7      8      9      Total Frequencies:  18     19    23    21    16    15    22    20    21    15     200 Under the hypothesis : “digits were equally distributed”, expected frequencies for all the | digits should be | 200 10 | - 20 each. So we have | | | | | Total | | |---|---|---|---|---|---|---|---|---| | Ohs. Freq = O | | 18 | 19 | 23 | 21 | 16 | 25 | 22 | 20 | 21.15 = 200 | | | Exp. Freq = E | : | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20.20 = 200 | | | (O-E) = | : | -1 | -1 | 3 | 1 | -5 | 5 | 2 | 0 | 1       -5 | | | (O-E)2 | : | 4 | 1 | 9 | 1 | 16 | 25 | 4 | 0 | 1      25 | | | (O-E)2 | : | 4 | 1 | 9 | 1 | 16 | 25 | 4 | 0 | 1      25 | 86 | | -- E | : | - 20 | - 20 | - 20 | - 20 | - 20 | - 20 | - 20 | - 20 | - 20    20 | - 20 | | 86 hence x2 = 20 - | 4.3000 | with (10-l) = | 9 d.f. | | | | | | | Table value of x2 for 9 d.f. for 5% L.S (level of significance) = 16.919. Now, since the calculated value x2 is much lesser than the table value of x2 = 0.05, 9 = 16.919 the hypothesis is accepted at 5% level of significance i.e. digits are equally distributed. Similarly, in regard-to the example 1 the expected frequencies as follow : | Success | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | Exp. freq. | 1 | 12 | 66 | 220 | 495 | 792 | 924 | 792 | 495 | 220 | 66 | 22 | 1 | Now, the condition (ii) is not fulfilled as the expected frequencies for success and 12 are less than 5 each and hence the test cannot be applied to this problem as such. To overcome this difficulty, the classes corresponding to successes 0 and 1 and the classes corresponding to successes 11 and 12 may be combined together, to form one class each so that after doing so, no class is left with frequency less than 5 and effective number of classes is not reduced by two. After doing so the data and the solution of the problem will be as follows: | Success | 0.1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11,12 | Total | |---|---|---|---|---|---|---|---|---|---|---|---|---| | Obs. Freq O | 7 | 60 | 198 | 430 | 731 | 948 | 847 | 526 | 257 | 71 | 11 | 4096 | | Exp. Freq. E | 13 | 66 | 220 | 495 | 792 | 924 | 792 | 495 | 220 | 66 | 13 | 4096 | | O – E | -6 | -6 | -22 | -65 | -61 | +24 | 55 | 41 | 37 | 5 | -2 | | | (O - E)2 | 37 | 36 | 484 | 4225 | 3721 | 576 | 3025 | 1681 | 1369 | 25 | 5 | 33.49 | | E | 13 | 66 | 220 | 495 | 792 | 924 | 792 | 495 | 220 | 66 | 13 | Effective No. of class = 11 (No. of classes after regrouping) Hence d.f. = (11 – 10) = 10 x2 0.05, 10 (table value) = 18.31 x2 0.01, 10 (table value) = 23.21 Since calculated value of x2 is greater than the tabulated value even at 1% I.s, the hypothesis is rejected i.e., the data is not consistent with the hypothesis that the dice were unbiased even at 1% level of significance. 14.4.2 Independence of Attributes: We have studied simple Correlation and Regression and we know that the coefficients of Correlation and Regression are meant, for studying the relationship between the two variables. This is possible only if both the variables are of quantitative nature. If even one of the two is of qualitative nature techniques of correlation and digression become non-applicable. Two way table presenting a bi-variate distribution is known as contingency table. In it at least one of the two variables is qualitative Various measures of association have been developed to study the relationship between the two variables in such a case. One of them is based on x2 —distribution which is meant to examine if the two variables or attributes (more appropriate, word to be used for qualitative variables) are independent or dot. If there exists no mutual relationship of any kind between two attributes A and B then they are said to be independent. If “A” stands for tossing a coin by right hand and “α” for tossing the coin by left hand. “B” i for getting head and “β” for getting tail, then, the two attributes tossing of the coin and the result of the loss will be said to be independent if there is absence of any relationship between them in the sense that the same proportion of heads is observed whether the coin is tossed with right hand or left hand. * you should know how to consult the table. Let the following table represent the observed frequencies, | Attribute | B | β | Total | |---|---|---|---| | A | a | b | R1 | | a | c | d | R2 | | Total | C1 | C2 | N | Now under the hypothesis that the two attributes are independent i.e. the proportion of heads should be the same whether, the coin is tossed by right hand or left hand, the corresponding expected frequencies, will be given as below : Expected frequency corresponding to observed, frequency ‘a’ denoted by E(a) will be given as E (a ) = R1 x C1 N Similarly: E (b) = R1 x C2 N E (c) = R2 x C1 N E (d ) = R2 x C2 N Thus the expected frequency corresponding to any cell in the contingency table is obtained, by multiplying the totals, of Row and Column in which the cell falls divided by the total number of frequencies N. Thus in a contingency table when each of the two attributes has two or even more than two classes the expected frequency of the (ij) the cell i.e. the cell corresponding to its ith row and jth column is obtained as E (i, j) = Ri x C N where        Rj is the sum of the ith row, and          Cj is the sum of jth column. Now when we have got observed as well as expected frequencies we can calculate the value of x2 with the help of the same formula viz. Degrees of freedom of the test are (m— 1)(n— 1) where m is the effective number of rows and n that of the columns. Rest of the testing procedure remains the same. The null hypothesis (i.e. H0). The two attributes are independent. The two conditions mentioned earlier have to be satisfied here also i.e. (i) the total number of observations should not be less than 50 and (ii) no cell frequency (expected) should be less than 5, In case of a contingency table of order m × n (m > 2) (m > 2) if any expected cell frequency is less than 5, then either two rows or two columns are pooled together, so that the new table has no expected cell frequency less than 5. If there are two or more cells having 189 expected cell frequency less than 5 pooling of both rows and columns may have to be done. Obviously the number of column? or rows or both will, be reduced. The reduced numbers of columns and/or rows are known as the effective number of columns and/or rows. This reduction in the number of columns and/or rows will bring about a reduction ‘in the number of degrees of freedom of the test also. This should be done such that the reduction in the number of degrees of freedom is minimum. The case of 2 × 2 table for such a situation will be dealt with later on. Example 3. Discuss on the basis of the following data given in respect of 1000 school boys if the two attributes general ability and mathematical ability are independent or not. | Maths Ability | General ability | Total | |---|---|---| | Good | Fair | Poor | | Good | 44 | 22 | 4 | 70 | | Fair | 265 | 257 | 178 | 700 | | Poor | 41 | 41 | 98 | 255 | | Total | 350 | 370 | 280 | 1000 | Assuming that the general ability and mathematical ability are independent, hypothetical or expected frequencies will be as given below : General ability | Good | Fair | Poor | Total | |---|---|---|---| | 70 x 350 | = 5 | 700 x 370 | 70-(24.5 + 25.9) = | 70 | | 1000 | 1000 =; | | 24.5 | | 25.9 | 19.6 | | | 70 x 350 | | 700 x 370 | 700- (259+245) = ; | 700 | | 1000 | ; | 1000 | | 245 | | 259 | 196 | | | 350-(24.5+245)=370+(25.94-259)= : 280-(19.6 +196) | = 290 | | 80.5 | | 75.1 | 64.4 | | | 350 | | 370 | 280 | 1000 | | 2   (44 *x | - 24.5) 24.5 | 2 (22 - 25.9)2 +    25.9 | +... + (98 - 84.4)2 = 72.1 84.4 | d.f of the test = (3-1) ( 3-1) = 4. Since the calculated value of x2 is much greater than the tabulated of x2, hypothesizes rejected i.e. me two attributes, general ability and mathematical ability are not independent i.e. they are associated. Note : If any of the expected cell frequencies would have been less than 5, two rows or two columns should have been merged together as per discussion given above. In a 2 × 2 table if the observed frequencies are represented as a, b, c and d as given below : Observed frequencies in a 2 × 2 table | | | | Total | |---|---|---|---| | | a c | b d | R1 R2 | | Total | C1 | C2 | N | then x2 with (2—1) (2—1) = 1 d.f. can be directly obtained by the formula 2_ (ad - bc) x N X = R1x R 2 x C1x C 2 In a 2 × 2 table, if any one of the expected cell frequencies is less than 5 then we have to apply Yate’s correction which is as given below ; Calculate the products ad and be if ad > be subtract 1.5 from a and b each and add 0.5 to b and c each and vice-versa. This will leave R1, R2, C1, C2 and N unaltered? Now work out the expected frequencies and follow the usual procedure. The D.F. of the test will remain = 1. Directly the value of /’ will be given by tire formula N12 < (ad - bc)f x N X2 =  -----------2^------ R1 R2 C1 C2 with I.d.f. Example 4: Examine if the vaccination has any effect on the possibility of survival in case of human beings on the basis of the data given below. | | Survived | Died | Total | |---|---|---|---| | Vaccinated | 40 | 4 | 40 | | Non-Vaccinated | 50 | 1 | 51 | | Total | 90 | 5 | 95 | Expected frequencies on the assumption that the possibility of survival is independent of vaccination, expected frequencies will be as given below: | | Survived | Died | Total | |---|---|---|---| | Vaccinated | 44 x 90 | (44-41.7) = 2.3 | 44 | | 95 = 41.7 | | Non-vaccinated | 90-41 = 48.3 | (5-2.3) = 2.7 | 51 | | Total | 90 | 5 | 95 | Since two of the expected frequencies are less than 5 each Yale’s correction will have to be applied and accordingly the observed, frequencies will be corrected as: *(ad-bc) indicates difference between ad and bc, such that the smaller product is subtracted from the greater i.e. (ad-bc) is always a positive quantity. | | Survived | Died | |---|---|---| | Vaccinated | 40.5 | 3.5 | | Non-vaccinated | 49.5 40 × 1 = 40 50 × 4 = 200 200 > 40 | 1.5 | hence 5 added to 40 and 1 each, and 5 subtracted from 50 and 4 each. Now the x2 can be calculated with the help of the formula x2 = i iO-< E by taking the corrected Observed frequencies. The test will have l.d.f. Alternatively, the problem can be solved by using the direct formula : ' t , N V * (ad ~ bc) - — ? x N χ R1 R2 C1 C2 2 (40 x1)~(50 x 4) - 951 x N 2 (200 - 40) - yV x 95 44 x 51 x 90 x 5 44 x 51 x 90 x 5 with I.d.f. 2 v (O - E)2 Note : Calculate X = i-------- and examine that it is E equal to the value obtained by the formula χ2 I N (ad - bc)l - — 2 ? x N R1 R2 C1 C2 v2 - y (O — E)2 hint for calculating X = i-------- E Complete the problem | 0 | E | O-E | (O-E)2 | (O-E)3/E | |---|---|---|---|---| | 44 x 90 40.5           95 44 x 5 3.5             95 51 x 90 49.5           95 51 x 5 1.5             95 | | | | Self-Check Exercise-2 Q1. What is Chi-square test of goodness of fit? Q2. A sample of 200 persons with a particular disease was selected. Out of these, 100 were given a drug and the others were not given any drug. The results are as follows: | | Numbers of Persons | |---|---| | Drug | No Drug | Total | | Cured | 65 | 55 | 120 | | Not Cured | 35 | 45 | 80 | | Total | 100 | 100 | 200 | Test whether the drug is effective or not. Q2.   The number of parts for a particular spare part in a factory was found to vary from day to day. In a sample study, the following information was obtained: | Day | Monday | Tuesday | Wednesday | Thursday | Friday | Saturday | Sunday | |---|---|---|---|---|---|---|---| | No. of Parts Demanded | 1124 | 1125 | 1110 | 1120 | 1126 | 1115 | 6720 | Test the hypothesis that the number of parts demanded does depend on the day of the week. 13.5    Summary In this unit, we discussed the Chi-square test, which is used to test hypotheses about the distribution of categorical data. We covered the Chi-square distribution and explored how to use the Chi-square test for goodness of fit and independence of attributes. Through these tests, we can determine if observed data fits an expected distribution and whether two categorical variables are independent of each other. 13.6    Glossary •    Chi-square Distribution: A statistical distribution used for hypothesis testing in categorical data analysis. •    Goodness of Fit: A test to see how well-observed data matches expected data based on a specific hypothesis. •    Independence of Attributes: A test to determine whether two categorical variables are independent or related. •   Observed Frequencies: The actual counts of occurrences in each category of a dataset. •   Expected Frequencies: The theoretical counts of occurrences in each category if the null hypothesis is true. 13.7    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 13.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 13.4.1. Answer to Q2. Refer to Section 13.4.2. (Answer: Σ (O-E)2/E = 3.522) Answer to Q3. Refer to Section 13.4.1. (Answer: Σ (O-E)2/E = 0.179) 13.8    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Mill, F. C. (1955). Statistical methods. Pitman and Sons: London. 13.9    Terminal Questions Q1. What is Chi-square goodness of fit? Q2. A survey among the women was collected to study the family life. The observations are as follow: | Family Life | |---| | | Satisfied | Unsatisfied | Total | | Educated | 70 | 30 | 100 | | Non-educated | 60 | 40 | 100 | | Total | 130 | 70 | 200 | Test whether there is any association between family life and education. Use Chi-square test, the table value of Chi-square for 1 degree of freedom at 5% level is 3.84. ***** UNIT-14 STANDARD ERROR OF MEAN AND STUDENT’S ‘T’ DISTRIBUTION STRUCTURE 14.1    Introduction 14.2    Learning Objectives 14.3    Tests Based on Normal Distribution 14.3.1    Testing the Significance of Sample Mean x Self-Check Exercise-1 14.4    Standard Error 14.4.1    Standard Error and Size of Sample 14.4.2    Standard Error and Precision Self-Check Exercise-2 14.5    Sampling Distribution 14.5.1    Simple Sampling of Variables Self-Check Exercise-3 14.6    Standard Error of the Mean 14.6.1    Sampling Error and Level of Significance 14.6.2    Standard Error of Coefficient of Correlation Self-Check Exercise-4 14.7    Student’s ‘t’ Test 14.7.1    The Student’s T-Test and Its Properties 14.7.2    Student’s ‘t’ Distribution 14.7.3    Chief Features of the t Probability Curve 14.7.4    Properties of t Distribution Self-Check Exercise-5 14.8    Snedecor’s F Distribution 14.8.1    Variance Between the Samples 14.8.2    Variance within the Samples 14.8.3    Properties of the F Distribution and F Curve Self-Check Exercise-6 14.9    Interrelationship Between x/σ, t, X 2 and F Self-Check Exercise-7 14.10    Summary 14.11    Glossary 14.12    Answers to Self-Check Exercise 14.13    References/Suggested Readings 14.14    Terminal Questions 14.1    Introduction In the previous units, we discussed the testing of hypotheses and Chi-square distribution. In this unit, we will explore the concepts of standard error and Student’s t-distribution, which are essential for hypothesis testing, especially when dealing with small sample sizes. We will also examine the interrelationship between standard error, the t-distribution, the Chi-square distribution, and the F-distribution. 14.2    Learning Objectives After going through this unit, you will be able to •   Understand the concept of standard error and its relationship with sample size and precision. •    Apply the Student’s t-test to test the significance of sample means. •    Explain the properties and uses of the t-distribution in hypothesis testing. 14.3    Tests based on Normal distribution Broadly, there are two kinds of problems which are solved with the help of tests on normal distributions. (1)    Whether or not a given sample belongs to a specified population with mean µ. by testing the significance of difference between population mean µ and sample mean X . (2)    Whether or not means of two samples, are significantly different. Besides these, there are some tests for populations and percentages also. 14.3.1    Testing the significance of the Sample mean x . The distribution of the sample mean x 7 of a simple random sample x1,x2 xn approaches to the normal distribution with mean µ and variance σ2/n, (refer to chapter 12, topic “Sampling distribution”.) (µ is the population mean and σ2 the population variance and µ, the sample size), as n becomes increasingly large, even if the parent population from which the sample has been drawn is not normally distributed. So far every population with mean µ, and variance σ2 the statistic Z= x-µ*I σ/ n behaves like a standard normal variate, provided that n is sufficiently large. σ2 is rarely .known. If the sample size is very large, then the sample estimate of variance. 2   Σ(x - x)2 S= n-1 can be taken as an approximation to the population variance without any significant error of approximation and then the statistic Z takes the form Ix-µ S/Vn where Z now also, follows normal distribution with mean O and S.D.I. For testing : null hypothesis is. set H0 ; x = µ i.e. sample belongs to the population. Z value is calculated and compared with 1.96, 2.33 and 2.51 for 5%, 2% and 1% levels of significance, respectively. If the calculated value of Z exceeds the value (s) specified above, the hypothesis is rejected otherwise it is accepted. Example 5. A sample of 900 members is found to have a mean of 3.4 cms. Could it be reasonably regard d’ as a sample from a large population whose mean is 3.25 cms and S.D. = 2.61 cms. Null sample of 900 members may be regarded as a sample from the population with mean = 3.25 * | | is called Mod, meaning there by that only the absolute value of the figure enclosed is to be considered. u = 900 (large x = 314 cms. u = 3.25 cms and a = 2.61 cms. Ix -^ _ 13.4 - 3.251 _ Z                            1.72 a / 4n   2.61/7(900) Z (calculated) 1.72 < 1.96 and hence hypothesis is accepted i.e. sample 900 members may be regarded as a sample from the population with mean 3.5 cms. Example 2. Testing the significance of difference between the means of two samples : (i)    Suppose the two sample of sizes n1, and viz. x1, x2 ……xn, and y1, y2   yn2 have been drawn from the sample population with standard deviation a. We want to test whether the difference (x - y) of their means is significant. Null hypothesis H0: x = y i.e. there is no significant difference between the sample means, Z = Ix- y| a 71 i Ï ] ni — + — n2 JJ will be a standard normal variate for sufficiently large n1 and n2. If Z < 1.96, we would accept H0: x — y at 5% l.s. If 1.96 < Z < 2.58 we would reject H0 : x — y at 5% l.s. but will accept at 1% l.s and so on. (ii)    If the two samples have been drawn from two populations having different variance a12 and a22 respectively then our problem may be to examine whether the two populations differ in their means or not, leaving apart the difference in their dispersion. Let the means and variances of the two populations be gp o12 and u.. a22respectively. Null hypothesis in this case will be H0 : g1, = u, i.e. two populations do not differ in their means. In this case Z = a Ix- y| ( a ] H ni + 2 ^L n2 J will behave like a standard normal variate i.e. Z - N (0, 1). Further procedure will remain the same as described above in (i). (iii)    Variances 012 and c22 are known only very rarely in practice and hence are to be replaced 2 by S1 2( x — x )2 n1 — 1 and S22 ^( j — y )2 n — 1 estimates of O12 and G22 respectively obtained from sample. Then Z = k-d (7 SL + St ' I n1    n2 J will behave like a standard normal variate provided n1, and n2 are large enough if parent populations are not normal n1 and n2need not necessarily be large if parent populations are normal. Further test procedure will remain the same as discussed above in (i). Example 6. The means of the samples of 1000 and 2000 are 67.5 and 68.0 inches respectively. Can the samples be regarded as drawn from the same population of S.D. 2.5 inches. H0 : x = y i.e. two samples have been drawn from the same population n1, = 1000, n2 = 2000; a = 2.5, x = 67.5, y = 68.0. Z = Ix -y\ 167.5 — 68.01 If1+1T [ln1    n2 J = 517 If 1 + 2 2.5- i +------ \ [l100   2000 Now Z cal. (5-17) > 2.58 and hence hypothesis is rejected even at 1% l.s. i.e. the samples cannot be regarded as drawn from the same populations. Example 7 : A random sample of 1000. men from Shimla shows their mean wage to be Rs. 2.50 per day with a standard deviation of Rs. 1.50. A sample of 1500 men from Chandigarh gives a mean wage of Rs. 2.68 per day with a standard deviation of Rs. 2.0. Discuss the suggestion that the wages vary between Shimla and Chandigarh ? H0 : x = y i.e., the wages do not vary between Shimla and Chandigarh. n1 = 1000, x = 2.50, S1 = 1.50 n2 = 1500, y = 2.68, S2 = 2.00 12.50 — 2.581 Z 7 = = 2.57 and (1.50)2    (2.00)2 1000 + 1500 1.90 < 2.57 < 2.58. Thus it may be concluded that at 5%, l.s. we reject the hypothesis i.e. wages vary between Shimla and Chandigarh, but at 1% l.s., we accept the hypothesis i.e., difference between wages of two places is insignificant and existence of variation in wages is not accepted at this level. Example 8. Balls are drawn from a tag containing an equal number of red and white balls, each ball being returned before drawing another. In 2250 drawing 1018 red and 1232 white balls have been drawn. Do you suspect some bias 90 the part of the drawer ? Solution : Your first job is to ascertain values of p, q and n. Red and white balls being equal 1 1 in number, probability of a red ball or p = 2 ; q = 2 n the total number of drawing = 2250. 1 The expected number of red balls in 2250 drawings = 2 × 2250 = 1125. The actual number of red ball = 1018s The numerical difference between expected and actual frequency = 1125 - 1018 = 107. The standard deviation of simple sampling = ^npq) = ^(2250) x (2 x 2) = 2.7 ( The difference 107 is about 4.5 times L, 7 time I the o and there is hardly any probability . that arose due is fluctuations of sampling. Almost definitely the drawer is biased against red balls (May be a victim of commission phobia!). The above solution is sufficient for the examination. But how shall we look at its terms of our normal curve. The expected number of red balls, 1125 lies at the point of maximum frequency of the normal curve. When sample of 2250 drawings each are taken, actual number of red balls in some sample will be greater than the expected number 1125 and in others less. The actual numbers of red balls in each of the many samples will be spread in a symmetrical normal way around the expected numbers. As usual 99.73% of such actual numbers of red balls will lie within mean ± 3o, o in this case = 23.7. Therefore 99.73% of sample will be such whose red ball drawing lie within 1125 ± 3 × 23.7 i.e., within 1054 and 1196. The number of red ball drawn in Example 8 lies outside these limits. Hence, it is extremely unlikely dial the difference has arisen due to sampling fluctuations. Example 9. A group of scientists reported 1705 sons and 1527 daughters do these figures confirm the hypothesis that the sex ratio is 1 : 1. Solution : Your first job is to write the values of p, q and n. The hypothesis is that the sex. ratio is 1:1. 1 1 :. p, the probability of a son = 2 and q = 2 n, the number of events = 1527 + 1705 = 3232. Number of observed male births = 1705. Number of expected male births. 1 according to the hypothesis = 2 × 3232 = 1616. The difference between expected and observed number = 1735 - 1616 = 89. The standard deviation of simple sampling = A(npq) = ^i3232 x 2 x ^j = A (808) = 21.43 approximately The observed difference 89 is more than three time (about 3.13 time) this standard deviation, therefore it is unlikely that it arose on account of fluctuations of simple sampling. Please note again that a categorical statement has been made. It has been stated clearly that probability of the difference having arisen due to sampling fluctuations is extremely small. The question asked was : “Do these figures-confirm to the hypothesis that the sex ratio is 1 : 1 ? After writing as we have done above, just state in the end that the figures do not confirm to the hypothesis that sex ratio is 1:1. This problem can also be solved in terms of proportion rather than numbers : 1 The expected male ratio = 2 = 5. 1705 3232 . The observed male ratio = The difference between the expected and the observed male ratio. 1705 1616    8 = .075 3232 3232  3232 The standard deviation of proportion. = ^ (pqn) 11  1 = 0088. X—X —+--- 2 2 3232 The difference between the observed and the expected proportion is more than three times (3.13 times) the standard deviation of proportion. Therefore, it is improbable that the difference arose on account of sampling fluctuations. The figures as such do not confirm to the hypothesis of 1 :1 sex ratio. You can reconstruct this problem also in term of a normal curve. Self-Check Exercise-1 Q1. Describe testing the significance of the sample mean. Q2. A random sample from Shimla of 1000 men shows their mean wage to be Rs. 2.50 per day with a standard deviation of Rs. 1.50. A sample of 1500 men from Chandigarh gives a mean wage of Rs. 2.68 per day with a standard deviation of Rs. 2.0. Discuss the suggestion that the wages vary between Shimla and Chandigarh. 14.4 Standard Error : In sampling problem, standard deviation of simple sampling has to be used again and again. It would be convenient if a shorter name could be used for this quantity. This shorter name is STANDARD ERROR. Do not attach much significance to the word Error here. The use of the word Error is justified here by the fact we usually regard the expected value at the true value and divergences from this excepted value in certain observations and samples are regarded as errors of estimation due to sampling effects. For our purpose standard error will just mean standard deviation of simple sampling though the term can be used in a slightly wider sense. From now onwards, we shall use the term Standard Error instead of standard error of simple sampling. In the solution of problems on sampling we should use p and q of the universe rather than of sample. This is what we have done up to now. But sometimes p and q of the universe are not known. In such cases p and q of the sample are taken asp and q of the universe. The assumption is justified if n is large, say 100 or more, and neither p nor q is very small. You will get only those problems where these assumptions hold. Another course open to us when p and q of the universe are not known is to the highest value of p × q. The highest value of p × q = 111 2 x 2 - 4. When either p or q is found to be rather small we can take p = f = 4 in order to calculate the standard error. The question on meaning of standard error and its importance and usefulness in sampling studies is very popular with the examiner. I trust you understand, this question and can answer it. The main use of standard error is to find whether the difference between observed and expected value or between one observed value and another is significant or not. If the difference is more than 3 times the standard error chances are that it could have arisen due to sampling fluctuations. Example 10. 400 children are examined in a town; and 150 are found to be under weights. Assuming conditions of simple, sampling, estimate the percentage of children who are underweight in that town. Population p and q are not known, we shall take sample p and q as population p and q. The probability of p, of getting an underweight children. 150 = 3 5 7 q - - 8 400 = 8 No. of children examined or n = 400 Standard error of the proportion of under-weight children / pq\^(3 I n ) V 8 51 x — x--- 8 400 - 0.24- 2.4% Thus the limits, within which the percentage of under-weight children will probably lie are 37.5 ± 3. S.E. or 37.5 ± 3 × 2.4, i.e. 30.3 and 44.7 per cent. The probability of his statement, being correct is 99.73% since 99.73 area is covered between mean 3 c. If we wanted the probability of our statement to be correct by 99%, the limits would have been 36.5 ± 2.5758 × 2.4 37.5 ± 19.6 × 2.4 would give us limits which are true in 95% cases. 14.4.1    Standard Error and Size of the Sample: We know that S.E. = ^f pq 1 V n J The value of S.E. thus depends on p and n (q is automatically covered in p since q is always (1—p). The size of the universe is not important in the value of S.E. It is affected by the size of the sample. Given the value p, greater the value of n, smaller shall be the value of S.E. and vice versa. The value of S.E. varies inversely as the square root of n. If n increase 4 times the value S.E. is 1 (■      1 1                      1                                                                                      1 reduced by 2 I i.e. ^4 J . Ifn remains 9 th of its former values, of S.E. shall go up, 3 times ifp = 2 1           /I 111 } q = 2 , n = 100, S.E. = 'I 2 x 2 - y^ I = .05 or 5% if n becomes 400. 1111   1 Ì S.E. = 'I^y-T^l = .005 or 2.5% \          Tvv J 14.4.2    Standard Error and Precision : From what we have studied already is should be clear that greater the standard error, greater is the departure of observed value from expected values. 3       5                                     it 3 5      1 In Example 10. Supposep remains — q = — but n = 135. Then S.E. = 'I s x 8 x 88           \ 88135 4.16% app. The limits of under-weight children will then be 37.5 ± 3 = 4.16 i.e. 25:22 and 49.9%. Making the statement that limits of under-weight children are 25.02 and 49.98% is much less precise than that they are 30.3 and 44.7%. In this way standard error gives us an idea about the unreliability of an estimate in a sample. The estimate and is sometimes called precision. This reciprocal of S.E. standard error. 1 S.E. is a measure of reliability measure Precision, however, is not very much used in sampling studies. Example 11. A random sample of 500 pineapples was taken from a large consignment and 65 were found to be bad, Estimate the proportion of had pineapples in the consignment, as well as the standard of the estimate. Deduce that the percentage of bad pineapples in the consignment almost certainly lies between 8.5 and 17.5 (UPSC) 65 The proportion of bad pineapples in the sample 500 = 13%. In the absence of any other information, the proportion of bad pineapples in the sample can be taken as an estimate of bad pineapples in the consignment as a whole. :. the estimated percentage of bad pineapples in the consignments = 13%. p the probability of bad pineapples = 0.13 q = 0.87 n = 500 Standard error of bad pineapples in sampling distribution. 11 pq l     /I 13x 87  i     /I 113 l      । or S.E. = ^1 — 1 = ^1        1 = ^1 — I = ^000226 = .015 or 1.5%. V n J    V  500  J    V500 J Whenever the true proportion of bad pineapples in the consignment, 99.73% changes are (which means it is almost certain) that it lies within a limit of three times the standard error on either side of the estimate. : The limits within which the proportion of bad pineapples almost certainly lie are. Estimate ± 3.S.E. or 13 % ± 3 × 15% or    13% ….. 4.5 % or    8.5% and 17.5 % Self-Check Exercise-2 Q1. Define standard error. Q2. A random sample of pineapples was taken from a large consignment and 65 were found to be bad. Estimate the bad pineapples in the consignment as well as the standard error of estimate. Deduce that percentage of bad pineapples in the consignment almost certainly lies between 8.5 and 17.5. 14.5 Sampling Distribution To understand properly the ideas and assumptions on which studies of this type are based, it is necessary to develop and understand some theoretical considerations. Please understand that in sampling of variables there is no question of p and q, or success and failure. Our members, of samples can now take any value out of (theoretically at least) infinite values, and not one of the two attributes as in sampling of attributes. Suppose we take a random sample of 121 Persons from adult male population of India and find the Arithmetic Mean of their heights. The A.M. could be some value say 65 ". We take suppose 200 such samples, each of 121 persons. We shall have 200 values of Arithmetic mean, like 67 ", 66", 64", 66", 68", 67", 65", 66", 67", and so on (Some values will be in fractions also). These values can be classified and grouped in a frequency distribution. This distribution will be called sampling Distribution of the Arithmetic Mean....... We can calculate standard deviation of heights of 121 persons in each sample. We shall thus get 200 standard deviations. The distribution of these values will be called Sampling Distribution of standard deviation. Coefficient of correlation between heights and weights in each sample could be calculated. We would get 200 coefficients of correlation. Their distribution would be called Sampling Distribution of correlation coefficient. We can thus have .sampling Distribution of any statistical measure, e.g. sampling distributions of median, mode, quartile mean deviation etc. There is one very important characteristic of all. sampling distributions and it is that they give a more or less, normal distribution. If the number of samples used in, the sampling distribution is large and the size of each sample is large the sampling distribution would be a normal distribution even though the parent distribution from which the samples have been drawn is not normal. This is an important and useful characteristic which forms the basis of sampling studies. Since the sampling distribution resembles a normal distribution, it can be used to estimate the values of the population from the values of sample and we can then lay down the limits within which the observed or actual ‘values will probably lie. You will remember that there is a mathematical relationship’ between the area, covered within the mean ordinates and the ordinates at various distances from the mean ordinate, 99.73% area is covered with in (mean × 3σ), 99% area is enclosed within (mean ± 2.5758σ), 95% area is enclosed within (mean ± 1.96σ). The area outside these limits are only 27%, 1% and 5% respectively. In sampling distribution of means, 99.73% chances are that the actual mean lies within mean ± 3σ (or 27% or 0027 that the actual mean lies outside these, limits); 99% chances are that the actual mean lies within mean ± 2.5758% σ (or 1% or .01 that it lies outside these limits)! or chances are 95% that the actual mean lies within mean ± 1.96σ for 5% (or .05 that it lie outside these limits). 14.5.1    Simple Sampling of variables : As in sampling of attributes, our study here can be made, only on the assumption of simple sampling. The conditions of ‘simple sampling’ in the case of variables are same as in attributes. Only they, have to be worded differently in this context. 1.    The drawing of each number of the sample is independent of draws of all other member, and each member of our sample is drawn ‘from the same records. 2.    The different samples are being drawn from the’ same record. The standard deviation of a sampling distribution is called the Standard Error, here also. In actual practice, ‘sampling distribution are not available and we have to estimate parameters or statistical measures for the universe from the values of one sample only. In such cases we shall estimate the mean of standard deviation or other measures of the universe and shall then establish the limits with in which these (values of the universe) can be expected to vary with some specified probability. Self-Check Exercise-3 Q1. Define a)    Sampling Distribution b)    Simple sampling of variables 14.6.    Standard Error of the Mean : Standard error of the mean is the standard deviation or the sampling distribution of means. It is calculated by the formula c(population) Standard Error of the Mean = n n here stands for the number of items in the sample. If standard deviation c of the population is not known, standard deviation of the sample can be substituted in its place, provided the Sample size is large say n > 30, Then S.D. (Sample) S S.E. of the mean =      r     _ ~/= nn Suppose in a sample of 121 persons, the average height is 68" and standard deviation is 5.5". Within what limits do we expect the average height of the, population to exist ? 5.5'' S.E. of the Mean = s _ 5-5'' _ 5-5 _ 5" n (121)    11    . We can now say that the average height of the population from which the sample was taken at random is expected to lie within the range (Mean ± 3 S.E., i.e, 68" ± 3 × 5 which is between 66.5 and 69.5". (Please note here we do not say that the height of the members of the population shall lie with, in 66.5" and 69.5". Only the average of their heights will be within this range. The heights of the individuals may well lie outside this range. 14.6.1    Sampling errors and levels of significance : The term sampling error is used to indicate the error at a certain level of significance. Let us first, be clear what we mean by ‘level of significance’. When we take the limits (Mean ± 3 S.E. 99.7% of the cases are covered in these limits; and only .27% cases are left out. The probability of parameter or population value lying within (mean ± 3 S.E.) is 99.73% and the probability of its lying outside the limits is 27%. This .27% is the level of significance when the limits are (Mean ± 3. S.E.) And at 27% level of significance, 3 is the critical value. Similarly, the probability of parameter lying between (Mean ± 2.5758 S.E.) is 99%. The level of significance of these limits is 1%. The critical value at 1% level of significance is 2.5758. For .5%level of significance the limits would be given by (Mean ± 1.96 S.E.), 1.96 is the critical value at 5% level of significance. You will note that the probability of our statement being correct is 95% at 5% level of significance, 99% at 1% level of significance and 99.73% at 27% level of significance. This may seem paradoxical and our commonsense may not easily accept it, but it is true that the level of significance is inversely related to the extent of precision. Further sometimes the term level of confidence is used in place of level of significance. When our level of confidence is 5% we shall be correct in 95% of cases, but when our level of confidence is less say 1% we shall be correct in 99% cases. The limits which are obtained at a certain level of significance or confidence are called the Confidence Intervals. In our illustration of heights in the previous section, the confidence intervals at 27% level of significance are 6 8" ± (3'.5) i.e. 66.5" and 69.5". At say 5% level of Confidence the confidence interval shall be (68% ± 1.96 × 5) i.e. 67.02" and 68.98". With the explanation of such terms as level of significance or confidence, critical values and confidence intervals, sampling error becomes an easy thing to understand. Sampling error is simply equal to the critical value multiplied by the standard error. Critical value at 5% level of confidence is 1.96. Therefore, sampling error at 5% level of confidence is (1.96 × S.E). Sampling error at 1%. level of significance is (2.5751 × S.E.), and at 27% level of confidence” (3 × S.E.). Find the sampling errors at the various levels of Significance for the illustration given in the previous section. By the way confidence intervals would be given by Mean ± Sampling error, which is the same thing as (Mean + Critical value = S.E.). The most difficult part of the sampling of variables is over. Please go over these pages again in order to understand the theoretical functions of the practical problems that follow. Example 12. An investigator wants to make a survey of the mean weekly wage of 10,000 workers of an industry. Since the study of all the workers is impossible, a representative sample of 400 workers is selected. The mean weekly wage of 400 workers is Rs. 30 and the standard deviation Rs. 2.50, If additional samples were taken by how much would the results differ from the above sample ? Solution : The last sentence of the problem amounts to this : “Within what range is the mean weekly wage of the population expected to lie?” Size of the universe. In the case 10,000 workers is irrelevant and shall not be used anywhere. Standard error of the mean weekly wages of the sample is S.E. of Mean a _ 2.5 _ 2.5 n 400   20 8 _ 0.125Re 99.73% chances are that the mean weekly wages of the additional samples would lie between Rs. (30 ± 3× .125) or Rs. (30 ± .375) or Rs. 29.625 and Rs. 30.375. The confidence interval will be different if we consider 5% or 1% level of significance instead of 0.27% as above. Example 13. It has been determined that the average pulse rates of male in the 20-25 age group is 72 beats per minute and that the standard deviation is 8 beats per minute. If a group of 100 distance runners, all in the given age group of 20-25 were examined and found to have an average pulse rate of 68, should this be regarded as significant deviation from the general average ? Solution: Standard deviation of the population, 8 is available should be used. S.E. of the average pulse rate of 100 distance runners is S.E. of mean = g _ 8 n 100 — _ .8. 10 The difference between the two average pulse rates in 4 beats per minute which is 5 times the standard error. Hence the deviation of average pulse rate of distance runners from the general average is significant. Example 14. The following are the data of height measurements for a random sample or individual from a certain population. Mean = 60",         n = 81, o = 4.5". What would be the limits of the 1% confidence interval for the true mean ? (It is given that Mean ± 2.57580o covers 99% area of a normal curve). Solution: Standard error of the mean height of 81individuals is _ ^ _ 4.5 _ 4.5 _ S.E. mean = = = -          = -5 n 81    9 At 1% level of confidence, the limits of confidence intervals will be given by (Mean ± 2.6778 + S.E.) or 66" ± 2.5758 × .5 or 66" ± 1.2879 or 64.712)" and 67.2879". They can be approximated to 64.7" and 67.3". 14.6.2    Standard error of coefficient of correlation: The standard error of the coefficient of correlation, or _   1-r2 S.E. (r) = An Example 15. Asample of 400 fathers and sons gives a correlation coefficient between their heights as +8. It additional samples were taken from the same universe, between what limits would the coefficients of correlation in the case of those samples vary ? The standard terror of the coefficient of correlation between the height of the father and the son is S.E. (r) = n 1 - r2     1 - 82 400 1-64 = .36 = 0.18 20    .20 The correlation coefficient of other samples would lie between r + 3 S.E. or 8 ± 3 × .018 i.e. between .746 and 84. Standard errorsof almost every statistical measures are given in your text books. Example 16. If 60 new entrants in a given university are found to have a mean height of 68.60 inches, and 50 seniors a mean height of 69.51 inches, is the evidence conclusive that the mean height of the seniors is greater than that of the new entrants ? Assume the standard deviation of height to be 2.48 inches. Solution: The observed difference between the mean heights of the two samples (new entrants and seniors) is 69.51-68.60 = .91" The two independent samples come from the same universe. Standard error of the difference of the two mean heights or S.E. (m1 = m2) = ( 1     1 Ï p2 -+- I n1   n2 J g p is given as 2.48" n1= 60, n2 = 50 ,ox2    ( 1      1 S.E. = (m1 - m2) = ^j(2-48) X! 60 + 50 ( 5 + 6 Ï ^ 6.1504 [ 300 J = ^('2255) The observed difference .91 is less than two times this standard error and could, therefore have arisen due to fluctuations of simple sampling. Self-Check Exercise-4 Q1. Define a)    Standard Error of the Mean b)    Sampling Error c)    Level of Significance Q2. A sample of 400 fathers and sons gives a correlation of coefficient between their heights +8. If additional sample were taken from the same universe, between what limits would the coefficient of correlations in the case of those samples vary? 14.7    Student ‘t’ Test 14.7.1    The Student ‘t’ Test and its Properties The probabilities of the ‘t’ distribution have been tabulated by W.S. Gosset, who wrote under the pseudonym Student which gave, the name to the ‘t’ W.S. Gosset found that the distribution of standard deviation of small samples departs systematically from a normal distribution. Therefore the technique, meant for large samples cannot be applied to small samples. It is well known that if the sample is sufficiently large (n > 30) the estimates are adequate for the applicationof the Z bˆ1 transformation Z = when the sample is small n < 30 and provided that the standard error of (b1 ) population of the parameter is normal, another test can be applied based on the students Y distribution. The general formula which transforms the value of any variable X into ‘t’ units is similar to the Z transformation, but the ‘t’ values depends in addition on the number of degrees of freedom and it includes the variance estimates S2x. The transformation formula (t statistic) is t = ^1 Sx A with n – 1 degrees of freedom where u. = value of the population mean S2x = Sample estimate of the. population variance3 S2x = S(x - x ) ----- . 1 n -1 n = sample size. The sampling distribution in this case, that is the distribution of the sample mean, is x = N (U.S2 x ) and the transformation statistic is ( x -u)/ 2 x2 SN, and has at ‘t’ distribution with (n -1) degrees of freedom. The ‘t’ distribution is always symmetric, with m equal to zero and variance (n — l)/n-3, which approaches unity when n is large. Clearly as n increases, the ‘t’ distribution approaches the standard Normal distribution Z » N (0. 1). The t distribution is independent of population parameters U and a2. The definition of t is independent of a’ and therefore, it will not be necessary to know o’ (before, using this distribution. This is the chief merit of the t distribution. Before its discovery it was usual to replace the unknown by S in. 7 x - M        x - m Z         togett < 60, the fact t is a symbolically a standard, normal variate can be used. 14.7.4 . Properties off H’ Distribution (1)    The graph of the t distribution is lower, at the centre and high at tails as shown in fig. 12.1. (2)    The t statistic ranges from - « to + ». (3)    For values of v > 60, the fact, t is asymptotically a Standard normal variate can be used. As N approaches unity, the t distribution approaches, normal curve. (4)    The t distribution is similar to the normal curve since it is single peaked at and symmetrical about, a zero mean, for the case in which are under the distribution is unity. (5)    The table of ‘t’ falls rapidly as the number of degrees freedom from 1 to 10 and vary, slightly as degrees of freedom increases from 10 to 40. (6)    All the moments of odd order about the origin vanish. The moment or order 2r about the origin (which is also, the mean) is w ^2 r = 2 J 0 t2         dt (t 2 Y1/2 w •J 0 12r - I v) d(t2 / v) ^2 r 1 + t 2 J ^ v+1) I v ) ( t 2 Y+1/2-1 vr w J 0 5 k v) ( t 2 A ^ v+1) 1 + — 2 ( tv J k v) fl (    1 v A     v r +—,--r , r < — k   2 2 )    2 t2 On letting 1 +--= — i.e. — 1 t 2 v Y v 1 - Y Y Hence ^2 r = (2r -1) (2r - 3)......1 (v - 2) (v - 4)(v - 2r) Self-Check Exerxise-5 Q1. Explain the t test for testing the significance of the difference between two sample means. Q2. State the properties of t distribution. Q3. A random sample of 27 pairs of observation from a normal population gave a correlation coefficient 0.6. Is this significant of correlation in the population? 14.8    Snedecor’s F Distribution In the preceding section we discussed the methods of determining whether two samples have come from the same universe or from two universes which are significantly different from each other. One of the method by which this study is done is by the calculation of standard error of the difference of the means of two samples; another method is the x2 test’ and in case of small samples the method that is generally followed is that of t test. Here we shall discuss Snedecor’s ‘F’ distribution which shows that the distribution of the ratio of independent estimates of the population variance. We know that variance Σ(x - x)2 V = σ2 = n 7 Σ(x - x)2 n -1 degrees of freedom in such cases are equal to n -1. In other words, in small samples. = Σ(x - x)2          Σ(x - x)2 σ =         & V= n-1               n-1 There are two types of variations in the data. One between the various samples and the other within the various samples. Now if the variations within the samples and between the various samples are not significantly different from, each other then the samples belong to the same universe. Suppose the values of the items in the four samples were as follows. | | Sample 1 | Sample 2 | Sample 3 | Sample 4 | |---|---|---|---|---| | | 4 | 6 | 12 | 10 | | | 6 | 8 | 16 | 10 | | | 2 | 6 | 14 | 10 | | | 6 | 10 | 8 | 6 | | | 2 | 0 | 20 | 4 | | Total | 20 | 30 | 70 | 40 | | Mean | 4 | 6 | 14 | 8 | Total number of items in the sample or N = 20. T = 20 + 30 + 70 + 40 =160 The grand mean of all the items of all the samples = 160/20 = 8. Total Variations | Squares of the deviations of various items from the grand average of 8 | |---| | Sample 1 | Sample 2 | Sample 3 | Sample 4 | | 16 | 4 | 16 | 4 | | 4 | 0 | 64 | 4 | | 36 | 4 | 36 | 4 | | 4 | 4 | 0 | 4 | | 36 | 64 | 144 | 16 | | Total 96 | 76 | 260 | 32 | Grand total of squares = 96 + 76 + 260 + 32 = 464 Degrees of freedom = 20— 1 = 19 14.8.1    Variance Between The Samples: We shall calculate the square of the deviations of the means of the various samples from the grand average, if the value of each item in the first sample be taken as 4, for the second sample as 6, in the third samples as 4 and in the fourth as, 8 and the squares of the deviations of those values of the grand average are calculated, they would be below: | Sample 1 | Sample 2 | Sample 3 | Sample 4 | |---|---|---|---| | 16 | 4 | 36 | 0 | | 16 | 4 | 36 | 0 | | 16 | 4 | 36 | 0 | | 16 | 4 | 36 | 0 | | 16 | 4 | 36 | 0 | | Total        80 | 20 | 180 | 0 | | | 280 | 280 | | | Variance between the samples is 41 | _ — = 93.7 | | 14.8.2    Variance within Samples Thus in sample 1, the deviation would be taken from 4, in sample 2 from 6, in sample 3 from 14 and in sample 4 from 8. These deviations would be squared & totaled. | | Sample 1 | Sample 2 | Sample 3 | Sample 4 | |---|---|---|---|---| | | 0 | 0 | 4 | 4 | | | 4 | 4 | 4 | 4 | | | 4 | 0 | 0 | 4 | | | 4 | 16 | 36 | 4 | | | 4 | 36 | 36 | 16 | | Total 16 | 56 | 80 | 32 | | Grand total of the sum of squares = 16 + 56 + 80 + 32 = 184 Variance with in the samples 184 _ 184 20 - 4 = 16 All these results can be tabulated as follows: | Source of variation | Sum of squares | D.F. | Variance | |---|---|---|---| | Between Sample | 280 | 3 | 280 _ 93.3 3 | | Within Sample | 184 | 16 | 184 _ 11.5 16 | | Total | 464 | 19 | „ 933 F_---_ 8.1 115 | If should be remembered that F = Variance between samples Variance within samples It has been noted that variance between samples is generally greater than variance within samples. Now if we look at Snedecor’s table for the value F for the given degrees of freedom at 50% level of significance. The calculated value of F is higher than this and as such the difference is significant. (l) (2) (3) (4) The variance ratio of F has a very important, property that its value remains unchanged, if all the figures are either multiplied or divided by a common factor or if a common factor is added to or subtracted from each figure. The numerator and denominator of the second member are independent X2 variates with v1 and v2 degrees freedom respectively. The distribution of F is independent of the population variance o’ and depends on v1 and v2 only. The F curve is J shaped if v2 < 2 and bell shaped for v1 > 2. For v1 > 4 the shape of the F curve is as shown in figure below. The distribution has highly positive skewness. Fig. 12.2 (5) The probability density of the distribution increases steadily at first reaching at the highest peak (corresponding to the model value) and then goes on ecreasing slowly so as to become tangential at infinity. But mode exists if and only, if v1 > 2 and is equal to v2 V’ + 2 v1 - 2 v1 Hence mode of F distribution is always less than 1. (6) v2                ( 11 ^ For Fv,, v, distribution mean d(u) =       and variance (u z 2)     + 1’ 2                          'r|/    v2 — 2                   'r2      7 v V1    V2 J (7)    The distribution free from population parameters. (8)    When v1, = 1, the F distribution becomes equal to that of t2 with v2 degrees of freedom. (9)    When v2 = ». it means S22 = o2. So that v1, F = xv12. (10) 22. When v, = ». it means S 2 = o’. So that -z- -o = xv 1           1          FF (11) For large Vj and v2 F » N 1 1 2 ;+; I V y v 2 14.8.3 Some Properties of the F Distribution and F curve Let x1, (i = 1,2 ….. n1) and x2 J (J = 1, 2 …… n) be the values of two independent random samples drawn from the same normal population with variance o’. Let x 1 and x 2be sample means and let. S12 1 n1 -1 n I( x1/- xi)2 i=1 and S22 1 n2 - 1 n I (X2 j - x2)2 j=1 Then we define the statistic F by the relation 2 F = SI S22 2 Hence v1F v2 (ni — 1) 4 <7 2 (n2 — 1) 4 7 where v1 = n1 - 1, v2 = n2 – 1. The numeratorand denominator of the second member are independent X2 variates with v1 and v2 S.f. respectively. Hence v1 F/v2 is a P2 (v1,/2, v2/2) variate so that the probability that a random value of F will fall in the interval dF is vv^ F-^sf 222      2 p1 (              \                  -(-1+-2) P ( 2’ 2" ) (V1F + V2)2 This distribution is called the distribution of the variance ratio F with v1 and v2 f The r th moment about the origin is given by v 2 ^r=e ( f )= v v2 2 1 f F + - J0      2 F F 1 1 + -1 F V    -2 2 -v1+v 2 2 5F where or A-r k V2 J k v. v_ ^ k 2,2 J 1 f Y 0 V2               y+--1 + - r - 1(1 - 7) 2 1 , v r    1    A        v2 d y 1 + — F = — and dF = —, -^y- v      Y            v Y2 μ1r v1 \-r k V2 J v —+y 2 v2 v1 \ k V2 J 2 y v — + y 2 dy ’ y < v2 2 v1 v. k 2’2 J k v1 J y v v r < — 2 in particular Mi i ~ +1 2 v2 2 1 v2 v2 k v1 J v2 - 2 which is independent of v1, and is always greater than unity. M2 = v1 2 - 2 v1    v2 22 2 V2  =   (v + 2)v22 k v,J   (v2 - 2Xv2 - 4)V1 M2 = M2 - M1 - 2 = 2v1 ( v + V2 - 2) x 14.9    Interrelationship Between σ , t, x2, and F As has been noted earlier in ‘t’ distribution properties that t distribution approaches the normal distribution as n approaches infinity. The normal distribution is therefore a special case of the t distribution. For the same set of data normal distribution yield the same, probabilities as do x2 values when n = 1 for of x2. More specifically, comparing Areas in Two Tails of the Normal Curve at Selected Values of s or σ from the Arithmetic mean and values of x2, that for a given probability x ^2 I ° J = x2, when n = 1 for x2. x2 For any given probability ~ = F. when n for X2 equals n1 for F and when n2 = ^ for f. This can be seen by comparing values of X2 and values of F (from their respective tables). Itis also to be noted that for any given probability, t2 = F. when n for t equals n2 for F and when n2 for is l. This is apparent from an examination of values of t and values of F (from their respective tables). F distribution is an inclusive distribution in that the other three distributions are merely special cases of F. APPLICATIONS (i)    Special Tests For large samples the sampling distributions of many statistics are normal distributions with mean us. and standard deviation cs. In each case the result hold for infinite populations or for sampling with replacement. (1)    Means. Here s = x , the sample mean : u = 4 x = U the population mean : cs = c x = c/ Nw where c is the population standard deviation and N is the sample size; The Z score is given by Z = x - ^ a / Jn when necessary the sample deviations s or sˆ is used to estimate s. (2)    Proportion: Here S = P. the proportion of “successes” in a sample us = u = P where p is the population of proportion of success N = the sample size; cs = cp = c/ 4pn7n where q = 1 - P. The Z score is given by Z = P - p PN /N x In case P = N where x is the actual number of successes in a sample, the z score becomes Z _ X - NP NPq i.e.,   ^a = u = N p oa = a = ^NPq , and S = x. l. Tests of Significance Involving sample Difference Let x1 and x2 be the means obtained in large sample of size N1 and N2 drawn from respectively, populations having means g1 and u. and standard deviation o1 and o2. Consider the null hypothesis that there is no difference between the populations means, i.e. h = b Hx - x2= Hx - Hx = h = ^2and 2   22 px - x = pP+ + P x = Pi2 + p2_ N1 N2 …. 1 The above equation is when S1 and S2 are the Sample means from the two populations S, which we denote by x1 and x2 , then the sampling distribution of the differences of means is given for infinite population with, mean and standard deviation gp o2 and g2 respectively as in equation I. H x - x2 = 0 and 0x - x2 = 2 £1 \ V Ni J + 2 £ V N 2 J Where we can, if necessary, use the sample standard deviations S1 and S2 (or sˆ1 and sˆ2 ) as estimates of o1 and o.. By using the standardized variable or Z score given by Z _ Xi - X2 - 0 oX1 - X2 = X - X 0X1 - X2 We can test the null, hypothesis, against alternative hypothesis (or the significance of an observed difference at an appropriate level of significance). (ii) . Differences of Proportions Corresponding results can be obtained for the sampling distributions of differences of proportions from two binomially distribute populations with parameters P1, q1, and P2, q2 respectively. In this case S1 and S2 correspond to the proportion of successes, P1 and P2, and equation ^s - S2= us - ^s and o2 - S2. J a 2, + a 2., yield the results. S1      S2 ^p1-p2 = ^p1-p2 = P1 - P2 ' P1qr + P2q2 N1 N 2 and ° ■ p = T^ï+^J = If N1 and N2 are large (N1, N2 > 30) the distributions of differences of means or proportions are very closely normally distributed. We noted that the sampling distribution of differences in proportions is approximately normally distributed with mean and standard deviation given by ^p1-p2 = 0 and °p1-p2 N P.+ N + p Where P = —1P---2---2 Where       —! + —2 pq —+— N N L N1 N 2 7 is used an estimate of the population proportion and q = 1- P. By using the standardized variable P - P -0 P - P Z _ P1 P2  0 _ P1 P2 O’P1-p2      O - °2 And can test observed differences at an appropriate level of significance and thereby test the null hypothesis. Tests of Means and Proportions Using Normal Distributions Example 1. Find the probability of getting 40 and 60 heads inclusive in 100 tosses of a fair coin. Solution : (a) according to binomial distribution the required probability is 40      60                   41      59 + 100 C60 100 C40| - I I - I +100 C4,1 - I I - I + 4012 7 L 2 7         41L 2 7 12 J (17             17 Since Np = 100 I 2 I = nq = 100 I 2 I are both greater than 5, the normal approximation to the binomial distribution can be used in evaluating this sum. The mean and standard deviation of the number of heads in 100 tosses are given by g = NP = 100 (1/2) = 50 a = 7Npq = 7100x 1/2 x 1/2 = 5 On a continuous scale, between 40 and 60 heads inclusive is the same as between 39.50 and 60.5 heads. 39.5 in a standard units (39.5 - 50) 5 - 2.10 _ (60.5 - 50) - 2.10 60.5 is standard units ------5----- Required probability = area under normal curve between Z = 2.10 and Z = 2.10. = 2 (area between Z = 0 and Z = 2.10) = 2 (0.4821) = 0.9642. (b) To test the hypothesis that a coin is fair, the following rule of decision is adopted: (1)    Accept the hypothesis if the number of heads in a single of 100 tosses is between 40 and 60 inclusive. Example 2. Reject the hypothecs otherwise (a)    Find the probability of rejecting the hypothesis when it is actually correct. (b)    Interpret ‘graphically the decision rule and the results of part (a). (c)    What conclusions would, you “draw if the samples of 100 tosses yielded 53 heads? 60 heads? (d)    Could you be wrong in your, conclusions to (c)? Explain. Solution: (a)    According to above problem the probability of not getting between 40 and 60 heads inclusive if the coin is fair = 1- 0.9642 = 0.0358. Then the probability of rejecting the hypothesis when it is correct = 0.0358. (b)    If a single sample of 100 tosses yields a Z score between -2.10 and 2.10. We accept the hypothesis otherwise, we reject the hypothesis and decide that the coin is not fair. The error made in rejecting the hypothesis when it should be accepted is the Type I error of the decision rule. Fig. 12.3 If a single sample of 100 tosses yields a number of heads whose Z score (or Z statistic) lies in the shaded region, which shows that the score differed significantly from what would be expected if the hypothesis were true. For this reason the total shaded area (i.e., Probability of a type I error) is called the level of significance of the decision rule and equals 0.0358 in this case. Thus we speak of rejecting the hypothesis at 0.0358 or 3.58% level of significance. (c) According to the decision rule, we would have to accept the hypothesis that the coin is fair in both cases. (d) Yes, we could accept the hypothesis when it actually should be rejected, as would be the case. For example, when the probability of the need is really 0.7 instead of 0.5. Self-Check Exercise-7 Q1. Describe the interrelationship between x/a, t, X 2 and F. 14.10    Summary In this unit, we explored the concepts of standard error and Student’s t-distribution, focusing on their significance in hypothesis testing, especially for small samples. We discussed how standard error relates to sample size and precision, and detailed the application of the Student’s t-test for testing sample means. Additionally, we covered the properties of the t-distribution and Snedecor’s F distribution and explained the interrelationship between x/a, t, X 2 and F distributions. 14.11    Glossary •   Standard Error: The measure of the variability of a sample statistic, often the sample mean, from the population parameter. •    Student’s t-Test: A statistical test used to determine if there is a significant difference between the means of two groups, particularly useful when sample sizes are small. •    t-Distribution: A probability distribution used in hypothesis testing for small sample sizes, characterized by heavier tails than the normal distribution. •    Snedecor’s F Distribution: A probability distribution used in the analysis of variance (ANOVA) to compare variances between groups. •    Sampling Distribution: The probability distribution of a given statistic based on a random sample, used to make inferences about the population parameter. 14.12    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 14.3.1. Answer to Q2. Refer to Section 14.3.1. Self-Check Exercise-2 Answer to Q1. Refer to Section 14.4. Answer to Q2. Refer to Section 14.4.2. Self-Check Exercise-3 Answer to Q1. Refer to Sections 14.5 and 14.5.1. Self-Check Exercise-4 Answer to Q1. Refer to Sections 14.6 and 14.6.1. Answer to Q2. Refer to Section 14.6.2. Self-Check Exercise-5 Answer to Q1. Refer to Section 14.7.1. Answer to Q2. Refer to Section 14.7.4. Answer to Q3. t = 3.75 (Significant) Self-Check Exercise-6 Answer to Q1. Refer to Section 14.8. Answer to Q2.Refer to Section 14.8.3. Answer to Q3: H0 : a12 = o/ : F = 1.54 < 3.73 ; Not Significant Self-Check Exercise-7 Answer to Q1. Refer to Section 14.9. 14.13    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Mill, F. C. (1955). Statistical methods. Pitman and Sons: London. 14.14    Terminal Questions Q1. Two samples are drawn from two normal populations. From the following data, test whether the two samples have the same variance at 5 % level (Use F0.05 – 3.68 of v1 = 9 and v2=7): | Sampel 1 | Sample 2 | |---|---| | 60 | 61 | | 65 | 66 | | 71 | 67 | | 74 | 85 | | 76 | 78 | | 82 | 63 | | 85 | 85 | | 87 | 86 | | - | 88 | | - | 91 | Q2. A dean of school of social science claims that the social sciences students have better English and Science students. A researcher conducts a test to verify the claim. She randomly selects 25 students from each educational stream for the test. The science and social science students on an average mark receive 55 and 60 marks with 4 and 9 as their standard deviation, respectively. Test Dean’s claim at 5% significance level. ***** UNIT-15TESTING HOMOGENEITY OF SEVERAL INDEPENDENT ESTIMATES OF POPULATION VARIANCE STRUCTURE 15.1    Introduction 15.2    Learning Objectives 15.3    Tests of Significance Based on t Distribution 15.3.1    Assumptions for the Test Self-Check Exercise-1 15.4    Testing the Significance of Observed Correlation Coefficient Self-Check Exercise-2 15.5    Testing the Significance of an Observed Partial Correlation Self-Check Exercise-3 15.6    Testing the Significance of an Observed Regression Coefficient Self-Check Exercise-4 15.7    Summary 15.8    Glossary 15.9    Answers to Self-Check Exercise 15.10    References/Suggested Readings 15.11    Terminal Question 15.1    Introduction In this unit, we will discuss methods for testing the consistency of several independent estimates of population variance. We will start by exploring how to test the significance of observed correlation coefficients, which measure the relationship between two variables. Additionally, we will cover testing the importance of observed partial correlation coefficients, which account for additional variables, and regression coefficients, which show the effect of one variable on another. Dear Student, Example 1: A random sample of 9 steel beams has an average compressive strength of 55,815 pounds , per square inch and a standard deviation of 200 pounds per square inch. Test the hypothesis that the true average strength of the Steel beams from which this sample was taken is 56,000 pounds per square inch, using a two sided alternative and a level of significance 0.05. Solution: We treat the hypothesis to be tested namely µ = 56,000 pounds per square inch (PSi) as the null hypothesis. The alternative hypothesis, being two sided is µ 56,000 pSi. Since we are using a small sample consisting of 9 steel beams, the appropriate distribution is the t distribution having (9.1) or 8 degrees of freedom, shown in figure 13.1 below :- Fig. 13.1 The t value corresponding to a level of significance 0.05 ‘two tailed’ test’ is found to be 2.31, from the table in which ‘t’ values are for different levels of significance and degrees of freedom are given. The difference between the sample mean and null hypothesis mean is now expressed in term. X-μ t= σ where x and µ are the sample and hypothesis means respectively and σ, the sample Vn-1 standard deviation. 558150-56000 200 8 -185× 8 200 -2.62 Since this ‘t’ value exceeds that t0.05 value for 8 degrees of freedom we reject the null hypothesis and conclude that the true mean compressive strength of the lot of the steel beams from which the sample has been taken is not, 56,600 psi. 15.2    Learning Objectives After completion of this unit, you will be able to •  Understand the concept of homogeneity of population variance. •    Learn the assumptions required for tests of significance. •    Test the significance of observed correlation and partial correlation coefficients. •    Test the significance of observed regression coefficients. 15.3    Tests of Significance based on t distribution We shall consider five tests of significance based on the t distribution. Example 2. Show that 95% fiducially limits for the mean µ, of the population are x ± St0.05/ Vn. . Deduce that for a random sample of 16 values with mean 41.5 inches and the sum of the squares of the deviations from the mean (135) (inches)2 and drawn from a normal population 95% fiducially limits for the mean of population are 39.9 inches and 43.1 inches. Solution: We have p or p X-μ/ n S X-St 0.05 — t0.05 = .95 — g — X = .95 So that we can say with a confidence .95 that the confidence interval X ± S t0.05 contains n the population mean u. The limits of this confidence or fiducially limits for u. In the particular case n = 16, v = n - l = 15. S = J — x 135 = 3 15 Also from the table, t0.05 = 2.13 n3 Sto.o5 = 4 x2.13 = 1.6 Approx. Therefore the required fiducial limits are 41.5 ± 11.6 i.e. 39.6 and 41.3 inches. Example 3. Two yields of the types ‘Type A’ and ‘Type B’ of cereals in kgs per hectare in 6 replications are given below. What comments would you make on the difference in the mean yields? You may assume that if there be 5 degrees of freedom and P = 0.2, t is approximately 1.476. | Replication | ‘A’ | ‘B’ | |---|---|---| | | Yield in kgs | Yield in kgs | | 1 | 20.50 | 24.86 | | 2 | 24.60 | 26.39 | | 3 | 23.06 | 28.19 | | 4 | 29.98 | 30.75 | | 5 | 30.37 | 29.97 | | 6 | 2383 | 22.04 | | Solution:— | | Replication A (yield | B (yield d       Deviation from Sq. of | | in kgs) | in kgs)              mean (5— d)     deviations | | | | (d- d )2 | | 1              20.50 | 24.86       4.36 | +2.717        7.38 | | 2            24.60 | 26.39       1.79 | +0.147        0.02 | | 3             23.06 | 28.19       5.13 | +3.487        12.15 | | 4 | 29.18 | 30.75 | 0.77 | - 0.873 | 0.76 | |---|---|---|---|---|---| | 5 | 30.37 | 29.97 | -0.40 | - 2.043 | 4.17 | | 6 | 23.83 | 22.04 | -1.79 | - 3.433 | 11.79 | 36.27 Sd = 9.86 d = mean of difference of the yields Sd 9.86 . = — =---= 1.643kgs. N0 S = S.D. of difference of yields E(d - d )2 n -1 36.7 = J —5— = 2.69 kgs. We set up the null hypothesis, that the difference in type has no effect on yield i.e., the population mean of the difference is zero then t = 1.643 - 0 2.69 Vä=1.489 This value of t is less than t0.05 for 5 d.f. and therefore, the difference is not significant at 20% level. Give two independent random samples from normal population with the same variance we have to test the hypothesis that the population means are p.1 & p2 respectively. For this case t = (x-^MMizEd S 11 + n1   n2 t = with (n1 + n2 -2) d.f. and the test of significance is carried out as d-0 15.3.1    Assumptions of the Test Note about assumption for the test: For testing significance of a sample mean, t test has a greater range of validity and precision than the normal approximation tests of x - h) a n1 <1 -96 but this is achieved by placing an important restriction viz., that the present population must be normal. If the parent population is not normal the distribution of the statistics Vn (x - u S) may be quite different from the r distribution, since, in that case, the distribution of x and S2 will not be independent and also the distribution will depend on parameters giving the departure of parent distribution from normality. For testing the significance of the between two sample, means we make an additional assumption that the variances of the two populations are the same. Before applying the t test it may be desirable to test this assumption by applying the F test. If the two variance, are different t test is no longer valid and another test ‘d’ based on confidence intervals applies by Behren can be used d = X1 - X2 V s2+s2 then we carry out the test of significance by using the tables of Sukhative and Fisher for various values of n1, n2 and S12/S22, the values of d corresponding to various probability levels. Self-Check Exercise-1 Q1. Two yields of the types ‘Type A’ and ‘Type B’ of cereals in kg per hectare in 6 replications are given below. What comments would you make on the difference in the mean yields? You may assume that if there be 5 degrees of freedom and P = 0.2, t is approximately 1.476. | Replications | ‘A’Yields in Kgs | ‘B’yields in Kgs | |---|---|---| | 1 | 20.50 | 24.86 | | 2 | 24.60 | 26.39 | | 3 | 22.06 | 28.19 | | 4 | 29.98 | 30.75 | | 5 | 30.37 | 29.97 | | 6 | 22.83 | 22.04 | 15.4    Testing the significance of observed correlation coefficient Given a random sample (x1, y1), (x1, y2)   (xn yn) from a bivariate normal population, here we have to test the hypothesis that the correlation coefficient of the population is zero i.e., variables are independent. If me hypothesis is true, it is foe case that in an uncorrelated population i.e. when P = 0; the distribution of correlation-coefficient, r, can be obtained easily as follows: Let the n values i (i = 1,2....n) be subjected to an orthogonal transformation yielding n undependent varieties n2, n2 ...... nn so that nn Z n2=Z y2 i=1 choosing n1 = _yn ^yi = xy^ (the sum of the squares of fee coefficient of y1' (s is unity) we get nnn Zn2=Zy2 =Z(yi - y+y)2 1=1         1=11 n = Z(Y1 - Y)2 = nY2 = nS2 + n i=1 n or nS = £2 n2 i=1 Also nr rs3= E( x, - x )(Y-Y ) n S1 n rs2 n Hence £ n2 = n (1 — r2)S2 i=3 Since n1/o2 are independent standard normal variates we conclude that nr2S22 2 and ^2 n (1 - r 2) S2 are distributed independently like X2 wife 1 and x – 2 degrees of freedom respectively. r2S22 _         nr2S2/&2 S2    nr2 S22/ &2 + n(1 - r 2)S22/ &22 = x,2+x 2 where x12 and x22 are distributed like X2 with 1 and (n – 2) d.f. respectively and therefore r2 is ^2 now r 2 β 1 n — 2 2 2 variate. Hence the distribution of r2 is 5p = (r 2 ) 1 —1(1 — r 2) n—4 ß 1 n — 2 2, 2 d(r) 0-1 < r2 < 1 Thus the distribution of r is dp = —-----=r (1 — r2) n—- dr — 1 < r < 1. a 1 n — 2           2 ß 2’ 2 If the hypothesis is true the critical ratio ‘t’ is defined by the expression. r4n - 2 t=~-r^ That the ‘t’ is a variate with (n-2) d.f. If our calculated value of (t) exceeds, to .05 for (n - 2) d.f. we say that the value of r is significant at 5% level of significance, f 111 < to .05 the sample is consistent with the hypothesis of an uncorrelated population. Example 4. A random sample of 15 from a normal universe gives a correlation coefficient of + 0.60. Is of the existence of correlation in the population significant? Here n = 15, r = - 0.60 ।   +^5-2 =+ 2.7o 71 - (-0.60)2 Also for 13 d.f. t0.05 = 2.16 .'. Sample correlation coefficient is significant as the calculated value of t is more than the table value. Hence our hypothesis that sample has been taken from an uncorrelated brvariate normal population appears to be incorrect. Example 5. A random sample of 18 pairs from a bivariate normal population showed a correlation coefficient of 0.3. Is this value significant of a correlation in the population ? Solution : We set up the hypothesis that the variables were uncorrelated in the normal population. = r^n -2 = 3718 - 2 =   1.2 t " 71 - r2 " T"1 - (3)2 " 71 - .09 1.2 = — = 1.26. .91 The number of degrees of freedom = n -2 = 18 - 2=16 The value of t from the table for 15 d.f. at 5% level of significance is 2.12. Thus the calculated value of t is less than the table value. Hence our hypothesis that sample has been taken from an uncorrelated bivariate normal population appears to be correct. Example 6. It was found that the correlation coefficient between two variables, calculated, from a sample of size 25 was 0.40. Does this show evidence of having come from a population with zero correlation? Solution: We set up the hypothesis that the sample has come from a normal population with zero correlation. r - 0 1 = 71 - r2 x V n - 2 .40 71 - (.40)2 x V 25 - 2 .40 = ■ '4O   x 4.7958 71 - .16 1.9183 .84 x 2.09 Approx. The number of degrees of freedom = 25 - 2 = 23. The table value of t for 23 d.f. at 5% level of significance is 2.07 so that the calculated value is higher than the table value. Hence our hypothesis appears to be incorrect. We are likely to conclude that the sample has not come from a population with zero correlation. Example 7. Two samples of sizes 23 and 28 give ‘r’ as 0.5 and 0.85 respectively. Is there any significant difference between the two correlation coefficients ? (1 + r A Solution : Z. = 1.1513 log... H I i                            11r < 1 +1/2A = 1.1513 log10 I 1—jj2 I = 0.55 ( 1 + r A Z2 = 1.1513 lOg1o |J ( 1 + 8 A = 1.1513 log10 I —I = 1.10 10 11 — 8 J t = Z1   Z2 11 + \ r1 — 3   r2 — 3 = 1.10 — .55 = .55 1     1     0.3 + 20 25 = 1.83 < 6.96 Hence the difference is not significant at 5% level. Example 8. Find the least value of r in a sample of 18 pairs from a bivariate normal population significant at 2% level. Solution : Substituting the value of ‘n’ in the formula V n — 2 V1 — r2 x r we get 4r t = TT7 Now for significance of ‘r’ at 2% level, t should be greater than the value of t from the table at 2% level for 16 degrees of freedom which is 2.58. 4r 41 - r2 > 2.58 r2 > (.645)2 (1-r2) r2 > 0416025 (1-r2) 1.416025 r2> 0.416025 0.416025 1.416025 > 0.54 Hence the required least value of r is 0.54 Ans. Example 9. Twelve pictures submitted in a competition were remarked by shown in the table below. | Picture | A | B | C | D | E | F | G | H | I | J | K | L | |---|---|---|---|---|---|---|---|---|---|---|---|---| | Rank assigned by first judge | 5 | 9 | 6 | 7 | 1 | 3 | 4 | 12 | 2 | 11 | 10 | 8 | | Rank assigned | 5 | 8 | 9 | 11 | 3 | 1 | 2 | 10 | 4 | 12 | 7 | 6 | be second judge Calculate p. Is there a lack of independence in these rankings? (Assume that on the hypothesis of independence of two sets of n ranking _ -\l n — 2 t = p i 2 follows the t distribution with (n - z) degrees of freedom V1 — P Given that Degrees of freedom         10           11           12 Value of t significant at 5% level of probability      2.23          2.20          2.18 Solution : Calculation of p, the rank correlation coefficient Rank difference d: 0.1 -3, -4, -2, 2, 2, 2, -2, -1, 3, 2. :. S d2 = 0 + 1 + 9 + 16 + 4 + 4 + 4 + 4 + 4 + 4 + 1 + 9 + 4 = 60 n - number of pictures ranked = 12 6Sd2           6 x 60 i 360 ” p = 1 - n(n2 — 1) = 1 - 12(122 — 1) = 1 - 143 x12 = -791 n — 2 t=p 1—Pp 12 — 2 •791 V — (-791)2 = 791 × 5.17 = 4.089 number of degrees of freedom = n – 2 = 12 – 2 = 10 The value of t for 10 degrees of freedom at 5% level of significance is 2.3. :. The calculated value of t is greater than the table value and hence deviation is significant. Therefore, the ranking is significant. Self-Check Exercise-2 Q1. A random sample of 15 from a normal universe gives a correlation coefficient of + 0.60. Is of the existence of correlation in the population significant? Q2. Find the least value of r in a sample of 18 pairs from a bivariate normal population significant at 2% level. 15.5    Testing the significant of an observed partial Correlation Fisher has shown that the sampling distributor of a partial correlation coefficient or order K (i.e., with K secondary scripts) is of the same form as that of correlation coefficient from a bivariate normal population with the sample size n reduced by K. In particular; if we are given a random sample from a multivariate normal distribution and we have to test the hypothesis that a particular partial correlation coefficient of order K in the population is zero, we use of the fact. t = —,     4 n - k - 2 l—n is a t variate with (n - k -2) d.f. Example 10. Show that in a sample of size 20 from a normal population, a correlation coefficient r12.3 = 0.5 is significant at 5% level. Hare r = 0.5, n = 20, k = 2 0.5            2    4 3   6.928 t = ,       716 = -i= =---=---- i 1 - (.5)2         V/75     3       3 = 2.31 Approx. As for 16 d.f. to .05 = 2.12. Also | t | + t0.05 the value of r123 is significant at 5% level. Self-Check Exercise-3 Q1. For given set of numerical data, the following results are obtained. r12.3 = 0.65 : n = 15 Test the significance of partial correlation coefficient at 5% level of significance. 15.6    Testing the Significance of an observed regression coefficient: Suppose we are given a random, sample x1 y1, x2 y2 ......(xn yn) from a bivariate normal population. The regression equation of y on x is obtained from the sample be y - y = b (x - x) So that the estimated value y corresponding to given x1 is y - y = b (x1 - x) We have to test the hypothesis that the regression coefficient of y on x or in the population is β. If the hypothesis is correct the t statistic is t — (b - B) (n - 2) £ x1 k j —x 7 Conforms to t distribution with (n - 2) degrees of freedom. Example 11. For a sample of size 30 where x takes the values 1,2, 3.... 30 it is found that. S(x1 - x) (y1 - y) = 599.62 and S(y1 - y )2 = 10.206. Test the significance of the regression coefficient of y and x. w               n(n2 — 1) 30x 899 Solution : Here 2(x2 - x )2 =--—— = ——— = 224 75 599.62 ------ — .2668 2247.5 b — 2( xi- x)(yi- y) 2(x1 - x )2 Also 2(y - yi) = 2(71 - y )2 - b2 S(V x )2 = 1020.6 - 159.97 = 860.59 Also the hypothesis gives B = 0. Substituting in t — (b - B) (n - 2) £(x, - x )2 _ i £ [(h- y,)2 ]1 t = 2.28 with v = 28 The value of t is significant at 5% level but not at 1% level. Example 12. The data showing the aptitude test scores of a random sample of salesman and their first year sales in rupees is given in Table 4 below. Table 4 - Aptitude Test Scorns and first year sales of Salesman | Salesman | Aptitude test score         First year sales (in 000’s) | |---|---| | A | 52                         108 | | B | 49                        92 | | C | 66                        136 | | D | 44                        109 | | E | 71                         107 | | F | 59 | 149 | |---|---|---| | G | 53 | 84 | | H | 74 | 190 | | I | 38 | 82 | | J | 66 | 157 | | K | 76 | 154 | | L | 32 | 84 | We can plot these values on a graph showing the test scores along the X-axis and the corresponding sales on the Y-axis, this diagram is called, “the scatter diagram.” Worksheet for Use estimation of constants a and b. | n | Aptitude test score Xi | First year Sale in (000 Rs) yi | xi (original value-55) | x12 | x1 y1 | y12 | |---|---|---|---|---|---|---| | 1 | 52 | 108 | -3 | 9 | -324 | 11,604 | | 2 | 49 | 92 | -6 | 36 | -552 | 8464 | | 3 | 66 | 136 | 11 | 121 | 1496 | 18,496 | | 4 | 44 | 109 | -11 | 121 | -1199 | 11,881 | | 5 | 71 | 167 | 16 | 256 | 2672 | 27,889 | | 6 | 59 | 149 | 4 | 14 | 596 | 22,201 | | 7 | 33 | 84 | 22 | 484 | -1848 | -7056 | | 8 | 74 | 190 | 19 | 361 | 3610 | 36,100 | | 9 | 38 | 82 | -17 | 289 | -1394 | 6,724 | | 10 | 66 | 157 | 11 | 121 | 1661 | 22,801 | | 11 | 76 | 154 | 21 | 441 | 3234 | 23,716 | | 12 | 32 | 84 | -23 | 529 | -1932 | 7056 | | | 2y1 = 1506 | 2x1 = 0 | 2x = 2784 | 2xi yi = | 6020 2y12=204, 048 | few? $0 ¿0 " Ji ^        i              S 70 Mmw£ nwrscate A straight line can be represented in the general form y = a + bx where a is the y intercept corresponding to x = 0 and ‘b’ is the slope of the line, rate of change of y for unit change in ‘x’. if we can obtain the values of the constants ‘a’ and ‘b’ in this equation, the relationship is determined (i)    Σy1, = n a + b Σxi (ii)    Σx1 y1 = 9 Σx1 + b Σx12 An intuitively understandable explanation of these equation is as under: y1, = a + bx2 i.e. y1 = a + bx1 y2 = a + bx2 yn = a + bxn Summing up both sides Y1 + Y2 + .......Yn = na + b (X1 + X2 + X3 ....... Xn) ΣY1 + na + b Σxi the first equation Similarly, if we multiply each equation by X, we get, X1 Y1 = ax1 = bx12 X2 Y2 = ax2 + bx22 Xn Yn + a× x + bxn2 Summing up both sides Σx1y1 = aΣx1 + bΣx12, the second equation. Solving these two equations we can obtain Σx12ΣYi- Σxi- ΣxiΣyi nΣx12 - (Σx1 )2 nΣxiyi- Σxiyi nΣxi2 - (Σxi )2 Σx2 ΣY -0xΣx y i i          ii nΣxi2 - (Σxi)2 Σxi2 ΣYi    ΣYi nΣx12 n b= nΣxi yi - 0x Σyi nΣxi2 - 0 Σxiyi Σxi2 Substituting these values ΣY a= i n 1506 = 125.5 12 b= Σxi yi 6020 Σxi2 2784 = 2.16 The regression equation can be written as Y 125.5 + 2.16 X where 125.5 is the value of the intercept on Y axis when x1 = 0, i.e., corresponding to the value 55 in the original data. The standard deviation (denoted by symbol Sy.x) of the Y variable is given by the expression. S(y JiT S yx n - 2 A more convenient form of the same relationship for case of calculation is given by S y:x Sy2 - aSy - bSxy \ n - 2 The denominator (n - 2) shows the number of degrees of freedom, since two degrees of freedom are lost because two quantities a and b have been estimated. The suffix Sy means that we are referring to the variability of the value of Y corresponding to given values of x. Sys for data is worked out as under S Sy2 - aSy - bSxy n - 2 12.04 x 0.48 -125.5 x 1506 - 2.16 x 6020 10 2.042 10 = 4204.2 = 14.3 We can estimate his first year sales by substituting X = 60 in the regression equation Y = 6.7 + 2.16 × .60 = 136.3 (in thousand rupees). For example we can compute the 95.5 percent confidence for the first year sales of salesman whose aptitude test score was 60. The 95.5 percent confidence interval will be 136.2 ± 2 × 14.3 (in thousands of rupees) Or 136.3 ± 28.6 (-do-) i.e. between 107, 700 and 164,900. Self-Check Exercise-4 Q1. For a sample of size 30 where x takes the values 1, 2, 3.....30. It is found that S(x1-x) (y1- y ) = 599.62 and S(y1- y )2 = 10.206. Test the significance of the regression of y and x. 15.7    Summary In this unit, we explored methods for testing the homogeneity of several independent estimates of population variance. We began by discussing the necessary assumptions for conducting tests of significance based on distribution. Next, we learned how to test the significance of observed correlation coefficients, which measure the relationship between two variables. We also covered testing the significance of observed partial correlation coefficients, which control for additional variables, and regression coefficients, which indicate the effect of one variable on another. 15.8    Glossary •  Population Variance: A measure of the spread of a set of values in a population. •    Correlation Coefficient: A statistical measure that describes the strength and direction of a relationship between two variables. •    Partial Correlation: The correlation between two variables while controlling for the effect of one or more additional variables. •    Regression Coefficient: A measure that indicates the amount of change in a dependent variable due to a change in an independent variable. 18.9 Answers to Self-Check Exercise Self-Chek Exercise-1 Answer to Q1. Refer to Example 3. Self-Check Exercise-2 Answer to Q1. Refer to Example 4. Answer to Q2. Refer to Example 8. Self-Check Exercise-3 Answer to Q1. t = 2.632, r12.3 = 0.65 is significant of partial correlation in the population. Self-Check Exercise-4 Answer to Q1. Refer to Example 11. 15.10    References/Suggested Readings 1    Siegel, S., & Castellan, N. J., Jr. (1988). Non-parametric statistics for the behavioral sciences. McGraw-Hill. 2    Kapur, J. N., & Saxena, H. C. (1969). Mathematical statistics. S. Chand & Co. 15.11    Terminal Questions Q1. A random sample of 27 pairs of observations from a normal population gives a correlation coefficient of 0.42. It is likely that the variables in the population are uncorrelated. Q2. If the regression coefficient of y on x is 0.40, then find the regression coefficient of x on y. ***** UNIT-16ANALYSIS OF VARIANCE STRUCTURE 16.1    Introduction 16.2    Learning Objectives 16.3    Analysis of Variance Self-Check Exercise-1 16.4    The Variance Ratio (F) 16.4.1    Two Criteria of Classification Self-Check Exercise-2 16.5    Comparison of Regression Analysis and Analysis of Variance Self-Check Exercise-3 16.6    Summary 16.7    Glossary 16.8    Answers to Self-Check Exercise 16.9    References/Suggested Readings 16.10    Terminal Questions 16.1    Introduction In this unit, we will discuss the Analysis of Variance (ANOVA), a statistical method used to compare the means of three or more groups to see if there are significant differences between them. ANOVA helps in understanding whether the observed differences in sample data are due to actual variation or random chance. This unit will guide you through the basics of ANOVA, the F-ratio, and its comparison with regression analysis, providing practical exercises to enhance your understanding. 16.2 Learning Objectives After going through this unit, you will be able to •    Understand the purpose and process of Analysis of Variance (ANOVA). •    Learn how to calculate and interpret the variance ratio (F). •    Identify the criteria for classifying data in ANOVA. •    Compare and contrast regression analysis with ANOVA. 16.3    Analysis of Variance The analysis of variance (ANOVA) is a statistical method developed by R.A. Fisher for the analysis of experimental data Initially, the main application (ANOVA) was confined to the analysis of agricultural experiments but later on, the use of this technique expanded to many other fields of scientific research. The total variance of a variable can be decomposed into different additive components with the help of analysis of variance which may be attributed to various separate factors. Let us explain with the help of an example that there are twenty plots of land on which rice is cultivated. It is proposed to study the yield per unit of land. Let us farther suppose that different seeds, different fertilizers and different means of irrigation are used. Thus the variation in yields may logically be attributed to the three factors. x1 = type of seed x2 = type of fertilizer x3 = type of irrigation. The total variation in yield can be broken down into three separate components : a component due to X1, another due to X2 and third due to X3 with the method of analysis of variance: Analysis of variance is conceptually the same as regression analysis, if we go by the above definition of (ANOVA). Inregression analysis total variation is in the explained variable is split into two components : the variation explained by the regression line, and the unexplained variations shown by the scatter of point around the regression line. However, it must be noted that tee are significant difference between regression analysis and analysis of variance. The main difference is that regression analysis provides numerical values for the influence of the various explanatory factors on the dependent variable, in addition to the information concerning the breaking down of the total variance of Y into additive components, while the analysis of variance provides only, the latter type of information. The main objective of regression analysis and ANOVA is the determination of the various factors causing variations of the dependent variable. The method of ANOVA is used in regression analysis for conducting various tests of significance, the most important being : (1)    The test of the overall significance of the regression. (2)    The test of the significance, of the improvement in fit obtained by the introduction of additional explanatory variables in the function. (3)    The test of equality of coefficients obtained from different samples. (4)    The test of the extra-sample performance of a regression, or test of the stability of the regression coefficients. (5)    The test of restrictions imposed on coefficient of a function. An experiment with four samples : During Cooking doughnuts absorb fat in various amounts. Mrs. X wished to learn if the amount absorbed depends on the type of fat use. For each of four fats, six of batches of doughnuts were prepared batch consisting of 24 doughnuts. Data of this kind are called a single or one way classification each fat representing one class. Before we start, it is to be noted that the totals for the four fats differ quite a lot; from 372 for fat 4 to 510 for fat 2. Grams of Fat Absorbed Per Batch | Fat | 1 | 2 | 3 | 4 | Total | |---|---|---|---|---|---| | | 64 | 78 | 75 | 55 | | | | 72 | 91 | 93 | 66 | | | | 68 | 97 | 78 | 49 | | | | 77 | 82 | 71 | 64 | | | | 56 | 85 | 63 | 70 | | | | 95 | 77 | 76 | 68 | | | S X | | | 432 | 510 | 456 | 372 | 1770 = G | |---|---|---|---|---|---|---|---| | X | | | 72 | 85 | 76 | 62 | 295 | | SX2 | | | 31,994 | 43652 | 35,144 | 23,402 | 134,192 | | (sx) N | 2 | | 31,104 | 43,350 | 34,656 | 23,064 | 132,174 | | Si2 = | Sx2 - I | (SX)2 ^ N J | 890 | 302 | 488 | 338 | 2018 Pooled S12 | | d.f. | | | 5 | 5 | 5 | 5 | 20 | | | | | Pooled S2= | 2,018/20 = | 100.9 | | | | | | | Sp 252/ | n = 22(100 | - 9)/6 = 5.80 | | | The analysis of variance is of great utility and flexibility and was developed by Fisher in the 1920’s. The analysis of variance performs two functions : (1)    It is an elegant and slightly quicker way of computing the pooled S2. In a single classification on this advantage in speed is minor, but in the more complex classifications, the analysis of variance, is the only simple and reliable method of determining the appropriate pooled error variances S2. (2)    It provides a new test, the F-test. This is a single test of null hypothesis that the population means g1, u2, u., u4, for the four fats are identical. This test is often useful, in a preliminary inspection of the results and has many subsequent applications. We can compute the total sum of squares of deviation for the 24 observations as _                                 (17702 642 + 722 + 682 + 702 + 682 - I 24 = 134, 192 - 130, 538 = 3654. 1.1 This sum of squares has 23 degrees of freedom. The mean square, 3654/23 = 158.9, is the first estimate of a2. The second estimate is pooled S2 already obtained. Within each fat, we computed the sum of squares between batches (890, 302 etc.), each with .5 d.f. These sum of squares were added to give 890 +302 + 488 + ….. = 2018 1.2 | Source of variation | Degrees of Freedom | Sum of Squares | |---|---|---| | Between fats Between batches within fats | 3 20 | 1,636 2,018 | | Total | 23 | 3,654 | The degrees of freedom and the sum of squares for the two components (between fats and within fats) add to the corresponding total figures. These results hold in any single, classification, the result for the difference is not hard to verify with classes and n observations per class, the d.f. are (a-1) for between fats and a (n-1) for with in fats, and (an-1) for the total. But (a-1) + a (n-l) = a-l + an-a- an- 1 The standard practices in the analysis of variance is to compute only the total sum of squares and the sum of squares between fats. The sum of squares within fats, leading to the pooled s2 is obtained by subtraction. The symbol T denotes a typical class’ total, while G = ΣT = ΣΣX (summed over both rows and columns) is the grand total. The first step is to calculate the correction for the mean, C = G an (1770) 2 24 130.538. This quantity is called the sum of squares between batches with in fats, more concisely the sum of squares within fats. The sum of squares is divided by its degrees of freedom, 20 to give me second estimate s2 = 2, 018/20 = 100.9. For the third estimate, consider the mean for the four fats, 72, 85, 76 and 62. These are also estimates σ2 of µ but have variances 6 since they are means of samples of 6, their sum of squares of deviations is 2 762 + 852 + 762 + 622 - (   ) = 272.75 4 with 3 degrees of freedom. The mean square, 272.75 3 is an estimate of µ consequently if we multiply by 6, we have the third estimate of µ we shall accomplish this by multiplying the sum of squares by 6, giving. (295) 2 6 {722 + 852 + 762 + 622 - ^4^ = 272.75} = 1636 1.3 1636 the mean square being 3 = 545.3. Since the total for any fat is six times the fat means, this sum of squares can be computed from the fat totals as 4322+5102+4562+3722   (1770) 2 - 6               24 = 132, 174- 130, 538 = 1656                           ……. 1.4 This sum of square is called the sum of squares between fats. This is done because c occurs both in formula 1.1 for the total sum of squares and in formula for the sum of squares between fats. The remaining steps should be clear from Table 3. Formula for calculate the Analysis of variance Table | Source of Variation | Degree of Freedom | Sum of Squares | Mean Square | |---|---|---|---| | ‘Between | a - 1 = 3 | | 'st2 V n > | -C = 1.636 | 543.3 | | classes (fats) Within classes (fats) | a (n-l) = 20 | Subtract 2,018 | 100.9 | | Total | an – l = 23 | SX2 - | C = 3,654 | | Self-Check Exercise-1 Q1. Q2. Explain the meaning of “Analysis of Variance”. A test was given to 5 students chosen at random from the M. A Economics Class of each of the three universities in Bihar. Their scores were found as follow: | University | Scores | |---|---| | A | 90 | 70 | 60 | 50 | 80 | | B | 70 | 40 | 50 | 40 | 50 | | C | 60 | 50 | 60 | 70 | 60 | Perform Analysis of Variance and show if there is any significant difference between the scores of students in the three universities. 16.4 The variances ratio F F = Mean square between classes Mean squares within classes should be good criterion for testing the null hypothesis that the population means are the same in all classes. The value of F should be around 1 when the null hypothesis holds, and should became large when the r1 differ substantially. The distribution was first tabulated by fisher in form 2 = log e F . In honour of Fisher, the criterion was named F by Snedecor. Fisher and Yates designate F as the variance ratio, It should be noted that when there are only two classes, the F test is equivalent to the t test, which is used to compare the two means. With two classes, the relation F = t2 holds. An experiment comparing two groups of equal size. Mr. X compared the 15 day mean comb weights of two lots of male chicks, one receiving harmone A, the other C. Day old chicks, 11 in number, were assigned at random to each of the treatments) To distinguish between the two lots, which were caged together, the heads of chicks were stained green and black respectively. The individual comb weights are given in Table. 1 Prove that Harmone A gives higher average comb weight than harmone C. Table 1 Testing the Differences Between the Means of Two Independent Samples Weights of Comb (rags) | Harmone | Harmone | |---|---| | A | C | | 57 | 89 | | 120 | 30 | | 101 | 82 | | 137 | 50 | | 119 | 39 | | 117 | 22 | | 104 | 57 | | 73 | 32 | | 53 | 96 | | 68 | 31 | | 118 | 88 | | Total                         1,067 | 616 | | N                        11 | 11 | | Mean                   97 | 56 | | SX2                        111,971 | 42,244 | | (SX)2 | | | 103,499 | 34,496 | | n | | | S2 '    SX               8,472 | 7,748 | | d.f.                                10 | 10 | 8472+7748 pooled s2             = 811 d.f = 20 10 +10 S x1-x2 252 n 2x811E 11    = 12.14mg. t = SX1 (X1 -X2) x2 41 12.14 = 3.38 Analysis of Variance with only two classes. The period S2 = 16220 2 = 811, has already been computed. To complete the analysis of variance, compute the between samples sum of squares. Since the sample totals were 1067 and 616, with n = 1, the Sum of squares is (1067)2 + (616)2 11 - (1683)2 = 9245.5 22 With only two samples, this sum of squares is obtained more quickly as (Σx1 - Σx2)2    (1067-616)2 =             = 9245.5 2×11 2n Analysis of Variance | Source of variation | Degrees of freedom | Sam of square | Mean square | |---|---|---|---| | Between samples | 1 | 9,245.5 | 9,2455 | | within samples | 20 | 16.2200 | 811.0 | 9245.5 F =        11.40 F = 3.48 811.0 The value of F is significant at 1% level showing that Harmone. A gives higher average comb weights than harmone C. The two sum of squares of deviations 8,472 and 7,748, make the assumption of equal σ2 appear reasonable. The 95% confidence limit for (µ1 - µ2) are x1 – x2 ± t0.005 Sx1–x2 or, in this example 41 - (2.086) (12.14) = 16mg. and 41 + (2.086) (12.14) = 66mg. Example : Below are given the yield of three strains of rice planted in five randomized blocks. Prepare the table of analysis of variance. | Blocks | |---| | Strains | I | II | III | IV | V | | A | 20 | 21 | 23 | 16 | 20 | | B | 18 | 20 | 17 | 15 | 25 | | C | 25 | 28 | 22 | 28 | 32 | We set up the hypothesis that there is no difference between the strains. Taking 20 as the origin, the given table becomes. | Blocks | |---| | Strains | I | II | III | IV | V | Total | | A | 0 | 1 | 3 | -4 | 0 | 0 | | B | -2 | 0 | -3 | -5 | 5 | -5 | | C | 5 | 8 | 2 | 8 | 12 | 35 | | Total | 3 | 9 | 2 | -1 | 17 | 30 | Here T = 30 Wj2 = 4 + 25 + 1 + 64 + 9 + 9 + 4 + 16 + 25 + 64 + 144 = 390 N = 15 Sum of squares between strains STi2 _ T2 nj _ 0 + 25 + (122)2 515 = 250 – 60 =190 Total sum of squares T2 =W - y = 390 (30)2 15 = 330 Sum of squares within strains = Total S.S - S.S between strains = 330 -190 = 140 M.S. = Analysis of Variance Table | Source of variation | D.F. | S.S. | M.S. | F | F at Level 1 %  5% | |---|---|---|---|---|---| | Between strains | 2 | 190 | 95 | 8.14 | 693   389 | | Within strains | 12 | 140 | 11.67 | | | | Total | 14 | - | - | - | - | S.S D.F The observed F being greater than the value of F at, 1% for 2, 12, d.f. we reject our hypothesis The difference between the strains is significant. 16.4.1    Two criteria of classification We classify blocks not only, according to type of soil but also according to the type of seed used on them. The question then arises whether the crop yield varies with the type of block or with the type of seed. Here the variations in the yield of crop may be attributed to (i)    Variation in crop yield due to type of soil. (ii)    Variation in crop yield due to type of seed. (iii)    Variation in crop yield due to error term. For carrying out the analysis of variance, we could compute total variations and variation between columns (type of soil). There is no variation with the columns (type of soil). But there is a variation between rows (variation due to type of seed) and residual variation : 1.    Total variation (Total sum of square’s) mn £x)2 N XX (x-x )2 X'2 -i=1 j=1 2.    Variation between columns (S.S between, columns) cc X[nc( Xc - x )| X 1 1 n X X k i 7 - Nc (^x )2 N c Where X 1 stands for summation of ‘C’ columns m X 1 stands for the summation of N items in a column Nc stands for the no of items in a column Xc stands for the column mean X stands for the grand mean 3.    Variation between rows (sum of squares between rows) RR X Nr( X, - X )2 -X 11 n X k 1 X 7 Nr (^x )2 N where R X i stands for the summations of ‘R’ rows. N Z 1 stands for the summation of ‘N’ item’s in a row. Nr stands for the no of items in a row. stands for the mean of row r 4.    Residual variation = Total variation - Variation between columns Variation within rows. Table of Analysis of Variance - Two criteria of classification | Sources of | Sum of | Degrees of | Estimates of | Variance | |---|---|---|---|---| | Variations | Squares | Freedom | Variance | Ratio-I | | (1) | (2) | (3) | (4) | (5) | | Bet columns | S.S.C. | C-1 | SSC -----= m. C -1    1 | F1 = m1 m3 | | Bet Rows | S.S.R. | R-1 | SSC = m. C -1    2 | F1 = -m2- m3m | | Residual | S.S.Re. | (C-1) (R-1) | SSRe | | | (C - 1)(R-1)   3 | | | Total | T.S.S. | N-1 | - | - | Example : Three varieties A, B, C, of a crop are tested in a randomized block design with four replications, the lay out being given in diagram appended. The plot yield in kgs. are also indicated therein. Analyze the experimental yield and state your conclusions. | A | 6 | C | 5 | A | 8 | B | 9 | |---|---|---|---|---|---|---|---| | C | 8 | A | 4 | B | 6 | C | 9 | | B | 7 | B | 6 | C | 10 | A | 6 | Solution | Blocks | |---| | Varieties | I | II | III | IV | Total | | A | 6 | 4 | 8 | 6 | 24 | | B | 7 | 6 | 6 | 9 | 28 | | C | 8 | 5 | 10 | 9 | 32 | | Total | 221 | 15 | 24 | 24 | 84 | T =84 YYxij2 = 36 + 49 + 64 + 16 + 36 + 25 + 64 + 36 + 100 + 36 + 81 + 81 = 624 T2(84) Correction Factor = — == 588 n 12 Total sum of squares = 624 - 588 = 36 Sum of squares between varieties SV12 4N 242+282+322 .oo 588 4 = 596 - 588 = 8 Yßi2   T2 3n Sum of squares between blocks 212 +152 + 242 + 242 =588 3 = 606 - 588 = 18 :. Error sum of squares = 36 - (18 + 8) = 10 | Analysis of Variance table | |---| | Sources of Variation | D.F. | S.S. | M.S. | F 5.143 | F at 5% level 4.757 | | Between verities | 2 | 8 | 4 | 2.4 | | | Between Blocks | 3 | 18 | 6 | 3.6 | | | Error | 6 | 10 | 1.667 | | | | Total | 11 | | | | | Since the calculated value of F is less than the table value in both cases, we conclude that the variation between the verities and between blocks are significantly not different from the variance due to random errors. Self-Check Exercise-2 Q1. Set up an analysis of variance for the adjoining data for three varieties of wheat each grown on 4 plots. State if the variety differences are significant. | Plot of Land | Variety of Wheat | |---|---| | | A | B | C | | 1 | 6 | 5 | 5 | | 2 | 7 | 5 | 4 | | 3 | 3 | 3 | 3 | | 4 | 8 | 7 | 4 | Q2. The following data are collected to compare the average yields per acre of three variables of wheat: | Varieties | Average Yield/Acre | |---|---| | A | 44 | 40 | 18 | 19 | | | | B | 24 | 21 | 14 | 18 | 29 | | | C | 19 | 17 | 17 | 18 | 20 | 23 | At the 5% level of significance, test whether the average yields of the three varieties of wheat differ significantly. 16.5    Comparison of Regression Analysis and Analysis of Variance As has been explain in the beginning that the basic difference between regression analysis and ANOVA is that while the former provides numerical values for the influence of the various explanatory, factors on the dependent variable, in addition to the information concerning the breaking down of the total variance of y into additive components, while the latter provides only the breaking down of the total variance into additive components. Firstly in both methods the total variation in y is split into two additive components: (a)    Regression analysis SY2 = SY2 + Se2 Total variation = (Explained by repressors) + (unexplained or residual) (b)    Analysis of variance k nj                       k                         k nj XX (Y - Y)2 = Xnj (Y - Y)2 +XX (Y - Y ) ji                   ji Total variation = Between + Within Secondly. The test performed in the method of analysis of variance concerns the equality between means of sub-groups of sub-samples of an enlarged population. That is, the null hypothesis being tested is H0 = ^1 = ^2 =     ^n and alternative hypothesis is H1 : g1 not all equal. The F* ratio is a test of significance of R2 _ SY2/K -1 F* = —:----- Se2/N - K R2YX1/K -1 (1 - R2YX1/N - K) If R2 is not statistically significant it means that there is do linear relationship between Y and X. Thirdly. In both methods we obtain an analysis of variance table, form which F ratios, can be computed and used for testing ‘hypothesis related to the aim of study. Fourthly : As. has been proved earlier that the individual regression coefficients that t and F tests are formally equivalent, the relationship between them being. Lastly. Regression analysis is more powerful method man the ANOVA method when studying economic relationship from market data which are not experimental but stochastic: It is generally believed that. ANOVA method is more appropriate for the study of the influence of qualitative factors on a certain Variable, because qualitative’ variables do not possess, numerical values, and hence their influence, cannot be assessed by regression analysis, while the ANOVA technique does not require information of the values of X’s but it is based solely on the values of Y. This argument losses its power due to the increasing use of dummy variables in regression analysis. However, the ANOVA technique may be incorporated into regression analysis, for carrying out tests of various hypothesis. Merits (1)    The calculations are comparatively simple in ANOVA than in regression analysis. (2)    The technique of analysis of variance makes use of all data. (3)    It does not involve a high degree of abstraction. (4)    It may be applied to all important planning phase of the enquiry as well as to the interpretive phase of enquiry. Self-Chek Exercise-3 Q1. Distinguish between the Regression Analysis and Analysis of variance. 16.6    Summary In this unit, we explored Analysis of Variance (ANOVA), a crucial statistical technique for comparing the means of three or more groups to determine if the differences observed are statistically significant. We delved into the calculation and interpretation of the variance ratio (F), and the criteria for classifying data in ANOVA. Additionally, we compared ANOVA with regression analysis to highlight their differences and applications. The unit included practical exercises to solidify understanding and a glossary to clarify key terms, ensuring a comprehensive grasp of ANOVA and its significance in statistical analysis. 16.7 Glossary •    Analysis of Variance (ANOVA): A statistical method used to compare the means of three or more groups to determine if there are significant differences. •   Variance Ratio (F): A ratio used in ANOVA to compare the variance between groups to the variance within groups. •   Regression Analysis: A statistical method for examining the relationship between a dependent variable and one or more independent variables. •    Significant Differences: Differences between groups that are unlikely to have occurred by chance, indicating a real effect. 16.8    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 16.3. Answer to Q2: F (Universities) = 3.33 4 F (2, 12); Not significant. Self-Check Exercise-2 Answer to Q1: F (Varieties) = 4/2.67= 1.5 4 F (2, 9); F0.05= 4.26 ; Not Significant Answer to Q2: F (Varieties) = 3.50; Not Significant Self-Check Exercise-3 Answer to Q1. Refer to Section 16.5. 16.9    References/Suggested Readings 1.    Chiang, A. C., & Wainwright, K. (2017). Fundamental methods of mathematical economics. McGraw Hill Education. 2.    Cramer, H. (1999). Mathematical methods of statistics. Princeton University Press. 3.    Dunn, O. J., & Clark, V. A. (1987). Applied statistics: Analysis of variance and regression. Wiley. 4.    Grewal, P. S. (1990). Methods of statistical analysis. Sterling Publisher. 5.    Schnecdor, G. W., & Cochran, W. G. (1989). Statistical methods. Wiley-Blackwell. 6.    Spiegel, M. R., & Stephens, L. J. (2008). Theory and problems of statistics. Schaum’s Outlines, McGraw-Hill. 16.10    Terminal Questions Q1. Use Analysis of Variance technique to the data given below: | Foundry | No. of Produced items | |---|---| | A | 84, 60, 40, 47, 34, | | B | 67, 92, 95, 40, 98, 60, 59, 108, 86 | | C | 46, 93, 100 | Q2. A certain company had 4 salesman A, B, C, and D. Each of whom was sent for a week into three types of areas K, O and S. The sales in kg per week are shown below. | District | Salesman | |---|---| | K | A | B | C | D | | 30 | 70 | 30 | 30 | | O | 80 | 50 | 40 | 70 | | S | 100 | 60 | 80 | 80 | Carry out the analysis of variance and interpret the results. ***** UNIT-17INDEX NUMBERS STRUCTURE 17.1    Introduction 17.2    Learning Objectives 17.3    Index Numbers Self-Check Exercise-1 17.4    Problems Involved in the Construction of Index Numbers Self-Check Exercise-2 17.5    Classification of Index Numbers 17.5.1    Price Index Numbers 17.5.2    Quantity Index Numbers 17.5.3    Value Index Numbers Self-Check Exercise-3 17.6    Methods of Construction of Index Numbers 17.6.1    Unweighted Indices 17.6.2    Weighted Indices Self-Check Exercise-4 17.7    Summary 17.8    Glossary 17.9    Answers to Self-Check Exercise 17.10    References/Suggested Readings 17.11    Terminal Questions 17.1    Introduction In this unit, we will discuss the concept of index numbers, which are vital tools in economic analysis for measuring changes in economic variables over time. Index numbers simplify complex data and make it easier to understand trends and patterns in various economic activities. We will delve into the construction, types, and methods of index numbers, as well as the challenges encountered in their development. 17.2    Learning Objectives After going through this unit, you will be able to •    Understand the fundamental concept and importance of index numbers in economic analysis. •    Identify and address the challenges in the construction of accurate index numbers. •    Distinguish between different types of index numbers (price, quantity, and value) and their respective construction methods. 17.3    Index Numbers Index numbers are the indicators which reflect changes over a specified period of time in (i) prices of different commodities, (ii) industrial production, (iii) sales, (iv) imports and exports, (v) cost of living etc. These indicators are of paramount importance to the management personnel of any government organization or industrial concern, for the purpose of reviewing position and planning 252 action if necessary and in the formulation of executive decisions. They reflect the pulse of an economy and serve as indicators of inflationary or deflationary tendencies. Economic index numbers measure the pressure of economic behaviour and are rightly termed as economic barometers or ‘barometers of economic activity’, since a look at some of the important indices little index numbers of whole sale prices, industrial production, agricultural production etc; gives a fairly good idea as to what is happening to the economy of a country. “Index numbers as statistical devices designed to measure the relative change, in the level of a phenomenon (variables or a group of variables) with respect to time, geographical location or other characteristics such as income profession etc.” In other words, these are the numbers which express the value of a variable at any given date called the given period as a percentage of the value of That variable at same standard date called the base period. The variable may be: (i)    The price of a particular commodity e.g. silver, iron etc. or a group of commodities like consumer goods, foodstuffs, etc. (ii)    The volume of trade, exports and imports, agricultural or industrial production, sales in a departmental store etc. (iii)    The national income of a country or cost of living of persons belonging to particular income group/profession etc. For example, suppose we want to measure the general changes in the price level of consumer goods. Obviously, these changes are not directly measurable as the price quotations of various commodities are available in different units, e.g. wheat and sugar in Rs. per quintal, petrol and kerosene oil in Rs. per liter, cloth in Rs. per metre etc. An average price of all these items expressed in different units is obtained by using the technique of index numbers. Self-Check Exercise-1 Q1. What is an index number? 17.4    Problems Involved in the Construction of Index The methods of Construction of index numbers warrant a careful study of the following problems: 1. The Purpose of Index Number: An index number which is properly designed for a purpose can be most useful and powerful tool, otherwise it can be equally misleading and dangerous. Thus the first and foremost problem is to determine the purpose of index number without which it is not possible to follow the steps in its construction. 2.    Selection of Commodities : Having defined the purpose of index numbers, select only those commodities which are relevant to the index. For example, if the purpose of an index is to measure the cost of living of low income group (poor families) we should select only those commodities or items which are consumed/ utilized by persons belonging to this group and due care should be taken not to include the goods/ services which are, ordinarily consumed by middle-income or high income group. For such an index, selection of commodities like cosmetics and other luxury goods like scooters cars; refrigerators, television sets etc, will be absolutely useless. 3.    Data for Index Numbers: The data, usually the set of prices and of quantities consumed of the selected commodities for different periods, places, etc., constitute the raw material for the construction of index numbers. The data should be collected from reliable sources such as standard trade journals, official publications periodical special reports from the producers, exporters, etc; or through field agency. The principles of data collection, viz., accuracy, comparability, sample representativeness and adequacy should be borne in mind. In any case the data should strictly pertain to what is being measured. 4.    Selection of Base Period : The period, with which the comparisons of relative changes in the level of a phenomenon are made is termed as base period and the index for this period is always taken as 100. The following are basic criteria for the choice of the base period. (i)    The ‘base period’ must be a ‘normal period’, i.e. a period free from all softs of abnormalities or chance fluctuations such as economic boom or depression, labour, strikes, wars, floods, earthquakes etc. If the base period be taken as a period of economic instability or depression in which the prices of various commodities and goods, due to their scarcity have been abnormally high then the comparison of price relatives in any given year will not be of much practical utility. (ii)    The base period should not be too distant from the given period. Since index numbers are Essential tools in business planning and in formulation of executive decisions, the base period should not be too for back in the past relative to the given period because due to dynamic pace of events these days, distant base period is likely to be entirely different from the given period. Moreover, if the base year is shifted far away from the given period, it is possible that the pattern of consumption of commodities may change appreciably. 5.    Type of Average to be used : Since index numbers are specialized averages a judicious choice of average to be used in their construction is of great importance. Usually the following averages are used : (i)    Arithmetic Mean (A.M.) : Simple or weighted, (ii)    Geometric Mean (G.M): Simple or weighted, (iii)    Median. Since in the construction of index numbers we deal with ratios or relative changes and since geometric mean gives equal weights to equal ratios of change, does not give undue weightage to extreme observation, Geometric mean is preferred over all other averages. 6.    Selection of Appropriate weights: Generally, various items/commodities say wheat, rice, kerosene, clothing, etc, included in the index are not of equal importance and proper weights should be attached to them to take into account their relative importance. There are two types of indices : (i)    ‘Unweighted Indices’, in which no specific, weights are attached to various commodities, and (ii)    ‘Weighted indices’, in which appropriate weights are assigned to various items. 7.    Choice of Formulae : A large number of formulae have been devised for constructing the index. The problem very often is that of selecting the most appropriate formula. The choice of the formula would depend not only on the purpose of the index number but also on the data available. Fisher has suggested that an appropriate index is that which satisfies time reversal test and factor reversal test. Theoretically Fisher’s method is considered as “Ideal” for constructing index numbers. Self-Check Exercise-2 Q1. Explain the various problems involved in the construction of index numbers. 17.5    CLASSIFICATION OF INDEX NUMBERS Index numbers may be classified into three categories depending on the nature of phenomena under study. 17.5.1.    Price Index Numbers : Price index numbers show the changes in the prices of commodities produced or consumed in a given period with reference to base period. These indices permit comparison of the prices, of commodities between regions, between cities of the same region and between two time periods. These are of two types : (a)    Whole sale price index numbers. (b)    Retail or Consumer price index numbers. Quantity index numbers show the changes in the quantity of goods produced or consumed or purchased in a given period with reference to base period. These indices permit comparison of the quantity produced or purchased of different commodities. 17.5.2    Quantity Index Numbers Quantity index numbers show the changes in the quantity of goods produced or consumed or purchased in a given period with a reference to the base period. These indices permit comparison of the quantity produced or purchased of different commodities. 17.5.3.    Value Index Numbers Value index numbers show, the changes in the value of commodities in a given period with reference to base period. Self-Check Exercise-3 Q1. Write short notes on a)    Price Index Numbers b)    Quantity Index Numbers c)    Value Index Numbers 17.6    Method of Construction of Index Numbers A large number of formulae have been devised for constructing index numbers. Broadly speaking they can be grouped under two heads : (a)    Unweighted Indices: and (b)    Weighted indices 1 7.6.1 Unweighted Indices In the unweighted indices weights are not expressly assigned whereas in the weighted indices weights are assigned to the various items. Each of these types may be further divided under two heads : (i) Simple Aggregative and (ii) Simple Average of Relatives. Types of Index Numbers Unweighted Index Number Weighted Index Number Simple Aggregative  Simple Average Weighted Weighted Average Index Number of Relatives Index Aggregative Index of Relatives Index Notations and Terminology: The notations and terminology used in constructing index numbers are given below: Base year—The year selected for comparison. It is denoted by the suffix zero, ‘0’. Current year—The year for which comparison is sought or required. It is denoted by the suffix ‘I’. P0-Priceof a commodity in the base year. P1-Price of a commodity in the current year. Q0 -Quantity of a commodity in the base year. Q1 -Quantity of a commodity in the current year. W-Weight assigned to a commodity according to its relative importance in the group. P01 - Price index number for the current year. P10-Price Index number for the base year. Q01-Quantity index number for fee current year. Q01-Quantity index number for the base year. The various methods of constructing index numbers are discussed below. 1 . Simple Aggregative Index: Under this method, the total of current year price for the various commodities or items is divided by the total of base year prices and the quotient is multiplied by 100. Symbolically P10 =^p1 X 100 ^Po Example. | With the data given below construct price index for 1988 taking 1986 as base. | |---| | Commodities | Price | | 1986 | 1988 | | A | 50 | 70 | | B | 40 | 65 | | C | 80 | 95 | | D | 110 | 130 | | E | 20 | 18 | | F | 15 | 14 | | G | 10 | 12 | We have, Sp0 = 325 and Sp1 = 404 Price index for 1988. Po = ^p- x 100 404 x 100 01 Zp0       325 = 124.31 Hence, there has been 24.31 percent increase in prices of commodities in 1988 as Compared to 1986 prices. Simple Average of Relatives Index : The average of relative price index is the ratio expressed in percentages of the price of a commodity in given period called current period to its price in another period called the base period. If P0 and P1 represent the price of commodity during the base period and the given period (current period) respectively then by definition: P1 Price Relatives = P0 The relatives get by their average either by the arithmetic mean method or geometric mean method By the arithmetic, mean method, the average of price relatives index is: S x 100 ; P = _kP 01        N Where N refers to the number of items (commodities) whose price relatives are thus averaged. When geometric mean, is used for averaging the price relatives, the formula for obtaining the Index become log P01 P S log -P x 100 L Pq . N S logP . . A or where P =    × 100 N           P0 | or P01 = antilog | [ | 1 S log P x 100 <     P0      7 | S log P = antilog N | |---|---|---|---| | N | | | L | J | | Example : From the data given below, construct index for 1988 taking 1987 as base year, by the average or relatives method using (a) arithmetic, mean and (b) geometric mean methods. Price Index for 1988 P01 Sp1 / P0 x 100 N 726.07 6 121.01 | Commodities | Price 1987 | Price 1988 | |---|---|---| | A | 70 | 90 | | B | 60 | 95 | | C | 123 | 135 | | D | 150 | 180 | | E | 20 | 20 | | F | 15 | 16 | Solutions. (a) Construction of index number by arithmetic mean method. | Commodities | Price 1987 (P0) | Price 1988 (P1) | Price relatives (P1/P0 × 100) | |---|---|---|---| | A | 70 | 90 | 128.57 | | B | 60 | 95 | 158.33 | | C | 120 | 135 | 112.50 | | D | 150 | 180 | 120.00 | | E | 20 | 20 | 100.00 | | F | 15 | 16 | 106.67 | SP1/P0 x 100 = 726.07 (b) Construction of index number by geometric mean method. | Commodities | Base year Price (P0) | Current year Price (P1) | Price relative (p1/p0 × 100) | Log P | |---|---|---|---|---| | A | 70 | 90 | 128.57 | 2.1091 | | B | 60 | 95 | 158.33 | 2.1996 | | C | 420 | 135 | 112.50 | 2.0511 | | D | 150 | 180 | 120.00 | 2.0792 | | E | 20 | 20 | 100.00 | 2.0000 | | F | 15 | 16 | 106.67 | 2.0280 | S log p = 12.4670 P = Antilog     p0 x I00 01              N or    P01 = Antilog Slog pN = 12.4670/6 = antilog 2.0778 = 119.62 17.6.2 Weighted Index Numbers The index numbers constructed by the methods of simple aggregative and simple average of relatives are unweighted in the true sense of term. An equal importance assigned to all items is included in the index computed by the above two methods. In other words, implicit weight is given, in the sense each price is assumed to be of equal importance. Implicit weighting is far from realistic in many cases; Construction of indices which are said to be useful, requires a conscious and obvious effort to assign to each commodity or item a weight is accordance with the importance in the total phenomena so as to describe the activity. Index numbers with weight are of two types, namely weighted aggregative index and weighted average of relative index. The index numbers calculated by the weighted aggregative method are of the simple aggregative type with the fundamental difference that weights are assigned to the items included in the index. In other words, the base period quantities or current period quantities of commodities involved in the computation are taken as weights, there are different methods of assigning weights; and as such a large number of formulae for constructing index numbers have been devised, of which some important or popular ones are : 1.    Laspeyres method 2.    Paasche method 3.    Darbish and Bowley’s method 4.    Fisher’s Ideal method 5.    Marshal Edgeworth method 6.    Kelly’s method. 1.    Laspsype’s Method : In this method weights are determined by quantities in the base period. This method is widely used because it keeps the quantities consumed in the base period as constant in the current period also, and finds the change in aggregate value. The main limitation of this method is that, it works under assumption that there is no change in quantities of the current year. The formula for constructing’ the index is p01  ±pq: x 100 ZpoQo Where P01 = price index, p0 = price in the base year, p1 = prices in the current year, q0 = quantity in the base year and q1 quantity in the current year. 2.    Paasche Method : In this method the current year quantities are taken as weigths. The formula for constructing the index is : P01 =^piqL x 100 ^Po^i i .e. multiply current year prices of various commodities. with current year weights and obtain Spq^ also multiply the base year prices of various commodities with current year weights and obtain Spoqr Divide Spiqi by Spq and multiply the quotient by 100. 3.    Darbish and Bowley’s Method : Darbish and Bowley have suggested a practical method of constructing index using weighted aggregates by combining the Lspeyre’s and Paasche’s methods. In other words, under this method, the index is the simple arithmetic mean of the Laspeyre’s and Paasche method. This method takes into account the influence of both the periods i.e., current as well as base period. The formula for constructing this index is : Pol = L + P 2 where L = Laspeyris Index and P = Paasche Index Zpq + ¿pq or P = ^poqo---¿pq. x 100 01             2 4.    Fisher’s ideal method: Living Fisher has successfully combined the Laspeyre’s and Paasche’s methods to develop ideal index. According to him, an ideal i.e. true or unbiased index number must satisfy two basic tests, namely time reversal and factor reversal tests. The geometric mean of the Laspeyre’s and Paasche methods satisfies these tests. Hence Fisher consider his method of constructing index number as an ideal one. The Fisher’s Ideal Index is given by the formula: P01 Zpqo \ zpq zpq zpq x 100 or P01 = 4 L x P This method is assumed to be free from bias and takes into account both the constant and fluctuating weights. 5.    Marshall-Edgeworth Method ? In this method also both the current year as well as base year prices and quantities are considered. The formula for constructing the index is : P = E( Po + q1) P1 01  ^( Po + qj PO X 100 or p0i = ,,q +pp: x 100 ?poqo +^P0Q1 6.    Kelly’s Method: T.L. Kelly has suggested the following formula for constructing index number P01 =^pq x 100 ^poq Here weights are the quantities which may refer so some period, not necessarily the base year or current year. Thus’ the average quantity of two or more years may be used as weights. If in the Kelly’s formula, the average of the quantities of two years is used as weights, the formula becomes. P01 _ P1q x 100 where q = q0 + q1 ?pq              2 Similarly the average of the quantities of three or more years can be used as weights. This method is known as fixed weight aggregative index. An important advantage of this formula is that like Laspqnres or Paaschs index, it does not demand yearly change necessitating corresponding change in the weights. Weightede Average of Relatives: In the weighted aggregative methods discussed above price relatives were not computed. However, like unweighted, relative method it is possible to compute weighted average of relatives. For purposes of averaging we may use either to arithmetic mean or the geometric mean. In order to compute the. weighted arithmetic mean of the relatives. Express each item of the period for’ which the index number is being calculated as a percentage of the same item in the base period. Then multiply the percentage as obtained for each item by the weight which has been assigned to that item. Symbolically the index number is P01 = ΣPV ΣV where P = Price relative and V = a Value weights i.e. p0q0 Instead of using arithmetic mean, the goemetric mean may be used for averaging relatives. The weighted geometric mean of relatives is computed in the same manner as the unweighted, geometric mean of relative index number except mat weights are introduced by applying them to the logarithms of the relatives. When this method is used the formula for computing the index is: P01 = Antilog ΣVlog P ΣV P1 where P = Xp0 100 P0 and   V = Value weight i.e. P0q0 for each item. Example: From the following data compute price index by applying weighted average of price relatives method using: (a)    arithmetic mean, and (b)    geometric mean. | Commodities | P0 (Rs) | q0 (kg.) | P1 (Rs.) | |---|---|---|---| | Sugar | 3.0 | 20 | 4.0 | | Flour | 1.5 | 40 | 1.6 | | Milk | 1.0 | | 101.5 | Solution: (a) Index Number Using Weighted Arithmetic Mean of price Relatives | Commodities | P0 | q0 | P1 | T = P0q0 | P = × 100 | PV | |---|---|---|---|---|---|---| | Sugar | Rs.3.0 | 20 kg | Rs. 4.0 | 60 | 4/3×100 | 8,000 | | Flour | Rs.1.5 | 50 kg | Rs. 1.6 | 60 | 1.6/1.5×100 | 6,400 | | Milk | Rs. 1.0 | 10kg | Rs.1.5 | 10 | 1.5/1.0×100 | 1,500 | | | | | SV = 130 | | SPV = 15,900 | | _ SPV _ 15,900 _ ΣV 130 This means that there has been a 22.3 per cent increase in prices over the base level. (b) Index Number Using Geometric Mean of Price Relatives | Commodities | P0 | q0 | P | V | P1 | Log P | V Log P | |---|---|---|---|---|---|---|---| | Sugar | Rs. 3.0 | 20kg | Rs.4.0 | 60 | 133.3 | 2.1249 | 127.494 | | Flour | Rs/ 1.5 | 40kg | Rs. 1.6 | 60 | 106.7 | 2.0282 | 121.692 | | Milk | Rs. 10 | 10kg | Rs. 1.5 | 10 | 150.0 | 2.1761 | 21.761 | SV = 130 SV Log P = 270.947 Poi = ΣV.log P Xv = Antilog 270.947 130 = Antilog 2.084 =121.3 Self-Chek Exercise-4 Q1. What are index numbers. How are they constructed? Q2. Distinguish between weighted and unweighted index numbers. Q3. Write short notes on a)    Laspeyres method b)   Paasche method c)   Fisher’s ideal method Q4. What are the differences between the Laspeyre’s and Paasche’s system of weights in compiling a price index? Calculate both Laspeyre’s and Paasche’s aggregative price indices for the year 2000 from following data: | Commodities | Quantity | Price Per Unit (Rs.) | |---|---|---| | 1999 | 2000 | 1999 | 2000 | | A | 3 | 5 | 20 | 25 | | B | 4 | 6 | 25 | 30 | | C | 2 | 3 | 30 | 25 | | D | 1 | 2 | 10 | 7.50 | 17.7    Summary In this unit, we discussed the fundamental concept of index numbers, their importance, and their applications in economic analysis. We explored the problems involved in constructing index numbers, ensuring we understand the common pitfalls and challenges. The unit also covered the classification of index numbers into price, quantity, and value index numbers, each serving a unique purpose in economic measurement. Additionally, we examined the methods of constructing index numbers, focusing on both unweighted and weighted indices, to provide a comprehensive understanding of how these tools are developed and utilized. 17.8    Glossary •    Index Numbers: Statistical measures used to express changes in a variable or a group of related variables over time, making it easier to compare different data points. •    Price Index Numbers: Index numbers that measure the relative change in the prices of goods and services over time, often used to track inflation. •  Quantity Index Numbers: Index numbers that measure the change in the quantity of goods produced, consumed, or sold over time, helping in understanding production and consumption patterns. •    Value Index Numbers: Index numbers that measure the change in the total value (price multiplied by quantity) of goods and services over time, combining both price and quantity changes. •    Unweighted Indices: Simple index numbers that do not take into account the relative importance or weight of the different items included in the index. •    Weighted Indices: Index numbers that assign different weights to items based on their relative importance, providing a more accurate measure of changes in the variables being studied. 17.9    Answers to Self-Check Exercise Self-Check Exercsie-1 Answer to Q1. Refer to Section 17.3. Self-Check Exercsie-2 Answer to Q1. Refer to Section 17.4 Self-Check Exercsie-3 Answer to Q1. Refer to Section 17.5. Self-Check Exercsie-4 Answer to Q1. Refer to Sections 17.3 and 17.6. Answer to Q1. Refer to Section 17.6. Answer to Q1. Refer to Section 17.6.2. Answer to Q1: 109.78 ; 109.72 17.10    References/ Suggested Readings 1.    Nagar, A. L., & Das, R. K. (1997). Basic statistics. Oxford University Press. 2.    Grewal, P. S. (1990). Methods of statistical analysis. Sterling Publisher 3.    Gupta, S. C. (2014). Fundamental of applied statistics. Sultan Chand & Sons. 4.    Gupta, S. P. (2021). Statistical methods. Sultan Chand & Sons. 17.11    Terminal Questions Q1. What is an index number? Explain the various problems involved in the construction of index numbers. Q2. What are Index numbers? How they are constructed? Explain the role of weights in the construction of general price index numbers. Q3. Given the following information: | Group | Food | Clothing | Fuel and lighting | Rent | Miscellaneous | |---|---|---|---|---|---| | Index | | | | | | | Number | 221 | 198 | - | 161 | 183 | | Weight | 35 | 14 | 15 | 8 | 20 | If the cost-of-living index is 193, find the index number of fuel and lighting. ***** UNIT-18TESTS FOR CONSISTENCY OF INDEX NUMBER STRUCTURE 18.1    Introduction 18.2    Learning Objectives 18.3    Tests for Consistency of Index Numbers 18.3.1    Time Reversal Test 18.3.2    Factor Reversal Test Self-Check Exercise-1 18.4    Fixed Base and Chain Base Index Numbers 18.4.1    Steps in Constructing a Chain Base Index Number 18.4.2    Conversion of Fixed Base and Chain Base Index Number 18.4.3    Merits of Chain Based Method 18.4.4    Limitations of Chain Index Self-Check Exercise-2 18.5    Summary 18.6    Glossary 18.7    Answers to Self-Check Exercise 18.8    References/Suggested Readings 18.9    Terminal Questions 18.1    Introduction In this unit, we will discuss the methods used to test the consistency of index numbers. Consistency tests are essential to ensure that index numbers accurately reflect the changes in economic variables over time. We will explore two primary tests: the time reversal test and the factor reversal test. Additionally, we will examine the differences between fixed base and chain base index numbers, the steps involved in constructing chain base index numbers, and the merits and limitations of the chain base method. 18.2    Learning Objectives After going through this unit, you will be able to •    Understand and apply the time reversal and factor reversal tests to evaluate the consistency of index numbers. •    Differentiate between fixed base and chain base index numbers, and comprehend their respective construction and conversion processes. •    Analyze the merits and limitations of the chain base method in the construction of index numbers. 18.3    Tests for Consistency of Index Numbers A large number of formulas leave been by different statistician for constructing index numbers and the problem is that of selecting the most appropriate in a given situation. Prof. Irving Fisher, a famous statistician, bad laid down two tests for a good index number. The tests are: 18.3.1    Time Reversal Test:— According to this test, an index number should show, the same relative movement from one period to another whichever may be taken as base. In other words, an index number should work, both ways as well as backward. An index number for current year on the basis of base year should be reciprocal of the index number for base year on the basis of current year. In the words of Fisher, “The test is that the formula for calculating an index number should be such that it will give the same ratio between one point of comparison and the other, no matter which of the two is taken as base.” When the time periods are reversed, one index number is the reciprocal of the other index and their product is always equal to one. Let p01 be an index number for die current year (1); based on base year (0) and P10 be an index number for base year (0) and based on current year (1) then : P01 – P10 = 1 If the product of two indices is not equal to unity men there is a bias in the formula being used. Below we discuss which formula satisfies time reversal test: (a)    Laspeyres Formula: Zpq0 P 01 = „ ZpoQo Changing 1 to 0 arid 0 to 1, we get zpq Pio -pqi ^PoQ! V * 1. Wi -^o x Po1 P1o   -pqo The product of two indices is not equal to unity, therefore, this formula does not satisfy time reversal test: (b)    Paasche Formula: P _ Zpqi P 01 = v Zpoqi Changing 1 to 0 and 0 to 1, we get : ^o P10 Zpqo -pq^ -p^ Po1 X P1o = -poQi SpiQo * 1- Like the Laspeyres formula, this formula also does not satisfy time reversal test (c) Fisher’s Formula : P10 = -pq0 y vpq x \ Zpoqo Zpq Changing 0 to 1 and 1 to 0, we get: P10 = Zpoqo \ ^1 X ^PoQo ^pq ^pq^ 2p1qL ^pq^ ^pq p n i X p    XXX 01       10 V ^pq ^pq  Ypq Vpq = 1. Thus this formula satisfies the time reversal test and is called the ideal formula. There are five methods which do satisfy the test : (1)    The Fisher’s Meal Formula (2)    Simple Geometric Mean of price relatives (3)    Aggregates with fixed weights. (4)    The weighted geometric mean of price relatives if we used fixed weights. (5)    Marshall-Edgeworth method. 2.    Factor Reversal Test : According to this test, the product of the price index and quantity index should be equal to the corresponding value index. In symbols: P01 x Q ^Pq ?pq In the words of Fisher. “Just as our formula should permit the interchange of the two times without giving inconsistent results, so it ought to permit interchanging the prices and quantities without giving inconsistent results i.e. the two results ‘multiplied together should give the true value ratio.” Put it in other words, the change in price multiplied by the change in quantity should be equal to the total change in value. The total value of a given commodity in a given year is the product of the quantity and the price per unit (value = p × q). If p1 and p0 represent prices and q1 and q0 the quantities in the current year and the base year respectively; and if P01 represents the change in price and Q01, represents the change in the current year then P01 X Q01 = ^pq ZpoQo In simple, words the tests is satisfied if the product of the price index and the quantity index computed from same data is equal to the ratio of the aggregate value in the current year to aggregate value in the base year. In case the product is not equal to the value index, there is a bias in the formula being used. Below we examine which formula satisfies this test : (a) Laspayres Formula : Zpq0 P = 01 ^pq ?Pq and Q = Q01 ^pqp P01 X Q Zpq0 Zpoqo Spoql ^ Zpq Zpoqo   %poqo Thus this formula does not satisfy the factor reversal test : (b) Paasche Formula : Zpq P = 01 ^poqi zpq and Q = Q01  Xpqo W1 ^P^  ^P^ Poi x Q01 —       x       ^ ZPoqi   Ppq.,   Zp0q0 Thus this formula also does not satisfy factor reversal test. (c) Fisher’s Formula: P 01 zpq0., W1 X Zp0 q0 ^pq and   Q 01 — ^poq1 v W1 X \ ^poqo ±/¥P (by changing p to q and q top) No’v P01 × Q01 | ^ppp X | ±/¥/1 X | ^POL x | ^Pq | |---|---|---|---| | Vpoqo | Xpq | Vpoqo | Zpqo | 2 jM lW)) Zpq = Zpq Thus Fisher’s formula satisfies factor reversal test. Example: The following figures relate to the prices and quantities of certain commodities. Construct an appropriate index, number and show if it satisfies the time reversal test. | | 1973 | 1974 | |---|---|---| | Commodities | Price | Quantities | Price | Quantities | | A | 30 | 50 | 32 | 50 | | B | 25 | 40 | 30 | 35 | | C | 18 | 50 | 16 | 55 | Index Number by Fisher’s Ideal Method 1973         1974 Siq> xSæ x loo \ zpq ^pq | Commodities | p0 | q0 | p1 | q1 | p1q0 | p0q0 | p1q1 | p0q1 | |---|---|---|---|---|---|---|---|---| | A | 30 | 50 | 32 | 50 | 1600 | 1500 | 1600 | 1500 | | B | 25 | 40 | 30 | 35 | 1200 | 1000 | 1050 | 875 | | C | 18 | 50 | 16 | 55 | 800 | 900 | 880 | 990 | | | | | | | W0 =3600 | W0 =3400 | Zp^l =3530 | -pq =3365 | P 01 — 3600 3530 -----x----- X 100 3400 3365 = 1.111 ×100 = 1.054 × 100 = 105.4 Time reversal test is satisfied when P01 × P10 = 1 Substituting the values of Sp1q0, 'Epoqo etc; 3600 3530 -----x----- 01    3400 3365 P10 Ypoq  Lpq _ 3365 3400 x       — x V Wi Wo 3530 3600 3600 3530 3365 3400 P01 x Pio — -----x-----x-----x----- 3400 3365 3530 3600 Hence time reversal test is satisfice the proof of the formula. Test the adequacy of index by the reversal factor and factor reversal for the following data: | Items                   1983                            1988 Price (Rs.)      Quantity (kg.) Price (Rs.)        Quantity (kg.) Rice            3                 5              4                 6 Oil              24                1                19                 1 Tea              6                 1               7                  1 Washing Powder        4                              5                1 Sugar           3                4              5                 5 Milk            2                2              3                 3 Solution: First, compute the required values. Items                  1983          1988 | |---| | | p0 | q0 | p1 | q1 | p1q0 | p1q1 | p0q0 | p0q1 | | Rice | 3 | 5 | 4 | 6 | 20.0 | 24.0 | 15.0 | 18.0 | | Oil | 24 | 1 | 19 | 1 | 19.0 | 19.0 | 24.0 | 24.0 | | Tea | 6 | 1 | 7 | 1 | 7.0 | 7.0 | 6.0 | 6.0 | | Washing Powder | 4 | 1 2 | 5 | 1 | 2.5 | 5.0 | 2.0 | 4.0 | | Sugar | 3 | 4 | 5 | 5 | 20.0 | 25.0 | 12.0 | 15.0 | | Milk | 2 | 2 | 3 | 3 | 6.0 | 9.0 | 4.0 | 6.0 | Time reversal test is P01 × Pl0= 1. According to Fisher formula, P01 X P10 = ^pq0 y Wi y ^p0qi y ^poqo X       X       X V Zpoqo ^pq Zpq Zpiq0 63.0 Laspeyres Formula, Splqo Poi x Pio = ^poqo 74.5 89.0 73.0 63.0 ----X ----X ----X ---- 73.0 89.0 74.5 X SPo^G W1 74.5 ----X 63.0 63.0 ----/ 1 89.0 Paasche Formula, | | Zpq v Zpoqo | |---|---| | P01 × P10 = = | X spq ^pq^ 89.0 63.0 :       X       ^ i | 73.0 74.5 Marshall Edgeworth, | P01 × P10 = | sm y Zpqi zpq spo^o -        X        X        X spoqo spq spqi spqi | |---|---| | = | 74.5 + 89.o   73.o + 63.o , :                X               = i 63.o + 73.o 89.o + 74.5 | |---|---| Simple Aggregative Method, | P01 × P10 = | Zpi vSo X Zp   Zpi | |---|---| 43 42 = -- X -- = 1 42 43    . Factor Reversal Test: Zp^! Poi x Qoi = Zpqo spq Poi X Q01 = Zpq ,, ¿qPo ,, ZqiPi ----- X ----- X ----- Zpq Zpq  2qo Po  sqüPi | 74.5 ----X 63.0 | 89.0 X 73.0 | 73.0 X 63.0 | 89.0 74.5 | |---|---|---|---| | (89)2 (63)2 | 89.0 = 63.0 | = i.4i | | Self-Check Exercise-1 Q1. What do you mean by tests of consistency for an index number? Q2. Explain the time reversal test and factor reversal test. Q3. For the following data prove that the Fisher’s Ideal Index satisfies both the Time Reversal Test and Factor Reversal Test and calculate its value. | Commodity | Base Year | Current Year | |---|---|---| | Price | Quantity | Price | Quantity | | A | 6 | 50 | 10 | 56 | | B | 2 | 100 | 2 | 120 | | C | 4 | 60 | 6 | 60 | | D | 10 | 30 | 12 | 24 | 18.4 Fixed Base and Chain Base Index Numbers Index numbers may be constructed by (1)    Fixed Base or (2)    Chain Base 1.    Fixed Base: In fixed baas method a definition or a parted of years is taken as base and prices of subsequent years are compared directly or independently with the price in base year. The same base year should be a normal year free from any abnormalities. It should also be not far in the past. The following formula used for computing fixed base index number. Index Number for a particular year Price of the Current × 100 Price of the Previous year Example : Construct index numbers for 8 years taking 1961 as base from the following date. | Year | Price (Rs.) | Year | Price (Rs.) | |---|---|---|---| | 1961 | 65 | 1965 | 86 | | 1962 | 70 | 1966 | 90 | | 1963 | 74 | 1967 | 95 | | 1964 | 80 | 1968 | 98 | Solution: Construction of Index Number Taking 1961 as base. | Year | Price (Rs.) | Index Number 1961= 100 | |---|---|---| | 1961 | 65 | 100 | | 1962 | 70 | 70 × 100 = 107.69 65 | | 1963 | 74 | 74 × 100 = 113.84 65 | | 1964 | 80 | 80 × 100 = 123.07 65 | | 1965 | 86 | × 100 = 132.30 | | 1966 | 90 | × 100 = 138.46 | | 1967 | 95 | × 100 = 138.46 | | 1968 | 98 | × 100 = 150.76 | This method, though convenient, has certain limitation. As time elapse conditions were once important become less significant and it becomes more difficult to compare accurate present conditions with those of a remote part. New items may have to be included and may have to be deleted in order to make the index more representative. In significant may be desirable to use the chain base number when this method is used to compare is made with s fixed base; rather than comparisons from year to year. 18.4.1    Step in constructing a Chain Base index In chain base method, there is no fixed base to compare the price of subsequent year, but price of each is compared with the price of the preceding year. In this way, we compute price relatives, there price relatives are called link relatives. These link relatives are chained together to get the Chain indeed. The following formulae may be used for computing link relatives and chain index: Link Relative   = Price of the Current year × 100 Price of the Previous year To get the chain indices, these link relatives are to be chained by the following formula: Chain Index Like relative for Current Year × Chain Indices of previous year 100 (1)    Find the link relative for each year by employing the above formula. (2)    Chain the link relatives to get the Chain Indices by me above formula. Example: Construct Chain base index numbers for the following data: Solution: | Year | Price (Rs.) | Year | Price (Rs.) | |---|---|---|---| | 1961 | 65 | 1965 | 86 | | 1962 | 70 | 1966 | 90 | | 1963 | 74 | 1967 | 95 | | 1964 | 80 | 1968 | 98 | Construction of Chain Index Numbers | Year | Price | Link Relatives | Chain Index | |---|---|---|---| | 1961 | 65 | 100 | 100 | | 1962 | 70 | × 100 = 107.7 | = 107.7 | | 1963 | 74 | × 100 = 105.7 | = 113.8 | | 1964 | 80 | × 100 = 108.1 | = 123.0 | | 1965 | 86 | × 100 = 107.5 | = 132.2 | | 1966 | 90 | × 100 = 104.6 | = 138.3 | |---|---|---|---| | 1967 | 95 | × 100 = 105.5 | = 145.9 | | 1968 | 98 | × 100 = 103.2 | = 150.6 | 18.4.2    Conversion for Fixed Base and Chain Base Index Numbers Index numbers computed with fixed base method can be converted into chain base index numbers and index numbers computed by chain base method can be changed into fixed base index numbers. In fixed base method year to year comparison is not possible because index numbers are tied to a distant past. In many situations, year to year comparison is most essential to study the behavior of the variables. To accomplish this purpose fixed base index numbers are converted into chain base numbers by the following rule Fixed base index for 1936 Chain index for 1936 =                             × 100 Fixed base index for 1935 and Fixed base index for 1937 Chain index for 1937 =                             × 100 Fixed base index for 1936 and so on. In chain base index numbers year to year comparison is possible but in many situations comparison with remote past becomes, necessary study in the growth pattern of the data. To achieve and objectives chain base index numbers are converted into fixed base index numbers by the following rule. Fixed index for 1950 × Chain index for 1951 Fixed Base Index for 1951 = 100 and Fixed Base Index For 1951× Chain Base Index for 1952 Fixed Base index for 1952 = 100 and so on. Example: From the following fixed base index number to find chain base index numbers. | Year : | 1965 | 1966 | 1967 | 1968 | 1969 | |---|---|---|---|---|---| | Fixed Base Index : | 425 | 446 | 457 | 480 | 496 | Solution: Conversion of Fixed Base Index Nos. to Chain Base Index Numbers. | Year | Fixed base Index Nos. | Fixed base index Nos. converted to Chain indices | Chain base index numbers | |---|---|---|---| | 1965 | 425 | - | | 100.00 | | 1966 | 446 | | × 100 | 107.94 | | 1967 | 457 | | × 100 | 102.46 | | 1968 | 480 | | × 100 | 105.03 | | 1969 | 496 | | × 100 | 103.33 | Example: Calculate fixed base index numbers of the following series of chain base index: Year           :   1950   1951    1952    1953    1954   1955   1956 Chain Indices   :   100    115     120     93      102     156    83 Solution: Conversion of Chain Base Index Numbers to Fixed Base Index Numbers. | Year | Chain Base Index Nos. | Chains Indices converted to Fixed base = 1950 | Fixed Base Index Nos. | |---|---|---|---| | 1950 | 100 | - | 100.00 | | 1951 | 115 | | 115.00 | | 1952 | 120 | | 138.00 | | 1953 | 93 | | 128.34 | | 1954 | 102 | | 130.90 | | 1955 | 156 | | 204.2 | | 1956 | 83 | | 169.48 | 18.4.3    Merits of the Chain Base Method: 1.    The chain base method has a great significance in practice because in economic and business data we are more often concerned with making comparisons with the previous period, and not with any distant past. The link relatives obtained by chain base method serve this purpose. 2.    Chain base method permits the introduction of new commodities and the deletion of old ones without necessitating either the recalculation of entire series or other drastic changes. Because of this flexibility, Chain index is used in many types of indices such as the consumer price index and the wholesale price index. 3.    Weights can be adjusted as frequently as possible. This flexibility is of great significance in many types of index numbers. 4.    Index numbers calculated by the chain base method are free to a greater extent from seasonal variations than those obtained by the other method. 18.4.4    Limitations of the Chain Index: The limitation of the chain index is that the whole percentages of previous year figures give accurate comparisons of year to year changes, the long range comparisons of chained percentages are not strictly valid. However, when the index number user wishes to make year to year comparisons, as is so often done, by the businessman, the percentages of the preceding year provide a flexible and useful tool. Self-Check Exercise-2 Q1. Distinguish between Chain Base Index Numbers and Fixed Base Index Numbers. Q2. How Chain Base Index Numbers are constructed? What are the Merits of Chain Base Index Numbers? Q3. From the following annual average prices of three commodities given in rupee per unit, find Chain Index numbers based on 1997: | Commodities | 1997 | 1998 | 1999 | 2000 | 2001 | |---|---|---|---|---|---| | X | 8 | 10 | 12 | 15 | 12 | | Y | 10 | 12 | 15 | 18 | 20 | | Z | 6 | 9 | 12 | 15 | 18 | 18.5    Summary In this unit, we discussed various tests for the consistency of index numbers, focusing on the time reversal test and the factor reversal test. These tests help verify the reliability and accuracy of index numbers in reflecting economic changes. We also examined fixed base and chain base index numbers, understanding their construction, conversion, and applications. The steps involved in constructing a chain base index number were detailed, highlighting the advantages and limitations of this method. 18.6    Glossary •   Consistency Tests: Methods used to verify the accuracy and reliability of index numbers in representing economic changes over time. •   Time Reversal Test: A test that checks whether an index number remains consistent if the time periods are reversed. •   Factor Reversal Test: A test that verifies whether the product of a price index and a quantity index equals the value index. •    Fixed Base Index Numbers: Index numbers that use a constant base year for comparison across different time periods. •    Chain Base Index Numbers: Index numbers that use the previous period as the base for each subsequent period, allowing for a continuous comparison over time. 18.7    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 18.3. Answer to Q2. Refer to Sections 18.3.1 and 18.3.2. Answer to Q3. Fisher Index Satisfies the Factor Reversal and Time Reversal Tests Self-Check Exercise-2 Answer to Q1. Refer to Section 18.4. Answer to Q2. Refer to Section 18.4.1 and 18.4.3. Answer to Q3: 100, 131.67, 166.05, 204.79, 212.36 18.8    References/Suggested Readings 1.    Nagar, A. L., & Das, R. K. (1997). Basic statistics. Oxford University Press. 2.    Grewal, P. S. (1990). Methods of statistical analysis. Sterling Publisher 3.    Gupta, S. C. (2014). Fundamental of applied statistics. Sultan Chand & Sons. 4.    Gupta, S. P. (2021). Statistical methods. Sultan Chand & Sons. 18.9    Terminal Questions Q1. Explain the Timer Reversal and Factor Reversal Tests. Examine whether Laspeyre’s and Paasche’s Index numbers satisfy these tests. Q2. Using the following information, calculate Paasche, Laspeyres and Fisher Index. Table: Price and Quantity Produced of different goods (Hypothetical Data) | Year | Prices | Quantity | |---|---|---| | Food | Cloth | Food | Cloth | | 2010 | 2 | 2 | 1 | 2 | | 2011 | 3 | 2 | 2 | 5 | | 2012 | 4 | 3 | 3 | 6 | Q3. Explain the difference between ‘Fixed’ and ‘Chain’ base index numbers. Write the formula to convert the Chain Base Index to Fixed Base Index. ***** UNIT-19ANALYSIS OF TIME SERIES STRUCTURE 19.1    Introduction 19.2    Learning Objectives 19.3    Time Series Analysis Self-Check Exercise-1 19.4    Uses of Time Series Analysis Self-Check Exercise-2 19.5    Components of Time Series Self-Check Exercise-3 19.6    Methods of Measurement of Trends 19.6.1    Freehand or Graphical Method 19.6.2    Semi-average Method 19.6.3    Moving Average Method 19.6.4    Least Square Method 19.6.4.1    Fitting of Straight Line by the Method of Least Square Method 19.6.4.2    Merits and Limitations of the Method of Least Square Self-Check Exercise-4 19.7    Summary 19.8    Glossary 19.9    Answers to Self-Check Exercise 19.10    References/Suggested Readings 19.11    Terminal Questions 19.1    Introduction In this unit, we will discuss the analysis of time series, a crucial technique in understanding how data points collected or recorded at specific time intervals are related. Time series analysis helps in identifying patterns, trends, and seasonal variations in data over time. We will explore the uses of time series analysis, its components, and various methods for measuring trends, including the freehand method, semi-average method, moving average method, and least square method. 19.2    Learning Objectives After completion of this unit, you will be able to •   Understand the fundamental concepts and uses of time series analysis in various fields. •    Identify and analyze the key components of a time series, including trends and seasonal variations. •    Apply different methods for measuring trends in time series data, including the freehand method, semi-average method, moving average method, and least square method. 19.3    Time Series Analysis In business and economics, it is not only sufficient to study the past pattern of changes but also essential to measure, analyse and understand the forces which are operating in a firm or in an 278 industry or the economic system as a whole. There are many approaches for studying as well as forecasting the operating forces but one of these approaches is known as ‘Analysis of time series.’ A time series is a set of observations taken at specified times, usually at equal intervals. For example, when quantitative data regarding agricultural reduction national income, birth rate are arranged in order of their occurrence, the resulting statistical series is called time series. Mathematically, a time series is defined by the values X1, X2 X3, ........ Xn of the variable X at times t1,t2,t3,  tn. Here the variable X is a function of time. (t) i.e.     X = f(t) It means that the variable X depends upon time. The problem of time series analysis can best be appreciated with the help of. the following example. | Year | Sales of Firm A (thousand units) | Year (thousand units) | Sales of Firm a | |---|---|---|---| | 1970 | 40 | 1974 | 73 | | 1974 | 42 | 1975 | 48 | | 1972 | 47 | 1976 | 45 | | 1973 | 41 | 1977 | 44 | If we observe the above series we find that generally the sales have increased but for two years a” decline is also noticed. There may be several causes responsible for increase or decrease from one period to another such as changes in the testes and habits of people, of population availability of alternate products etc. It may be very difficult to study, the effect of various factors that have led either to an increase or decrease in the sales. The statistician, therefore, tries to analyse the effect of the various forces under four broad heads : (1)    Changes that have occurred as a result of general tendency of the data to increase or decrease, known as ‘secular movements’. (2)    Changes that have taken place during a period of 12 months as a result of change in climate, weather conditions, festivals etc. Such changes are called ‘seasonal variations’. (3)    Changes that have taken place as a result of booms and depressions. Such changes are classified under the head ‘cyclical variations.’ (4)    Changes that have taken place as a result of such forces that could not be projected like floods, earthquakes famines etc. Such changes are classified under the head irregular or erratic variations’. These are called components of time series and shall be discussed in detail. Self-Check Exercise-1 Q1. Explain the meaning of time series analysis. Q2. Define the time series mathematically. 19.4    Uses of Time Series Analysis : The time series analysis is of great significance to the economist and businessman and researcher etc. for the reasons given below: (i)    One can understand by observing data over a period of time the changes that have taken place in the past. Such analysis will be extremely useful in predicting the future behaviour. (ii)    Time series, analysis helps us in forecasting events, with which are can plan for the future. In other words, it helps in planning future operation. (iii)    Time series analysis facilitates comparison of different, points of time or different units performance and thereby drawing important conclusions. Self-Check Exercise-2 Q1. Explain the uses of time series analysis. 19.5    Components of Time Series : The combined force that caused fluctuations in the value of a phenomenon in time series may be broadly classified into four categories, commonly known as the components of time series. All or some of which are present in a given time series at varying degree, the components of time series, which often called elements of series are: (a)    Secular Trend (b)    Seasonal Variations (c)    Cyclical Variations (d)    Irregular Variations’. (a)    Secular Trend: The tendency of time series data to increase or decrease or stagnate during a long period. It is called the secular trend or simple trend. Thus, the trend is secular or long term due to the basic tendency to grow or decline over a period of time. Either increase or decrease would not be in the same direction throughout the given period, but different tendencies of increase, decrease or stability would be taken in different sections of time. Such tendencies are the result of the forces in an evolutionary manner and do not reflect sudden change. For example the growth or decline in economic time series is the interaction of forces like advances in production technology, large scale production, improved marketing management and , business organization—all of which are continuous but gradual process. The secular trend is broadly divided into two heads, namely, linear or straight line trend and non-linear trend. If the time series values plotted on graph appear more or less straight, then it is termed as linear trend, otherwise non-linear. (b)    Seasonal Variations : A number of forces which repeat periodically over a 12 month period give rise to seasonal variations. In other words, a variation which is of periodic in nature and whose repeating cycle is of relatively short duration, say weak, month or quarter. The amplitude of seasonal variations may vary, but their period is fixed being one year. The factor that cause seasonal variations are : (a)    natural forces ; and (b)    man-made conventions. The climatic change, plays an important role in seasonal movements. Changes in natural forces like weather, climate, rainfall, humidity, heat etc; act on different, products and industries differently. While the nature is primarily response for seasonal variations in time series, the people’s habits, fashions, customs, conventions also have their impact on seasonal variations. The study of seasonal variations is very useful to decision making in the sense, planning future operations regarding purchase, production, inventory control, personnel needs, selling and advertising programmes. In the absence of knowledge of seasonal variations, a seasonal upswing or seasonal slump may be mistaken as indicator of better business or deteriorating business conditions respectively. Thus, to understand the behaviour of the phenomenon in a time series properly, the time series data must be adjusted for seasonal variations. (c)    Cyclical Variations: In most of the economic and business time series data, there is periodic up and down movement in the sense that recurrent variation usually last longer than a year and these are more or less regular. These variations are known as cyclical variations. In other words the cyclical variations are long term variation that recur in, rises and declines, in activities and may or may not follow exactly since patterns after equal intervals of time. On the complete period which normally lasts from a years is termed as a ‘cycle’. These oscillation in any business activity are the outcome of the so called ‘Business Cycles’ which are four phase cycles consisting prosperity, decline, depression and improvement. (d)    Irregular Variations : Irregular variations also called erratic accidental or random refer to variations, which do not recur in a definite pattern. In other words, these fluctuations are purely random and are the result of such unforeseen as well as unpredictable forces which operate in erratic and irregular manner. The non-recurring factors like floods, famines, droughts, wars, earthquakes etc, cause the fluctuations in a most powerful manner. Self-Check Exercise-3 Q1. Explain the components of time series analysis with examples. Q2. Distinguish between trend, seasonal variation, cyclical fluctuations in a timer series. 19.6 Methods of Measurement of Trends: Trend component in time series can be studied and measured with the help of the following four methods, namely. (i)    Freehand or Graphic Method. (ii)    Semi-average Method, (iii)    Moving-average Method, (iv)    Least Square Method. This method is simple and flexible in studying trend. The procedure to obtain straight line trend is to plot the given time series on graph and draw a straight line, carefully on the plotted dots which will best fit to the data. The line drawn should be smooth. 19.6.1    Freehand or Graphical Method Example: Fit a trend line ‘to the following data by the graphic method: | Year | Production (in ‘000 tones’) | |---|---| | 1981 | 18 | | 1982 | 20 | | 1983 | 22 | | 1984 | 19 | | 1985 | 21 | | 1986 | 25 | | 1987 | 23 | | 1988 | 21 | Solution : Fitting a tread line by the graphic method Years 19.6.2    Semi-Average Method: This method has more objective method as compared with graphic method. In this method the whole time series is divided into several equals with reference to time. If we ‘are given data from 1966’ to 1985 i.e. 20 years, the two equal parts will be the first ten years i.e. from 1966 to 1975 and the second part from 1976 to 1985. In case of odd number of years, say 9, 11, 13, 15 etc., two equal parts will be made by ignoring the middle year. For example, if data are given for 11 years.fbr 1976 to 1986, the two equal parts would be from 1976 to 1980 and from 1982 to 1986, in this the middle year 1981 will be omitted. After dividing the given series into two parts, next calculate the arithmetic mean of time series for each part These arithmetic means are called ‘Semi-average’. These semi-averages are plotted as points against the middle points of the respective time period covered by each part. The line joining points gives the straight line trend fitting the given data : | Determine trend of the following data by the method of semi-averages. Year              Seles (000 Units) 1980                   20 1981                    24 1982                   22 1983                    30 1984                   28 1985                    30 1986                   34 1987                   36 Solution: | |---| | | Year                Seles          Semi-Average (‘000 Units) | | | 1980               20 1981                24 1982               22 1983               30 1984               28 | | | | | | | | | | | | | | | | 1985               30 1986               34 | | | | | | 197                 36 | | | | The first part semi-average 24 is to be plotted against die mid-years or the first part i.e. 1981 and 1982, and the second semi-average 32 is to be identified against the mid-years of the second part i.e. 1985 and 1986. The trend line is shown in diagram given below : mo'S---1----1----1----""---”----^--* 1980 81 82 83 84 85 86 87 88 Years 19.6.3    Moving Average Method: The trend is computed by smoothening fluctuations of the time series data by means of a moving average. The moving average is a series of successive, averages secured from a series of values by averaging ‘groups’ of ‘n’ successive values of the series. It is necessary to select a period for moving average such as 3.yearly moving average 5 yearly moving average, 8 yearly moving average, etc. Example: Calculate trend values taking a 3 yearly and 5 yearly period of moving average from the following data: | Year        : | 1970 | 1971 | 1972 | 1973 | 1974 | 1975 | 1976 | |---|---|---|---|---|---|---|---| | Production | | | | | | | | | ‘000 units : | 5 | 7 | | 912 | 10 | 11 | 8 | | Year        : | 1977 | 1978 | 1979 | 1980 | 1981 | 1982 | 1983   1984 | | Production | | | | | | | | | ‘000 units : | 12 | 13 | 17 | 20 | 10 | 15 | 12      14 | Solution: Computation of trend values by 3 yearly and 5 yearly moving average method. | Year | Production (‘000 units) | 3 Yearly Total | Moving Average | 5 Yearly Total | Moving Average | |---|---|---|---|---|---| | 1970 | 5 | - | - | - | - | | 1971 | 7 | 21 | 7.00 | - | - | | 1972 | 9 | 28 | 9.33 | 43 | 8.60 | | 1973 | 12 | 31 | 10.33 | 49 | 9.80 | | 1974 | 10 | 33 | 11.00 | 50 | 10.00 | | 1975 | 11 | 29 | 9.67 | 53 | 10.00 | | 1976 | 8 | 31 | 10.33 | 54 | 10.80 | | 1977 | 12 | 33 | 11.00 | 61 | 12.20 | | 1978 | 13 | 42 | 14.00 | 70 | 14.00 | | 1979 | 17 | 20 | 16.67 | 80 | 16.00 | | 1980 | 20 | 55 | 18.33 | 83 | 16.60 | | 1981 | 18 | 53 | 17.67 | 82 | 16.40 | | 1982 | 12 | 41 | 13.67 | - | 15.80 | | 1984 | 14 | - | - | - | - | Merits: (i)    The moving average method is simple as compared with least squares method. (ii)    The moving average if happens to coincide with the cyclical movements, such variation are automatically removed. Limitation: (i)    In moving average method, the trend values cannot be computed for all the years. (ii)    No predetermined or guiding principles are available to select the period of moving aver-age. One has to use his own judgement. Therefore, great care has to be exercised in selecting the period of moving average. 19.6.4 Least Squares Method: The method of least squares is the most widely used method of fitting a line of the best fit to a series of data and is the most popular method of calculating trends in time series. It is a mathematical device employed for measuring the trend line which represents the movements of data most satisfactorily. The method of least squares is given this name because its method of calculating gives its certain important mathematical properties which are not shared by other methods. These properties are: 1.    The sum of the vertical deviations of the actual values of (y) from the fitted straight line when deviations above the trend line are given positive signs and deviations below the trend, line are given negative signs, is equal to zero, Is symbols : 2.    The sum of the squared, deviations of actual values (y) from trend values (yc), is less than the sum of squared deviations from any other straight line. In symbols: = minimum. This property is not shared by other straight lines. When a line is fitted td meet the first property is automatically met It is because of second property that the name “Least Squares” is derived. The major weakness of this method is that it does not indicate what type” of trend should be fitted to the data. Generally this is to be decided by the investigator but once the decision has been made, the method of least squares supplies formula which may be used for estimating the time of the best fit It is mathematical in character, therefore, it requires the solving of a set of simultaneous equations, called normal equations, to determine the values of constants involved in the equations. 19.6.4.1 Fitting of Straight Line by the Method of Least Squares The simplest method of computing the trend values by the method of least squares is the straight line, The straight line series as the exact representation of the trend, in case the ‘series is increasing or decreasing by constant amounts. The equation of straight line may be written as yc = a + bx Where, yc = Computed value of the trend, x = Independent variable which represents time, a = Value of trend when x is zero (y intercept), b = Slope of the line Here ‘a’ and ‘b’ are constants, once their values, are determined they do not change. The value of ‘b’ represents the amount by which the trend increases or decreases for each unit of turn When ‘b’ is positive, the trend increases by constant amounts and when ‘b’ is negative, the trend decreases by constant amounts. The values of ‘a’ and ‘b’ are computed by solving a set of simultaneous equations commonly called normal equations. The number of normal equations depends upon the number of constants in the equation. In a straight line, there are two constants so we need two normal equations for computing the values of ‘a’ and ‘b’. Normal equations are : Σy = Na + b Σx                      …… (1) Σxy = aΣx + b Σx2                         ..….. (2) In above equations y represents original values in a time series, N stands for the number of years of x represents time. There are two methods of fitting straight line trend: 1.    Direct Method: When this method is adopted for computing the values of constants ‘a’ and ‘b’, the first year is always taken as origin and its value is taken equal to zero, the next year is given value 1 and third, year 2 and so on. After computing the desired values, normal equations are solved simultaneously to get the values of ‘a’ and ‘b’. Example: Fit a straight line trend to the following data: Year        :    1947    1948   1949   1950   1951   1952  1953   1954  l955 Production   :     11       13      15     14     15      16     16     17     18 (‘000 tons) Solution: Calculation of straight line trend by the method of least squares. | Year | Production Y | Origin, 1947 X | XY | X2 | |---|---|---|---|---| | 1947 | 11 | 0 | 0 | 0 | | 1948 | 13 | 1 | 13 | 1 | | 1949 | 15 | 2 | 30 | 4 | | 1950 | 14 | 3 | 42 | 9 | | 1951 | 15 | 4 | 60 | 16 | | 1952 | 16 | 5 | 80 | 25 | | 1953 | 16 | 6 | 96 | 36 | | 1954 | 17 | 7 | 119 | 49 | | 1955 | 19 | 9 | 144 | 64 | | N = 9 | ΣY = 315 | ΣX = 36 | ΣXY = 584 | ΣX2 = 204 | The equation of straight line is Y = a + bX Normal Equations are ΣY = Na + bΣx ΣXY = aΣX + bΣX2 Substituting the above computed values, we get 135 = 9a + 36b                     ….. (1) 584 = 36 a + 204 b                 ..….(2) Multiplying equation (1) by 4 and subtracting it from equation (2), we get: 584 = 36 a + 204 b 540 = 36 a + 144 b 44 = 60b or b = 0.73 Substituting the value of (b) in equation (1) 135 = 9a + 36 × 0.73 or    9a = 108.72 a = 12.8 The required equation is Y = 12.8 + 0.73 X 2.    The Short-cut Method In this method, origin is always taken in the middle of the time series in such a way that Σx becomes zero. The small letter ‘x’, representing time, has been substituted to distinguish it from capital letter ‘X’ which is employed in the direct method. Normal equations are : Σy = Na + b Σx Σxy = a Σx + bΣx2 By taking origin in the middle of time period covered, Σx becomes zero. Therefore, the above equations are reduced to the form : Σy = Na the second equation becomes. Σxy = bΣx2. or Hence, we can compute the values of ‘a’ and ‘b’ directly. If the series consists, of an odd number of years, the origin would be the middle of the series. On the contrary if the series consists of an even number of years, the origin falls in the centre of two middle years. (a)    Odd Number of Years: When we are supplied with the data for an odd number of years, the origin is always taken in middle of time period so that Σx becomes zero. The values of Parameters ‘a’ and ‘b’ calculated with the aid of formula : (b)    Even Number of Years: If the time series consists of an even number of years the middle of the series falls between two years. This middle would necessitate the use of decimals to measure the distance of any years from the origin. In case of even years also Σx will be zero if the origin is placed, mid way between the two middle years. For example, if the years are 1973, 1974, 1975, 1976, 1977 and 1978 we can take deviations from the middle year 1975.5. If the deviations would be -2.5, -1.5, -0.5, + 0.5, +1.5, -2.5 for the various years and the total Σx would be zero. Hence both in odd as well as in even number of yearswe, can use the simple procedure of determining the values of the constant a and b. Example: Fit a straight, line trend by the method of least squares to the following data. Assuming that the same rate of change, continues, what would be the predicted earnings for the year 1972. | Year            : | 1963 | 1964 | 1965 | 1966 | 1967 | 1968 | 1969 | 1970 | |---|---|---|---|---|---|---|---|---| | Earnings | | | | | | | | | | (Rs. In lakhs) : | 38 | 40 | 65 | 72 | 69 | 60 | 87 | 95 | | Solution: | | | | | | | | | Fitting of straight line trend by the method of Least Squares. | Year (Rs. In lakhs) By 2 Y | Earnings from 1966 X | Deviations multiplied XY | Deviations X2 | |---|---|---|---| | 1963 | 38 | -3.5 | -7 | -266 | 49 | | 1964 | 40 | -2.5 | -5 | -200 | 25 | | 1965 | 65 | -1.5 | -3 | -1.95 | 9 | | 1966 | 72 | -0.5 | -1 | -72 | 1 | | 1967 | 69 | +0.5 | +1 | +69 | 1 | | 1968 | 60 | +1.5 | +3 | +180 | 9 | | 1969 | 87 | +2.5 | +5 | +435 | 25 | | 1970 | 95 | +3.5 | +7 | +665 | 49 | | N = 8 | ΣY =526 | | ΣX = 0 | ΣXY = 616 | ΣX2 = 168 | Y0 = a + bX a = b = Y = 65.75 + 3.667X For 1972, X will be + 11 When X is + 11, Y will be Y = 65.75 + 3.667 (11) = 65.75 + 40.337 = 106.087 Thus the estimated earnings for the year 1972 are 106.087 lakhs. 19.6.4.2 Merits and Limitation of the Method of Least Squares Merits: 1.    This is a mathematical method of measuring trend and as such there is no possibilities of subjectivesness. 2.    The line obtained by this method is called the line of best fit because it is his line from which the sum of the positive and negative deviations is zero and the sum of the squares of the deviations is least, i.e. Σ(y-yc) = 0 and Σ(y-yc)2 is least. Limitations: Mathematical curves are useful to describe the general movement of a time series but it is doubtful whether any analytical significance should be attached to them, examine in special cases, it is seldom possible to justify on theoretical grounds any real dependence of a variable on the passage of time. Variables do change in a more or less systematic manner over time, but this can usually be attributed to the operation of other explanatory variables. Thus many economic time series show persistent upward trends over time due to a growth of population or to a general rise in prices, i.e., national income and the trends element can to a considerable extent be eliminated by expressing these series; per capita or in terms of constant purchasing power. For these reasons mathematical trends are generally best regarded as tools for describing movements in time series rather than as theories of the cases of such movements. Self-Check Exercise-4 Q1. What do you mean by term moving average? Q2. Explain the methods of measurement of trends. 19.7    Summary In this unit, we discussed the fundamental concepts of time series analysis and its importance in analyzing data over time. The unit covered the key components of a time series, including trends, seasonal variations, cyclical variations, and irregular variations. We explored multiple methods for measuring trends: the freehand or graphical method, the semi-average method, the moving average method, and the least square method. Specifically, we detailed how to fit a straight line using the least square method, along with its merits and limitations. 19.8    Glossary •    Time Series Analysis: A statistical technique that deals with time series data, or data that is observed over a period of time, to identify patterns, trends, and seasonal variations. •   Trend: The long-term movement or direction in a time series data, representing a sustained increase or decrease over time. •    Seasonal Variations: Regular fluctuations in time series data that occur within specific periods, such as months or quarters, due to seasonal factors. •    Moving Average Method: A technique used to smooth out short-term fluctuations and highlight longer-term trends by averaging data points over a specific number of periods. •   Least Square Method: A mathematical approach used to find the best-fitting line through a set of data points by minimizing the sum of the squares of the vertical deviations from each data point to the line. 19.9    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 19.3. Answer to Q2. Refer to Section 19.3. Self-Check Exercise-2 Answer to Q1. Refer to Section 19.4. Self-Check Exercise-3 Answer to Q1. Refer to Section 19.5 Answer to Q2. Refer to Section 19.5. Self-Check Exercise-4 Answer to Q1. Refer to Section 19.6.3. Answer to Q2. Refer to Section 19.6. 19.10    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Grewal, P. S. (1990). Methods of statistical analysis. Sterling Publisher. 3.    Gupta, S. P. (2014). Statistical methods. Sultan Chand & Sons. 4.    Nagar, A. L., & Das, R. K. (1997). Basic statistics. Oxford University Press. 19.11    Terminal Questions Q1. Explain the meaning of time series analysis. What are the various components of time series analysis? Q2. Discuss the methods of measurement of trend. Q3. Production of rice in a district during the last 10 years is given below: | Years | Production (in tons) | |---|---| | 1960 | 11,200 | | 1961 | 12,300 | | 1962 | 10,600 | | 1963 | 13,400 | | 1964 | 13,800 | | 1965 | 14,500 | | 1966 | 11,600 | | 1967 | 14,300 | | 1968 | 13,600 | | 1969 | 15,400 | Using 3-yearly moving average indicate the trend in the production of rice in the district. ***** LESSON-20MEASUREMENTS OF NON-LINEAR TRENDS STRUCTURE 20.1    Introduction 20.2    Learning Objectives 20.3    Non-Linear Trends 20.3.1    Fitting of Second-Degree Equation Self-Check Exercise-1 20.4    Logarithmic Straight-Line Trend or Exponential Curve 20.4.1    Methods of Fitting the Exponential Curve Self-Check Exercise-2 20.5    The Gompertz Curve Self-Check Exercise-3 20.6    Logistic Curve 20.6.1    The Method of Fitting of the Logistic Curve Self-Check Exercise-4 20.7    Summary 20.8    Glossary 20.9    Answers to Self-Check Exercise 20.10    References/Suggested Readings 20.11    Terminal Questions 20.1    Introduction In this unit, we will discuss the measurement of non-linear trends, which are essential for analyzing data that do not follow a straight-line pattern. Understanding non-linear trends allows us to better model and predict complex behaviours in various fields such as economics, biology, and environmental science. We will explore methods for fitting different types of non-linear equations, including second-degree equations, exponential curves, the Gompertz curve, and logistic curves. 20.2 Learning Objectives After going through this unit, you will be able to •    Understand the concept and significance of non-linear trends in data analysis. •    Learn methods for fitting second-degree equations and exponential curves to non-linear data. •   Apply the Gompertz and logistic curves to model complex growth patterns in various fields. 20.3    Non-Linear Trends The second degree equations alternatively called second degree parabola, is one of the most useful and simplest methods of fitting non-linear trends. The method of fitting second degree equations is only a little more, complicated than a straight line since it involves the addition of one more constant ‘C’. The equation of second degree parabola may be written as: yc = a + bx + cx2 Where, yc     stands for trend values a      stands for the y-intercept b stands for the slope at the origin c stands whether the curve is concave-upwards or downwards x stands for time It is named second degree equation because of the fact that two is the highest power to which ‘x’ appears in the equation. It is the square value ‘x’ which gives curvature to thetrend line whether the curve is concave upwards or downwards depends upon the value of ‘c’. (i)    If the value of ‘c’ is negative, the curve will have a downward bulge. (ii)    If the value of ‘c’ is positive, the curve will have an upward bulge. Fitting of Second Degree Equation The second degree equation like the straight line may be fitted by two methods namely: (a)    Direct Method In second degreeparabola there are three constants, so we need three normal equations for computing the values of constants. Normal Equations are : Σy = Na + bΣx + cΣx2 Σxy = aΣx + bΣx2 + cΣx3 Σx2y = aΣx2 + bΣx3 + cΣx4 To find the values of constants a, b and c, the above equations may be solved simultaneously. (b)    Short-cut By taking origin in the middle of the series “x becomes zero and the above equations are reduced to the form : Σy = Na + cΣx2                       …… (1) Σxy = aΣx2                             …  (2) Σx2y = aΣx2 + cΣx4                       …  (3) The value of ‘b’ is given by the equation (2) The values of ‘a’ and ‘b’ may be determined by the simultaneous solution of equations (1) and (3). These values may also be computed directly : Example: Fit a second degree parabola to the following data and estimate to value for 1986. Year             :     1979    1980     1981     01982    1983 Sale in (‘000 Rs.)  :      10        12        13        10         8 Solution: Computation of second degree parabolic trend equation. | Years (1) | Sales (‘000 Rs.) Y (2) | x∞ X - 1981 (3) | xy (4) | x2 (5) | x2y (6) | x2 (7) | x2 (8) | Trend value (9) | |---|---|---|---|---|---|---|---|---| | 1979 | 10 | -2 | -20 | -4 | 40 | -8 | 16 | 10.086 | | 1980 | 12 | -1 | -12 | 1 | 12 | -1 | 1 | 12.057 | | 1981 | 13 | 0 | 0 | 0 | 0 | 0 | 0 | 12.314 | | 1982 | 10 | +1 | +10 | 1 | 10 | +1 | 1 | 10.857 | | 1983 | 8 | +2 | +16 | 4 | 32 | +8 | 16 | 7.686 | | N= 5 | y=53 | x = 0 | xy = - 6 | Σx=10 | Σx2y=94 | Σx2 = | 0 Σx4 | =34 | The second degree parabolic trend elation is : y = 12.314- 0.6x - 0.857x2 Trend values shown in column (9) of She table are transits above equation in the following manner: For 1979 : x = -2, yc = 12.314 - 0.6 (-2) -0.857 (-2)2 = 10.086 For 1980 : x = -1, yc = 12.314-0.6(-1) -0.857(-1)2 = 12.057 and soon Trend value for 1986, x = 5, ye = 12.31-0.6 (5)-0.857(5)2 = 12.314-3.0-21.425 = -12.111. It is of the form : y = a + bx + cx2 + dx3 For computing the values of a, b, c, and d, we need four normal equations as shown below : Σy = Na + bΣx + cΣx2 + dΣx3            …… (1) Σxy = aΣx + bΣx2 + cΣx3 + dΣx4          …… (2) Σx2y = aΣx2 + bΣx3 + cΣx4 + dΣx4          …… (3) Σx2y = aΣx3 + bΣx4 + cΣx3 + dΣx4          …… (4) The value of a, b, c and d, may be determined by solving a set of four equations simultaneous. By making Σx = 0, the above equations may be reduced to the farm: Σy = Na + cΣx2                      …… (5) Σxy = bΣx2 + dΣx4                     …… (6) Σx2y = aΣx3 + cΣx4                      …… (7) Σx2y = dΣx4 + dΣx4                     …… (8) The values of ‘a’ and ‘c’ may be computed by solving equations (5) and (7) and the values of ‘b’ and ‘d’ of may be determined by solving equations (6) and (8). Self-Check Exercise-1 Q1. Define Non-linear trends. Q2. Explain the equation of second-degree parabola: yc = a + bx +cx2 Q3. Fit a parabola of the second-degree order to the following data: | Years | Sales (in million tonnes) | |---|---| | 1960 | 100 | | 1961 | 105 | | 1962 | 115 | | 1963 | 100 | | 1964 | 112 | | 1965 | 118 | 20.4 Logarithmic Straight Line Trend or Exponential Carve The straight line trend is fitted when the data is ‘increasing or decreasing by constant amounts from one period to another. Such a series show linear trend when plotted on a graph paper. If a series does not show a linear trend when plotted on a arithmetic graph paper but depicts a linear trend on a semi log paper then log arithmetic straight line trend is fitted. The equation of log straight line is of the form : y = ab2 c where : (a)    If the value of ‘b’ is a positive number greater than one, then the trend of the series is upward and the amount of change is undergoing ‘a constant percentage of increase. (b)    If the value of ‘b’ is a positive number smaller than one then trend is downward and the amount of change shows a constant percentage of decrease. These trends are also called exponential trends because x appears as an exponent in the equation. These trends are also named semi-logarithmic linear trends because we plot the values of y against absolute value of x. Method of Fitting the Exponential Curve: The exponential curve is of the form : y= ab2 In log form, fee equation may be written as : log y = log a + x log b Normal Equation are Σlog y = N log a + log bΣx               …… (1) Σx log y = log aΣx + log bΣx2            …… (2) By taking origin in the centre so that Σx becomes zero, the above equation may be reduced to the form: Σ log y = N log a Or           log and          Σx log y = log b Σx2 or            log By taking antilog of log a and log b, we can find the values of a and b. Steps: The following steps are involved in fitting an equation of the type y = abx (i)    Write the log values of y, the dependent variable (log.y). (ii)    Take origin in the middle of time period to make Σx = 0 and write down the deviations (x). (iii)    Multiply log y values with deviations obtained in step 2 to get x log y. (iv)    Apply the rule of addition to the products obtained in step (iii) to get Σx log y. (v)    Square the deviations to get Σx2. Example: Fit a logarithmic straight line trend to the following data showing the production (‘000 tons) of a sugar factory during 1979-85 | Years | :   1979 | 1980 | 1981 | 1982 | 1983 | 1984 | 1985 | |---|---|---|---|---|---|---|---| | Production | | | | | | | | | (‘000 tons) | : 12 | 10 | 14 | 18 | 20 | 24 | 30 | Solution: Fitting of logarithmic straight line trend by the method of Least Squares. | Year | Production (‘000 tons) y | log y | Origin, 1979 | Trend Values yc | |---|---|---|---|---| | x | x2 | x log y | | 1979 | 12 | 1.08 | -3 | 9 | -3.24 | 10.00 | | 1980 | 10 | 1.00 | -2 | 4 | -2.00 | 12.02 | | 1981 | 14 | 1.15 | -1 | 1 | -1.15 | 14.45 | | 1982 | 18 | 1.26 | 0 | 0 | 0 | 17.38 | | 1983 | 20 | 1.30 | 1 | 1 | +1.30 | 20.89 | | 1984 | 24 | 1.38 | 2 | 4 | +2.76 | 25.12 | | 1985 | 30 | 1.48 | 3 | 9 | +4.44 | 30.20 | | N= 7 | | Σlog y = 8.65 | Σx2 | = 28   Σ(x log y) | = 2.11 | | The logarithmic straight line is of the form y = abx or log y = leg a + x log b….... (i) Normal equations are Σlog y = N log a + log b (Σx) Σ(x log y = log a Σx + log b (Σx2) Since Σx = 0, the above equations are reduced to the form Log a and log b Thus the required equation is log y = 1.24+ 0.08 x Computation of Trend Values Year 1979, log y = 1.24 + 0.08 (-3) = 1.0 y = Antilog (1.0) = 10.00 Year 1980, log y = 1.24 + 0.08 (-2) = 1.08 y =Antilog (1.08) = 12.02 Year 1981, log y = 1.24 + 0.08 (-1) = 1.16 y = Antilog(1.16) =14.45 Year 1982, log y = 1.24 + 0.08 (0) × 1.24 y = Antilog (1.24) = 17.38 Year 1983, log y = 1.24 + 0.08 × 1 = 1.32 y = Antilog (1.32) = 20.89 Year 1984, log y = 1.24 + 0.08 × 2 = 1.40 y = Antilog (1.48) = 30.20 Example: Fit the curve y = aebx to the following data ‘e’ being 2.7183. x:0     2     4 y :      5.012         10            31.620 Solution: Fitting of the curve : The form of the equation is y = aebx In logarithmic form this equation can be written as log y = log a + (b log e)x. Putting y = log 10 y and B, = b log 10e The above equation is reduced to the form y = A + Bx Normal equations are Σy = NA + BΣx                     …… (1) Σxy = AΣx + BΣx2                    …… (2) The required calculation are shown in the following table | x | y | log y = y | xy | x2 | |---|---|---|---|---| | 0 | 5.012 | 0.7 | 0.0 | 0 | | 2 | 10.000 | 1.0 | 2.0 | 4 | | 4 | 31.620 | 1.5 | 6.0 | 16 | | Σx = 6 | - | Σy = 3.2 | Σxy = 8.0 | Σx3 = 20 | Substituting these values in equation (1) and (2) 3.2 = 3A + 6B ..... (3) 8.0 = 6 A + 20 B ..... (4) Multiplying equation (3) by 2 and subtracting from equation (4), we get 8.0 = 6 A + 20 B -6.4 = 6A + 12 B 1.6 = 8 B B = 1.6/8 = 0.2 Substituting the value of B in equation (3), we get 3.2 = 3A + 6 × 0.2 or 3A = 3.2-1.2 3A = 2.0 A = 0.67 Now A = log 10a 0.67 = log 10a Taking antilog, we get a = 4.677 and B = 6 log 10e or 0.2 = 0-4343 b or b = 0.46                    (log 10e = -4343) Hence the required equation is 0.46x y = 4.677e Self-Check Exercise-2 Q1. Discuss the steps involved in fitting the exponential curve. 20.5 The Gompertz Curve: The Gompertz curve may become to describe an increasing series winch is increasing by a decreasing percentage of growth or decreasing series which is decreasing by a decreasing percentage of growth. It is mainly employed in the study of economic and social trends because it portrays a process of cumulative expansion to a maximum value. It is defined-by the equation y = ka bx After taking log s the equation takes the shape of modified exponential form as shown below Log y = log k + b2 log a The shape of the Gompertz curve depends upon the values of k, log a and b. Actually its shapes depends upon whether 1.    log a is positive or negative 2.    The value of ‘b’ is greater or lesser than one. The Gompertz curve like the modified exponential curve also takes four shapes depending upon the values of k, log a and b. The Method of Fitting the Gompertz Curve Two methods may be employed for fitting this curve 1.    The Method of partial total 2.    The method of selected points. The procedure of fitting the Gompertz Curve by these methods is the same as decrease for fitting the modified exponential curve, examine a minor difference that we use log values of ‘y’ in place of given values of ‘y’ for computing the three constants, i.e. k, a and b. Example : Fit a Gorapertz curve to the following | Year (‘000 Rs.) | Savings | |---|---| | 1968 | 5.00 | | 1968 | 7.07 | | 1970 | 8.41 | | 1971 | 9.17 | | 1972 | 9.58 | | 1973 | 0.79 | Solution: Fitting of Gompertz Curve The Gompertz Curve is stated in the form y = ka bx in log form, it may be expressed as | log y = log k + bx log a | |---| | Year | x | Savings | log y | bx | log a(b) | log y = log k + log a (bx) | | 1968 | 0 | 5.00 | 06990 | 1.00 | -0.03 | 0.6992 | 5.00 | | 1969 | 1 | 7.07 | 0.8494 | 0.50 | -0.84 | 0.8492 | 7.06 | | Σ1log y | | 1.5484 | | | | 1.5484 | | | 1970 | 2 | 8.41 | 0.9248 | 0.25 | -0.075 | 0.9242 | 8.39 | | 1971 | 3 | 9.17 | 0.9624 | 0.125 | -0.0375 | 0.9617 | 9.15 | | Σ2log y | | 1.8872 | | | | 1.8859 | | | 1972 | 4 | 9.28 | 0.9814 | 0.0625 | -0.01875 | 0.9804 | 9.559 | | 1973 | 5 | 9.79 | 0.9908 | 0.0312 | -0.00937 | 0.9898 | 9.768 | | Σ3log y | | 1.9722 | | | | 19702 | | Calculation of Constants where,       Σ3 log y = 1.9722, Σ2 log y = 1.8872 Σ1 log y = 1.5484, n = 2 Substituting these values or b = 0.25 log a = (Σ2, log y - Σ1, log y) (1.8872 × 1.5484) and Thus the required equation is log y = 0.9992 + (0.5)x (.03) Trend values computed with the aid of the above equation are shown in the last column of the above table. Self-Check Exercise-3 Q1. Define Gompertz curve. Q2. Fit a Gompertz curve to the following data | Years | Saving (‘000 Rs.) | |---|---| | 1968 1969 | 5.00 7.07 | 19708.41 19719.17 19729.58 19730.79 20.6 The Logistic Curve : The logistic curve was introduced by Raymond Pearl and L.G. Reed, therefore, it is also called Pearl-Reed growth curve. It was employed by them for predicting the growth of population in the United States. The Logistic Carve The logistic curve is stated in the form 1/y = k + ab x The logistic curve is identical with the modified exponential except that ‘y’ in that equation is changed in to l/y in the case of logistic curve. The shape of the logistic curve resembles an elongated ‘S’ rising from a lower asymptote of zero to an upper asymptote indicated by k as has been depicted. 20.6.1    The Method of Fitting the Logistic Curve The method of fitting the logistic curve is the same which has been employed for fitting the modified exponential curve with a minor modifications. Here we employ reciprocals of y rather than original values of y. First, we find the reciprocals of original values and then the series is divided into three equal parts. The observations falling in each group are totaled. S1    stands for the total of observations in the first part. S2    stands for the total of observations in the second part. S3    stands for the total for observations in the thud part. The values of three constants a, b and k may be combated by the following expression: and The logistic curve gives a fairly good representation of the stages of slow initial growth acceleration and retardation in the life history of an industry. This curve may be used for finding the tread of economic series that decreases at a constant rate at later stages. The phenomenon for which this growth curve has been used is population growth, or the number of cells in an organism or the number of individuals in the region. Both the Gompertz curve and the logistic curve are identical because both the curves can be employed to describe increasing series which are increasing by a decreasing percentage of growth or decreasing series which are decreasing by decreasing percentage of decline. But both the curves differ in one respect, the Gompertz curve involves, a constant ratio of successive first differences of log y values whereas for the logistic curve it is the constant ratio of successive first differences of 1/y values. Self-Chek Exercise-4 Q1. Define the Logistic Curve. Q2. Explain the method of fitting the Logistic Curve. 20.7    Summary In this unit, we discussed the concept of non-linear trends and their importance in accurately analyzing data that exhibit curved patterns. We explored the process of fitting second-degree equations, which are useful for modeling quadratic trends in data. The unit also covered the logarithmic straightline trend, or exponential curve, explaining various methods for fitting these curves to data. We examined the Gompertz curve, a type of sigmoid function often used in modeling growth processes, and the logistic curve, which is commonly used in population studies and other fields. 20.8    Glossary •    Non-Linear Trends: Patterns in data that do not follow a straight line, indicating more complex relationships between variables. •    Second-Degree Equation: A polynomial equation of the form y = ax2 + bs + c used to model quadratic trends in data. •    Exponential Curve: A type of non-linear trend represented by the equation y = abx, often used to model rapid growth. •    Gompertz Curve: A sigmoid function used to model growth that starts rapidly and then slows over time. •    Logistic Curve: Another type of sigmoid function used to model growth that starts rapidly and then slows over time.A logistic curve is an S-shaped curve that is used to model population forecasting. It increases initially, rapidly then, and slowly later on and levels off or reaches a saturation value after a period of time. 20.9    Answers to Self-Check Exercise Self-Check Exercise-1 Answer to Q1. Refer to Section 20.3. Answer to Q2. Refer to Section 20.3. Answer to Q3: Y= 107.656 + 2.743 X + .232 X2 ; Origin: 1962.5 Self-Check Exercise-2 Answer to Q1. Refer to Section 20.4.1. Self-Check Exercise-3 Answer to Q1. Refer to Section 20.5. Answer to Q2. Refer to Section 20.5. Self-Check Exercise-4 Answer to Q1. Refer to Section 20.6. Answer to Q2. Refer to Section 20.6.1. 20.10    References/Suggested Readings 1.    Croxton, R. E., Cowden, D. J., & Klein, S. (1967). Applied general statistics. Prentice Hall. 2.    Grewal, P. S. (1990). Methods of statistical analysis. Sterling Publisher. 3.    Gupta, S. P. (2014). Statistical methods. Sultan Chand & Sons. 4.    Nagar, A. L., & Das, R. K. (1997). Basic statistics. Oxford University Press. 20.11    Terminal Questions Q1. Define a)    Gompertz curve. b)    Logistic Curve Q2. Fit an equation of the type y = a + bX + cX2 to the following data: | Years | Production (in ‘000 tons) | |---|---| | 1968 | 70 | | 1969 | 72 | | 1970 | 88 | | 1971 | 80 | | 1972 | 90 | ***** 303