--- title: "Vol 3 2" book: "SLM EME EDUCC 111" category: "General" publisher: "Ratan Prakashan Mandir Pvt. Ltd." type: "Educational Material" --- According to Latest Syllabus Read For Sure Success In University Examination RATAN TEXT BOOK EDUCATION MEASUREMENT AND EVALUATION Vol-3 M.A.Education (Sem-III) Dr. Nisha Tiwari Published by Ratan Prakashan Mandir Pvt. Ltd. 2nd Floor, Centre Plaza, Parinay Kunj, Lajpat Kunj Marg, Agra-282002 Copyright Authors & Publishers Revised Edition ISBN :978-93-0970-877-0 Price 95.00 only Printed at : KIDS INTERNATIONAL PVT. LTD. C-60, 61, 62, 63, EPIP, Shastripuram, Agra - 282007 Ph. : +91 9719004921 UNIT - 11 CHARACTERISTICS OF MEASUREMENT AND EVALUATION TOOLS-I Structure 11.1   Introduction 11.2   Learning Objectives 11.3  Basic Characteristics of Good Measuring Tools 11.3.1    Reliability Self-Check Exercise-1 11.3.2    Validity Self-Check Exercise-2 11.3.3    Objectivity, Usability and Practicability Self-Check Exercise-3 11.4    Norms for Interpretation of Test Scores 11.4.1    Developing and Using Test Norms to Compare Performance Self-Check Exercise-4 11.5    Summary 11.6    Glossary 11.7  Answers to Self-Check Exercise 11.8  References/Suggested Readings 11.9    Terminal Questions 11.1    INTRODUCTION Dear Learner, Evaluation is an important aspect of educational process. It is very important to evaluate achievement of the students. In the present lesson, we will learn about the basic characteristics of evaluation tools which includes; reliability, validity, objectivity, usability and practicability.' The lesson will also throw light on norms for interpreting test scores. 11.2    LEARNING OBJECTIVES After studying this unit, you will be able to: •     List down different characteristics of measurement and evaluation tools. Explain reliability and ways of its estimation. Discuss different types of validity. Define objectivity. Explain various procedures of developing and using test score norms. 11.3    BASIC CHARACTERISTICS OF GOOD MEASURING TOOLS Whenever a test or other measuring device is used as part of the data collection process, the validity and reliability of that test is important. Just as we would not use a math test to assess verbal skills, we would not want to use a measuring device for research that was not truly measuring what we purport it to measure. After all, we are relying on the results to show support or a lack of support for our theory and if the data collection methods are erroneous, the data we analyze will also be erroneous. Here we mention some qualities of a test and the techniques of its estimation. 11.3.1    Reliability Reliability is the consistency of your measurement, or the degree to which an instrument measures the same way each time it is used under the same condition with the same subjects. In short, it is the repeatability of your measurement. A measure is considered reliable if a person's score on the same test given twice is similar. It is important to remember that reliability is not measured, it is estimated. A good instrument will produce consistent scores. An instrument’s reliability is estimated using a correlation coefficient of one type or another. Reliability is synonymous with the consistency of a test, survey, observation, or other measuring device. Imagine stepping on your bathroom scale and weighing 140 pounds only to find that your weight on the same scale changes to 180 pounds an hour later and 100 pounds an hour after that. Based on the inconsistency of this scale, any research relying on it would certainly be unreliable. Consider an important study on a new diet program that relies on your inconsistent or unreliable bathroom scale as the main way to collect information regarding weight change. Would you consider their results accurate? A reliability coefficient is often the Statistic of choice in determining the ' reliability of a test. This coefficient merely represents a correlation, which measures the intensity and direction of a relationship between two or more variables. Test-Retest Reliability: Test-Retest reliability refers to the test’s consistency among different administrations. To determine the coefficient for this type of reliability, the same test is given to a group of subjects on at least two separate occasions. If the test is reliable, the scores that each student receives on the first administration should be similar to the scores on the second. We would expect the relationship between he first and Second administration to be a high positive correlation. One major concern with test-retest reliability is what has been terrified the memory effect. This is especially true when the two administrations are dose together in time. For example, imagine taking a short 10-question test on vocabulary and then terv minutes later being asked to complete the same test. Most of us will remember our responses and when we begin to answer again, we may just answer the way we did on the first test rather than reading through the questions carefully. This can create an artificially high reliability coefficient as subjects respond from their memory rather than the test itself. When a pre-test and post-test for an experiment is the same, the memory effect can play a role in the results. Parallel Forms Reliability: One way to assure that memory effects do not occur is to use a different pre- and post-test. In order for these two tests to be used in this manner, however, they must be parallel or equal in what they measure. To determine parallel forms reliability, a reliability coefficient is calculated on the scores of the two measures taken by the same group of subjects. Once again, we would expect a high and positive correlation is we are to say the two forms are parallel. Inter-Rater Reliability: Whenever observations of behavior are used as data in research, we want to assure that these observations are reliable. One way to determine this is to have two or more observers rate the same subjects and then correlate their observations. If, for example, rater A observed a child act out aggressively eight times, we would want rater B to observe the same amount of aggressive acts. If rater B witnessed 16 aggressive acts, then we know at least one of these two raters is incorrect. If there ratings are positively correlated, however, we can be reasonably sure that they are measuring the same construct of aggression. It does not, however, assure that they are measuring it correctly, only that they are both measuring it the same. Self-Check Exercise-1 1.    What is reliability in the context of measurement? a)    The accuracy of a measurement. b)    The consistency of a measurement. c)    The ability to measure different things. d)    The validity of a measurement. 2.    What does a high positive correlation in test-retest reliability indicate? a)    The test measures different constructs each time. b)    The test is not reliable. c)    The test produces similar scores on different occasions. d)    The test has low reliability. 11.3.2    Validity Validity is the extent to which a test measures what it claims to measure. It is vital for a test to be valid in order for the results to be accurately applied and interpreted. Validity isn’t determined by a single statistic; but by a body of research that demonstrates the relationship between the test and the behavior it is intended to measure. There are three types of validity: It is the strength of our conclusions, inferences or propositions. Validity refers to the degree in which our test or other measuring device is truly measuring what we intended it to measure. The test question “1 + 1 =     _______" is certainly a valid basic addition question because it is truly measuring a student’s ability to perform basic addition. It becomes less valid as a measurement of advanced addition because as it addresses some required knowledge for addition, it does not represent all of knowledge required for an advanced understanding of addition. On a test designed to measure knowledge of American History, this question becomes completely invalid. The ability to add two single digits has nothing do with history. For many constructs, or variables that are artificial or difficult to measure, the concept of validity becomes more Complex. Most of us agree that “1 + 1 = _______" would represent basic addition, but does this question also represent the construct of intelligence? Other constructs include motivation, depression, anger, and practically any human emotion or trait. If we have a difficult time defining the construct, we are going to have an even more difficult time measuring it. Construct validity is the term given to a test that measures a construct accurately and there are different types of construct validity that we should be concerned with. Three of these, concurrent validity, content validity, and predictive validity are discussed below. Concurrent Validity: Concurrent Validity refers to a measurement device’s ability to vary directly with a measure of the same construct or indirectly with a measure of an opposite construct. It allows you to show that your test is valid by comparing it with an already valid test. A new test of adult intelligence, for example, would have concurrent validity if it had a high positive correlation with the Wechsler Adult Intelligence Scale since the Wechsler is an accepted measure of the construct we call intelligence. An obvious concern relates to the validity of the test against which you are comparing your test. Some assumptions must be made because there are many who argue the Wechsler scales, for example, are not good measures of intelligence. Content Validity: Content validity is concerned with a test’s ability to include or represent all of the content of a particular construct. The question “1 + 1 =_______” may be a valid basic addition question. Would it represent all of the content that makes up the study of mathematics? It may be included on a scale of intelligence, but does it represent all of intelligence? The answer to these questions is obviously no. To develop a valid test of intelligence, not only must there be questions on math, but also questions on verbal reasoning, analytical ability, and every other aspect of the construct we call intelligence. There is no easy way to determine content validity aside from expert opinion. Predictive Validity: In order for a test to be a valid screening device for some future behaviour, it must have predictive validity. The entrance tests must possess predictive validity. The main concern with these, and many other predictive measures is predictive validity because without it, they would be worthless. Self-Check Exercise-2 1.    Which type of validity refers to a test’s ability to measure all aspects of a particular construct? a)    Concurrent Validity b)    Predictive Validity c)    Content Validity d)    Construct Validity 2.    Which type of validity is important for tests used as screening devices for future behavior? a)    Concurrent Validity b)    Predictive Validity c)    Content Validity d)    Inter-Rater Reliability 11.3.3    Objectivity, Usability and Practicability It refers to fairness and uniformity in the test scoring procedure. Examiner / rater bias is therefore non-existent in MCQ based tests. That is why they are called objective tests. Further, the analysis, of the test data is undertaken statistically which further assesses various dimensions of the tests to make them more precise and accurate instruments. MCQs are practically useful and efficient especially in a large scale testing situation unlike descriptive tests which are more resource intensive and demand time and money. The tests should be easy in administration and scoring and interpretation. These qualities fall under practicality. The tests should be feasible & usable. Quality of being usable means the test can be used in context to the objective to be achieved. Usability (practicality) includes ease in administration, scoring, interpretation and application, low cost, proper mechanical make - up. It should measure the objective to be achieved. Self-Check Exercise-3 1.    Why MCQ-based tests are considered objective? a)  They are easy to administer. b)  They are less time-consuming. c)  They eliminate examiner/rater bias. d)  They require less statistical analysis. 2.    Which quality is NOT mentioned as part of the practicality of a test? a)    Ease in administration b)    High cost c)    Scoring and interpretation d)    Proper mechanical make-up 11.4    NORMS FOR INTERPRETATION OF TEST SCORES Norms means distribution of scores on a particular test by a well-established group of people. There are different types of norm samples: National norms: based on particular population. Local norms: based on a region within a country Within group norms: usually based on national norms for some well identified sub-group within the population. It is important to know about the sample used to norm a test because an individual’s score on a test has meaning only in the context of the standardization sample. In psychological testing, scores on tests seldom have any concrete meaning, so that interpretation of a score on a test is always relative to how others scored on that test. 11.4.1    Developing & Using Test Norms to Compare Performance A test norm is a set of scalar data describing the performance of a large number of people on that test. Test norms can be represented by two important statistics: Means and Standard Deviations. The Standard Deviation (SD) tells us how spread out the distribution is around a central point. The standard deviation is, in effect, an average of the departure of people’s scores from the group mean. This is a measure of how spread out the group’s scores typically are. The standard deviation gives an indication of the degree of dispersion of the scores from the mean, and an estimate of the variability of results in the total sample. It is usefully employed in comparing the variability of different groups. It is used to develop Test Norms because, unlike percentiles, measures based on how much an individual deviates from the mean can be mathematically manipulated and compared. Let us suppose that we have two distributions, Group A and Group B, with the same mean but the performance of people Group A is more spread out. Group A will have a larger standard deviation than Group B, which has smaller individual differences. In Group B, each score is closer to the mean and this means there is a smaller standard deviation. A.    Describing Performance using Z-Scores The Z score, based on the mean and standard deviation, indicates how many standard deviations above or below the mean a score is. A Z score is merely a raw score, which has been changed to standard deviation units. The standard Z score is calculated by the formula: ( X-X ) Z = SD WhereZ - standard Z score X = each individual raw score X (bar) = mean score SD= standard deviation Usually when standard scores are used they are interpreted in relation to the normal distribution curve. Standard Z scores in standard deviation units are marked out either side of the mean. Those above the mean are positive, and those below the mean negative in sign. Therefore, by the calculation of the standard Z score an individual’s score can be viewed in relation to the rest of the distribution. The standard score is very important when comparing scores from different tests. Before these scores can be properly compared they must be converted to a common scale such as the standard Z score. The Z-score can then be used to express an individual’s score on different tests in terms of norms. One important advantage in using the normal distribution curve as a basis for test norms is that the standard deviation has a precise relationship with the area described by the curve. One standard deviation above and below the mean will include approximately 68% of the sample, as illustrated here. Z scores have a mean of 0 and a standard deviation of 1. Because of this, Z scores can be rather difficult and cumbersome to handle because most of them are decimals and approximately half of them can be expected to be negative. To remedy these drawbacks various transformed standard score systems have been derived. These simply entail multiplying the obtained Z score by a new standard deviation and adding it to a new mean. Both of these steps are devised to eradicate decimals and negative numbers. B.    Describing Performance Using Z-Scores The t in t score stands for transformed. A T score is a transformed z score. T scores and stens (see below) take the Z score and transform it into a more effective and user- friendly format. However critically, both still describe the level of deviation a particular score has from the mean. The T score (transformed score) is the most common standardized norm system for ability tests. The T score is derived from the Z score, transformed to a new scale, so that the: T score mean is set at 50 and the standard deviation at 10. Calculate the T score using the following formula: Z score X10+50 If Z score = 2 T score = 2 x 10 + 50 = 20 + 50-70 Although the T score can easily be calculated, it is most often used by reference to a norm table. In order to identify the relevant T score, you find the raw score in the body of the table and read across to the right hand side. The advantage of T scores over percentiles is that they are a linear scale with equal units of measurement. Thus they can be mathematically manipulated and give results which can be directly compared both within and between individuals. Their disadvantage is that they are not so easily explained or meaningful to those not knowledgeable about testing. C.    Describing Performance Using Stens Stens or Standard Tenths are the most commonly used scales for comparing individual differences. The normal distribution is divided into ten stens. Unlike percentiles, stens are equal units of measurement, and unlike percentiles they are not influenced by clustering around the midpoint. A Sten is based on a mean of 5.5 and a standard deviation of 2. Calculate the sten using the following formula: Z score x 2 + 5.5 If Z score =2 Sten              =2x2+5.5 = 9.5 Normally, the sten is taken to the nearest whole number, with a minimum value of 1 and. a maximum of 10. Stens are commonly used as the norm system for personality questionnaires. However, they are also a useful system for ability tests as they encourage us to think in bands of scores but still provide sufficient discrimination between applicants. Self-Check Exercise-4 1.    What does a standard deviation (SD) indicate in a set of test norms? a)    The average score of the group. b)    The degree of dispersion of scores from the mean. c)    The highest score in the group. d)    The median score of the group. 2.    How is a Z score interpreted in relation to the normal distribution curve? a)    It indicates the raw score of an individual. b)    It indicates how many standard deviations a score is from the mean. c)    It shows the percentage of people scoring above the mean. d)    It measures the average deviation of scores from the mean. 11.5    SUMMARY Evaluation in education is critical for assessing student achievement, relying on robust measurement tools characterized by reliability, validity, objectivity, usability, and practicability. Reliability, reflecting consistency in test results, can be estimated through methods like test-retest, parallel forms, internal consistency, and interrater reliability. Validity ensures the test measures what it purports to, while objectivity minimizes scorer bias. Various procedures develop test score norms, enhancing interpretative accuracy. Methods to ascertain reliability include the parallel form method, Kuder-Richardson formula, test-retest method, and split-half method, each with distinct advantages and limitations. These principles ensure that educational assessments are both accurate and fair. 11.6    GLOSSARY Reliability: it is the degree of consistency of a measure. Validity: Validity refers to how accurately a method measures what it is intended to measure. 11.7    ANSWERS TO SELF-CHECK EXERCISE 1, 2, 3 & 4. Self-Check Exercise-1 1.    b) The consistency of a measurement 2.    c) The test produces similar scores on different occasions. Self-Check Exercise-2 1.    c) Content Validity 2.    b) Predictive Validity Self-Check Exercise-3 1.    c) They eliminate examiner/rater bias. 2.    b) High cost Self-Check Exercise-4 1.    b) The degree of dispersion of scores from the mean 2.    b) It indicates how many standard deviations a score is from the mean. 11.8 REFERENCES/ SUGGESTED READINGS •  Education Measurement and  Evaluation: J. Swarupa Rani, Discovery PublishingHouse. •  Measurement and Evaluation  in Teaching: Norman Edward Gronlund Macmillian •    Measurement and Assessment in Teaching: Robert L. Linn Pearson Education India. •    Program Evaluation and performance measurement: James C. Me. David, Laura R.L. Hawthorn, Sage Publication 11.9 TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    Explain the concept of reliability in the context of educational evaluation tools. Provide an example to illustrate your explanation. 2.    Discuss the different types of reliability mentioned in the text. Why is interrater reliability particularly important for performance-based assessments? 3.    What are the methods for estimating the reliability of a test? Provide a brief description of one method. 4.    Describe the 'Split-Half Method' for calculating the reliability of a test and mention one of its limitations. UNIT-12 TYPES OF TESTS Structure 12.1    Introduction 12.2  Learning Objectives 12.3  Types of Tests Self-Check Exercise-1 12.4    Standardized Test and Diagnostic Test Self-Check Exercise-2 12.5    Practical Test and Mastery Test Self-Check Exercise-3 12.6    Summary 12.7    Glossary 12.8  Answers to Self-Check Exercise 12.9  References /suggested Readings 12.10    Terminal Questions 12.1    INTRODUCTION Dear Learner, For measuring various characteristics of the students, we employ different types of instruments. Tests are the devices used to measure abilities, achievement or skills of the individuals. In various fields like education, psychology, healthcare, and employment, tests are used to measure knowledge, skills, attitudes, and behaviours. Different types of tests serve distinct purposes, and understanding these differences is crucial for appropriate test selection and interpretation. 12.2 LEARNING OBJECTIVES After studying this unit, you will be able to: •      List down different types of tests. •     Differentiate between speed and power tests. •      Explain ability and achievement tests. •     Compare objective and subjective type of tests. •      Define mastery tests. •      Describe standardized test. 12.3    TYPES OF TESTS On the basis of various dimensions, tests can be divided into different types. Different types of tests are explained below: Speed Test and Power Test: Some tests have very easy items but there is a limited amount of time to answer them. Such speed tests are used to see how quickly students can work on skills they have already mastered. One example is a test of keyboard skills. The teacher may wish to find out how fast students can maintain accurate work when typing data on a typewriter or computer keyboard. In contrast, power tests are concerned with identifying skills which have been mastered. Power tests require adequate samples of student behaviour -so having sufficient time to attempt most of the items is an essential prerequisite for such tests. Aptitude or Ability Test and Achievement Test: Achievement tests may be used to assess the extent to which curriculum objectives have been met in an educational programme. Such tests should have tasks which relate to the learning that students have to demonstrate. Since future learning depends to some extent on past learning, success on such achievement tests may provide evidence of future success (provided that other conditions such as good teaching, adequate health care, and stable family circumstances are maintained). Tests which are constructed specifically to gather evidence about ability to learn are referred to as aptitude or ability tests. Results on such tests are used to predict future success on the basis of success on the specially selected tasks in the aptitude test. Often these tasks differ from the usual school learning requirements and depend to some extent on learning beyond the school curriculum. Of course, teaching students the test items and the corresponding answers may result in an increase in score without actually changing a student’s (real) aptitude. Objective and Subjective Test: The term ‘objective’ can have several meanings when describing a test. It can mean that the score key for the test needs a minimum of interpretation in order to score an item correct or incorrect. In this sense, an objective test is one which requires task responses which can be scored accurately and fairly from the score key without having knowledge of the content of the test. For example, a multiple-choice test can be scored by a machine or by a clerical worker without either the machine or the clerical worker having had to reach a high level of expertise on the material being tested. A less common usage relates to the extent of agreement between experts about the correct answer. If there is less argument about the correct answer the item is regarded as more objective. However the choice of which items (whether objective in their answer format or not) are to appear on a test is subjective, in that it depends on the personal preferences and experiences of those constructing the test. Objective tests are especially well suited to certain types of tasks. Because questions can be designed to be answered quickly, they allow lecturers to test students on a wide range of material.... Additionally, statistical analysis on the performance of individual students, cohorts and questions is possible. The capacity of objective tests to assess a wide range of learning is often underestimated. Objective tests are very good at examining recall of facts, knowledge and application of terms, and questions that require short text or numerical responses. But a common worry is that objective tests cannot assess learning beyond basic comprehension. There are, however, limits to what objective tests can assess. They cannot, for example, test the competence to communicate, the skill pf constructing arguments or the ability to offer original responses. Tests must be carefully constructed in order to avoid the de contextualisation of knowledge and it is wise to use objective testing as only one of a variety of assessment methods within a module. However, in times of growing student numbers and decreasing resources, objective testing can offer a viable addition to the range of assessment types available to a teacher or lecturer. Objective tests are appropriate when: •    The group to be tested is large and the test may be reused. •    Highly reliable scores must be obtained as efficiently as possible. •    Impartiality of evaluation, fairness, and free from possible test scoring influences are essential. Essay tests are appropriate when: •  The group to be tested is small and the test is not to be reused •  You wish to encourage and reward the development of student skill in writing - •    You are more interested in exploring student attitudes than in measuring his/her achievement Either essay or objective tests can be used to: •    Measure almost any important educational achievement a written test can measure •    Test understanding and ability to apply principles. •    Test ability to think critically. •    Test ability to solve problems. •    In general, question types fall into two categories: 1.    Objective, which require students to select the correct response from several alternatives or to supply a word or short phrase to answer a question or complete a statement. 2.    Subjective or essay, which permit the student, to organize and present an original answer. Examples: short-answer essay, extended-response essay, problem solving, performance test items. Subjective items require students to write and present an original answer. It includes short - answer essay, extended - response essay, problem solving, and performance tasks. Advantages of subjective teste are higher learning skills are utilized by learners, for example synthesis, analysis and evaluation. Brevity and consciousness, precious of expression is developed among learners. It can quickly and easily constructed and eliminates guessing. Disadvantages of subjective tests are subjectivity the same piece of work can get different marks. Students with poor language prowess tend to fail. Time is consumed when answering these questions, usually it is limited in scope, and thus it doesn’t cover much content. The examiner's judgment determines the final grade. The research recommends taking both subjective and objective teste for a classroom test. The universal adoption of a combination objective / subjective testing format would tend to sharpen writing and organizational skills. Comparison between Objective and Subjective Tests: Objective tests are so called according to their scoring which depends on personal judgments or opinions. Subjective tests are so called because their scoring depends on personal judgments or opinions, the techniques used in objective tests multiple - choice items (MCI), True / false items, matching items, transformation sentences, re-arrangement items and fill the blanks or gap filling. On the other hand, the techniques used in subjective tests include: essay writing, composition writing, letter writing, reading aloud, completion type and answer these - questions. To answer an objective test, the testee has to select his answers from two, three, four or even more alternatives while objective tests which [has] only one correct answer. Besides, to answer an objective test, the testee has to plan and write his own answer by using his own words and expressions. Furthermore, objective tests need much time and effort to write the questions because the examiner has to provide the answers as well as the question so that objective test requires more careful preparations than 46 other types of test. But in subjective test the examiner needs to write few questions without answers. In objective tests, it seems that kind is more, reliable because it gives a stable scoring. But in subjective test, it seems, it is not reliable because it doesn’t give a stable scoring. Objective tests are used to test structures, vocabulary, comprehension, and sound discrimination. On the other hand, subjective tests are used to test ideas, culture, coherence and creativity. Objective tests encourage guessing and it is difficult to write simple to answer, easy to score, suit for a large number of testees, and this type of test can be scored by a machine. Besides, subjective test doesn't encourage guessing easy to write, difficult to score and suit for a small number of testee. This type of test can't be scored by a machine. Objective test can be used to test specific area of language, while, subjective test can be used to evaluate overall achievements. Furthermore, objective tests require recognition more than production but subjective tests require production as well as recognition. An objective test is a type of discrete point test, but subjective test is a type of an integrative point test. Objective test is a type of close - ended atomistic, a system - referenced and it applicative test and replicative test. On the other hand, subjective test is a type open - ended test, a holistic and replicative and it isn’t applicative. Objective test depends on students' knowledge. It is a valid test which student's need a short time to answer than subjective test. But subjective test depends on student's experience. It is invalid test and students need a long time to answer than objective test. Finally, a good classroom test should be contained both subjective and objective. Self-check Exercise-1 1.    What is the primary purpose of an aptitude or ability test? a)    To assess curriculum objectives met in an educational program. b)    To predict future success based on specific tasks. c)    To evaluate the quality of teaching methods. d)    To measure students' past achievements. 2.    Which of the following is an advantage of objective tests? a)    They can assess the skill of constructing arguments. b)    They require extensive interpretation to score. c)    They allow for testing a wide range of material efficiently. d)    They are subjective and rely on the examiner's judgment. 12.4    STANDARDIZED TEST AND DIAGNOSTIC TEST Standardized Test: The term ‘standardized’ also has a number of meanings with respect to testing. It can mean that the test has an agreed format for administration and scoring so that the task is as identical as possible for all candidates and there is little room for deviation in the scoring of candidate responses to the tasks. Another meaning refers to the way in which the scores on a tests are presented. For example, if scores are given as a raw score divided by some measure of dispersion like the standard deviation, the resulting score scale is said to be in terms of standardized scores (sometimes called standard scores). Finally, the term can refer (loosely) to a published test which was prepared by standard (or conventional) procedures. The usage of ‘standardized’ has become somewhat confused because published tests often present scores interpreted in terms of deviation from the mean (or average) and have a standard procedure for administering tests and interpreting results. Diagnostic Test: This term refers to the use made of the information gained from administration of the test. The implication is that the test results will assist in identifying both the topics which are not known and in providing information on potential sources of the student’s difficulty. Teachers may be expected to provide appropriate teaching for each difficulty exposed by the use of a diagnostic test. For example, a simple open-ended mathematics question about area, given to junior secondary level classes provided a range of correct and incorrect answers. Self-check Exercise-2 1.    What is one meaning of a 'standardized' test? a)    A test that has a flexible format for administration. b)    A test that presents scores in terms of deviation from the mean. c) A test that has no specific scoring method. d)    A test that is personalized for each candidate. e) 2.    Which of the following best describes the purpose of a diagnostic test? a)    To compare scores among a large group of students. b)    To identify topics that are not known and sources of difficulty for students. c)    To measure the overall achievement of students in a standardized manner. d)    To provide a raw score without interpretation. 12.5    PRACTICAL TEST AND MASTERY TESTS Practical Test: In some senses, an essay test is a practical task. The essay item requires a candidate to perform. This performance is intended to convey meaning in a practical sense by writing prose to an agreed format. However, the term 'practical test’ goes beyond performance and other tasks used in traditional pencil-and-paper examinations. The term may refer to practical tasks in trade subjects (such as woodwork, metalwork, shipbuilding, and leather craft), in musical and dramatic performance, in skills such as swimming or gymnastics, or may refer to the skills required to carry out laboratory or field tasks in science, agriculture, geography, environmental health or physical education. Mastery Tests: These tests are generally criterion-referenced tests with a relatively high score requirement. Students who meet this high score are said to have mastered the topic. It is assumed that the mastery test has sufficient items of high quality to ensure that the score decision is well founded with respect to the domain of interest. For example, in Mathematics the domain might be ‘all additions of pairs of one-digit numbers where the total does not exceed 9'. A mastery test of this domain should have a reasonable sample of all possible combinations of one-digit numbers because the mastery decision implies that all can be added successfully even though all are not tested. This simple example should not be taken to imply that mastery testing is limited to relatively trivial skills. A more complex example is the regular testing of airline pilots. Safety requirements result in high standards being set for mastery in many areas. Failure to reach mastery will result either in further tuition under the guidance of an experienced tutor or in withdrawal of the permission to fly. Self-check Exercis-3: 1.    Which of the following is an example of a practical test? a)    A multiple-choice test in mathematics. b)    An essay test in English literature. c)    A swimming skills assessment. d)    A standardized test in history. 2.    What characterizes a mastery test? a)    It is usually a norm-referenced test with moderate scoring requirements. b)    It is a criterion-referenced test with a relatively high score requirement. c)    It involves subjective evaluation of performance. d)    It only tests trivial skills. 12.6    SUMMARY In this lesson, we studied about various types of test. After going through this lesson, you would have understood the differences between various types of tests. You have also become clear about the concept of norm-referenced and criterion-referenced types of teste and tee differences between these two types. 12.7    GLOSSARY Diagnostic Test: It is a form of pre-assessment or a pre-test where teachers can evaluate students’ strengths, weaknesses, knowledge and skills before their instruction. Practical Test: An assessment that evaluates an individual's ability to apply their knowledge and skills in a real-world setting. 12.8    ANSWER TO SELF-CHECK EXERCISE 1, 2 & 3. Self-check Exercis-1 1.    b) To predict future success based on specific tasks. 2.    c) They allow for testing a wide range of material efficiently. Self-check Exercis-2 1.    b) A test that presents scores in terms of deviation from the mean. 2.    b) To identify topics that are not known and sources of difficulty for students. Self-check Exercis-3 1.    c) A swimming skills assessment. 2.    b) It is a criterion-referenced test with a relatively high score requirement 12.9    REFERENCES /SUGGESTED READING •    Ebel, Robert L.(1966) “Measuring Educational Achievement, Prentice Hall of India Pvt. Ltd. Pp.481 •    Gronlund, N. E. (1976). Measurement and'Evaluation in Teaching. McMillan, USA. •    Hopkins, C.D. and Antes, R.L. (1990).Classroom measurement and evaluation. Itasca, Illinois: Peacock. •    Izard, J. (1991). Assessment of learning in the classroom. Geelong, Vic.: Deakin University. 12.10    TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    Name various types of tests. 2.    What is the difference between speed and power tests? 3.    Write down the advantages of objective type tests over subjective type tests. 4.    What do you mean by standardized test? 5.     Explain mastery tests. *********** UNIT-13 CRITERION-REFERENCED TESTS AND NORM-REFERENCED TESTS Structure 13.1    Introduction 13.2  Learning Objectives 13.3 . Meaning of Criterion Referenced Test (CRT) 13.3.1    Features of Criterion Referenced Test (CRT) 13.3.2    Advantages and Disadvantages of Criterion Referenced Test (CRT) Self-Check Exercise-1 13.4    Meaning of Norm Referenced Test (NRT) 13.4.1    Features of Norm Referenced Test (NRT) 13.4.2    Advantages and Disadvantages of Norm Referenced Test (NRT) Self-Check Exercise-2 13.5    Summary 13.6    Glossary 13.7  Answers to Self-Check Exercise 13.8  References /Suggested Readings 13.9    Terminal Questions 13.1    INTRODUCTION Dear Learner, Assessments and tests are essential tools in various fields, including education, employment, and certification. There are two primary categories of tests: Norm-Referenced Tests (NRTs) and Criteria-Referenced Tests (CRTs). Understanding the differences between these two types of tests is crucial for appropriate test selection, interpretation, and use. Norm-Referenced Tests (NRTs) compare individual performance to that of a larger group, known as the norm group. These tests aim to rank individuals relative to others, providing information on their relative standing within the group. NRTs are often used for: Selection and admission processes, ranking and comparison purposes, identifying strengths and weaknesses relative to others. Criteria-Referenced Tests on the other hand, evaluate individual performance against specific criteria, standards, or learning objectives. These tests aim to determine whether individuals have met predetermined criteria, regardless of how others perform. CRTs are often used for: Evaluating mastery or proficiency, diagnosing strengths and weaknesses, Certifying competence or licensure 13.2    LEARNING OBJECTIVES After studying this unit, you will be able to: Explain criterion-referenced tests. Explain norm-referenced test. Describe the features of CRT and NRT. • Enlist the advantages and disadvantages of CRT and NRT. 13.3    MEANING OF CRITERION REFERENCED TEST (CRT) Definition: A test in which questions are written according to specific predetermined criteria. A student knows what your standards are for passing and only competes against him or herself while completing. A criterion-referenced test is designed to measure how well test takers have mastered a particular body of knowledge. These tests generally have an established “passing” score. Students know what the passing score is and an individual’s test score is determined by knowledge of the course material. He runs a second try out, having established that 10 seconds in the 100 yard dash is competitive in the event. He now picks those who run the dash in 10 seconds or less. This is criterion-referenced testing. He knows the runners he selected can compete. He gets the funds. 13.3.1    Features of Criterion Referenced Test (CRT) Features of Criterion reference tests are as follows:- (i)    Criterion-referenced test place a primary focus on the content and what is being measured. Norm-referenced tests are also concerned about what is being measured but the degree of concern is less since the domain of content is not the primary focus for score interpretation. In norm-referenced test development, item selection, beyond the requirement that items meet the content specifications, is driven by item statistics. Items are needed that are not too difficult or too easy, and that are highly discriminating. These are the types of items that contribute most to score spread, and enhance test score reliability and validity. (ii)    With criterion-referenced test development, extensive efforts go into insuring content validity. Item statistics play less a role in item selection though highly discriminating items are still greatly valued, and sometimes item statistics are used to select items that maximize the discriminating power of a test at the performance standards of interest on the test score scale. A good norm-referenced test is one that will result in a wide distribution of scores on the construct being measured by the test. Without score variability, reliable and valid comparisons of candidates cannot be made. A good criterion-referenced test will permit content-referenced interpretations and this means that the content domains to which scores are referenced must be very clearly defined. Each type of test can serve the other main purpose (norm- referenced versus criterion-referenced interpretations), but this secondary use will never be optimal. For example, since criterion-referenced tests are not constructed to maximize score variability, their use in comparing candidates may be far from optimal if the test scores that are produced from the test administration are relatively similar. Because the purpose of a criterion-referenced test is quite different from that of a norm referenced test, it should not be surprising to find that the approaches used for reliability and validity assessment are different too. (iii)    With criterion-referenced tests, scores are often used to sort candidates into performance categories. Consistency of scores over parallel administrations becomes less central than consistency of classifications of candidates to performance categories over parallel administrations. Variation in candidate scores is not so important if candidates are still assigned to the same performance category. Therefore, it has been common to define reliability for a criterion-referenced test as the extent to which performance classifications are consistent over parallel-form administrations. For example, it might be determined that 80% of the candidates are classified in the same Notes way by parallel forms of a criterion-referenced test administered with little or no instruction in between test administrations. This is similar to parallel form reliability for a norm referenced test except the focus with criterion-referenced tests is on the decisions rather than the scores. Because parallel form administrations of criterion-referenced tests are rarely practical, over the years methods have been developed to obtain single administration estimates of decision consistency that are analogous to the use of the corrected split-half reliability estimates with norm-referenced tests. (iv)    With criterion-referenced tests, the focus of validity investigations is on (1) the match between the content of the test items and the knowledge or skills that they are intended to measure, and (2) the match between the collection of test items and what they measure and the domain of content that the tests are expected to measure. The “alignment” of the content of the test to the domain of content that is to be assessed is called content validity evidence. This term is well known in testing practices. Many criterion-referenced tests are constructed to assess higher-level thinking and writing skills, such as problem solving and critical reasoning. Demonstrating that the tasks in a test are actually assessing the intended higher-level skills is important, and this involves judgments and the collection of empirical evidence. So, construct validity evidence too becomes crucial in the process of evaluating a criterion referenced test. (v)    Probably the most difficult and controversial part of criterion-referenced testing is setting the performance standards, i.e., determining the points on the score scale for separating candidates into performance categories such as “passers” and “failers.” The challenges are great because with criterion-referenced tests in education, it is common on state and national assessments to separate candidates into not just two performance categories, but more commonly, three, four, or even five performance categories. With four performance categories, these categories are often called failing, basic, proficient, and advanced. 13.3.2    Advantages and Disadvantages of Criterion Referenced Test (CRT) Advantages of CRT 1.    Mastery of Subject Matter. • Criterion-referenced tests are more suitable than norm-referenced tests for tracking the progress of students within a curriculum. Test items can be designed to match specific program objectives. The scores on a criterion referenced tests indicate how well the individual can correctly answer questions on the material being studied, while the scores on a norm-referenced test report how the student scored relative to other students in the group. 2.    Criterion-Referenced Tests can be Managed Locally. •    Assessing student progress is something that every teacher must do. Criterion-referenced tests can be developed at the classroom level. If the standards are not met, teachers can specifically diagnose the deficiencies. Scores for an individual student are independent of how other students perform. In addition, test results can be quickly obtained to give students effective feedback on their performance. Although norm-referenced tests are most suitable for developing normative data across large groups, criterion-referenced tests can produce some local norms. Disadvantages of CRT •    Criterion-referenced tests have some built-in disadvantages. Creating tests that are both valid and reliable requires fairly extensive and expensive time and effort. In addition, results cannot be generalized beyond the specific course or program. Such tests may also be compromised by students gaining access to test questions prior to exams. Criterion referenced tests are specific to a program and cannot be used to measure the performance of large groups. Analysing Test Items •    Item analysis is used to measure the effectiveness of individual test items. The main purpose is to improve tests, to identify questions that are too easy, too difficult or too susceptible to guessing. While test items can be analyzed on both criterion-referenced and norm-referenced tests, the analysis is somewhat different because the purpose of the two types of tests is different. Self-Check Exercise-1 1.    What is the primary focus of a criterion-referenced test (CRT)? a)    Comparing a student’s performance to that of other students. b)    Measuring how well students have mastered a particular body of knowledge. c)    Selecting items that maximize score variability. d)    Ensuring a wide distribution of scores. 2.    Which of the following is an advantage of criterion-referenced tests (CRT)? a)    They are inexpensive and quick to develop. b)    They can be generalized beyond the specific course or program. c)    They are more suitable for tracking student progress within a curriculum and diagnosing specific deficiencies. d)    They are ideal for measuring the performance of large groups. 3.Fill in the blanks: (i)    Criterion reference tests have an established ______ scores. (ii)    ______ use criterion reference test to monitor student performance in their day to day activities. (iii)    Criterion reference tests are also used in the training programs to assess 13.4 MEANING OF NORM REFERENCED TEST (NRT) Definition: This type of test determines a student’s placement on a normal distribution curve. Students compete against each other on this type of assessment. This is what is being referred to with the phrase, ‘grading on a curve’. Norm-referenced tests allow us to compare a student’s skills to others in his age group. Norm-referenced tests are developed by creating the test items and then administering the test to a group of students that will be used as the basis of comparison. Statistical methods are used to determine how raw scores will be interpreted and what performance levels are assigned to each score. Many tests yield standard scores, which allow comparison of the student’s scores to other tests. They answer questions such as, “does the student’s achievement score appear consistent with his cognitive score?” The degree of difference between those two scores might suggest or rule out a learning disability. After the norming process, the tests are used to assess groups of students or individuals using standardized, or highly structured, administration procedures. These students’ performance is rated using scales developed during the norming process. Educators use norm-reference tests to evaluate the effectiveness of teaching programs, to help determine students’ preparedness for programs, and to determine diagnosis of disabilities for eligibility for IDEA special education programs or adaptations and accommodations. 13.4.1    Features of Norm Referenced Test (NRT) Norm-referenced tests (NRTs) compare a person’s score against the scores of a group of people who have already taken the same exam, called the “norming group.” When you see scores in the paper which report a school’s scores as a percentage -“the Lincoln school ranked at the 49th percentile” -- or when you see your child’s score reported that way -- “Jamal scored at the 63rd percentile” -- the test is usually an NRT. Most achievement NRTs are multiple-choice tests: Some also include open-ended, short-answer questions. The questions on these tests mainly reflect the content of nationally-used textbooks, not the local curriculum. This means that students may be tested on things your local schools or state education department decided were not so important and therefore were not taught. Creating the bell curve. NRTs are designed to “rank-order” test takers -- that is, to compare students’ scores: A commercial Norm-referenced test does not compare all the students who take the test in a given year. Instead, test-makers select a sample from the target student population (say, ninth graders). The test is “normed” on this sample, which is supposed to fairly represent the entire target population (all ninth graders in the nation). Students’ scores are then reported in relation to the scores of this “norming” group. To make comparing easier, test makers create exams in which the results end up looking at least somewhat like a bell-shaped curve :Testmakers make the test so that most students will score near the middle, and only a few will score low (the left side of the curve) or high (the right side of the curve). Scores are usually reported as percentile ranks: The scores range from 1st percentile to 99th percentile, with the average student score set at the 50th percentile. If Jamal scored at the 63rd percentile, it means he scored higher than 63% of the test takers in the norming group. Scores also can be reported as “grade equivalents,” “stanines,” and “normal curve equivalents.” One more question right or wrong can cause a big change in the student’s score : In some cases, having one more correct answer can cause a student’s reported percentile score to jump more than ten points. It is very important to know how much difference in the percentile rank would be caused by getting one or two more questions right. In making an NRT, it is often more important to choose questions that sort people along the curve than it is to make sure that the content covered by the test is adequate: The tests sometimes emphasize small and meaningless differences among test takers. Since the tests are made to sort students, most of the things everyone knows are not tested. Questions may be obscure or tricky, in order to help rank order the test takers. Tests can be biased: Some questions may favor one kind of student or another for reasons that have nothing to do with the subject area being tested. Nonschool knowledge that is more commonly learned by middle or upper class children is often included in tests. To help make the bell curve, test makers usually eliminate questions that students with low overall scores might get right but those with high overall scores get wrong. Thus, most questions which favor minority groups are eliminated. NRTs usually have to be completed in a time limit: Some students do not finish, even if they know the material. This can be particularly unfair to students whose first language is not English or who have learning disabilities. This “speededness” is one way test makers sort people out. 13.4.2    Advantages and Disadvantages of Norm Referenced Test (NRT) Advantages of NRT To compare students, it is often easiest to use a Norm-Referenced Test because they were created to rank test-takers: It there are limited places (such as in a “Gifted and Talented” program) and choices have to be made, it is tempting to use a test constructed to rank students, even if the ranking is not very meaningful and keeps out some qualified children. NRT’s are a quick snapshot of some of the things most people expect students to learn: They are relatively cheap and easy to administer. If they were only used as one additional piece of information and not much importance was put on them, they would not be much of a problem. Disadvantages NRT The damage caused by using NRTs is far greater than any possible benefits the tests provide: The main purpose of NRTs is to rank and sort students, not to determine whether students have learned the material they have been taught. They do not measure anywhere near enough of what students should learn. They have very harmful effects on curriculum and instruction. In the end, they provide a distorted view of learning that then causes damage to Norm-Referenced Measures (NRM) Most appropriate when one wishes to make comparisons across large numbers of students or important decisions regarding student placement and advancement. Norm-referenced measures are designed to compare students (i.e., disperse average student scores along a importance placed upon high scores, the content of a standardized test can be very influential in the development of a school’s curriculum and standards of excellence. The testing profession, in its Standards for Educational and Psychological Measurement, states, “In elementary or secondary education, a decision or characterization that will have a major impact on a test taker should not automatically be made on the basis of a single test score.” Any one test can only measure a limited part of a subject area or a limited range of important human abilities: A “reading” test may measure only some particular reading “skills,” not a full range of the ability to understand and use texts. Multiple-choice math tests can measure skill in computation or solving routine problems, but they are not good for assessing whether students can reason mathematically and apply their knowledge to new, real-world problems. Most NRTs focus too heavily on memorization and routine procedures: Tests like these cannot show whether a student can write a research paper, use history to help understand current events, understand the impact of science on society, or debate important issues. They don’t test problem solving, decisionmaking, judgement, or social skills. Tests often cause teachers to overemphasize memorization and de-emphasize thinking and application of knowledge: Since the tests are very limited, teaching to them narrows instruction and weakens curriculum. Making test score gains the definition of “improvement” often guarantees that schooling becomes test coaching. As a result, students are deprived of the quality education they deserve. Norm-referenced tests also can lower academic expectations: NRTs support the idea that learning or intelligence fits a bell curve. If educators believe it, they are more likely to have low expectations of students who score below average. Self-Check-Exercise-2 1.    Which feature is associated with norm-referenced tests (NRTs)? a)    Items are selected to ensure they align with the curriculum being taught locally. b)    Scores are often reported as percentile ranks. c)    Focuses on ensuring all students meet a predetermined standard. d)    Emphasizes individual mastery of subject matter. 2.    What is a significant disadvantage of norm-referenced tests (NRTs)? a)    They are expensive and time-consuming to develop. b)    They often cause teachers to overemphasize memorization. c)    They are not suitable for large-scale assessments. d)    They provide immediate feedback to students and teachers. 3.    Fill in the blanks: (i)    Norm-referenced tests (NRTs) determine a student's placement on a ________ distribution curve. (ii)    Scores on NRTs are usually reported as ________ ranks, which range from 1st percentile to 99th percentile. (iii)    One disadvantage of NRTs is that they often cause teachers to overemphasize ________ and de-emphasize thinking and application of knowledge. 13.5    Summary In summary, NRTs compare individuals to a norm group, while CRTs evaluate individuals against specific criteria or standards. NRTs focus on relative standing, while CRTs focus on absolute achievement. Both NRT’s and CRT’s used to evaluate the performance of learners and determine whether they have failed or excelled in their tests. It is after this that the students can be held accountable and told to re-sit their tests. 13.6    Glossary Norm Group: A large, representative sample of individuals used as a comparison group for NRTs. Standards:  Specific learning objectives or criteria used to evaluate performance on CRTs. Mastery: Demonstrating a predetermined level of proficiency or competence on a CRT. 13.7    ANSWER TO SELF-CHECK EXERCISE 1 & 2. Self-Check-Exercise-1 1.    b) Measuring how well students have mastered a particular body of knowledge 2.    c) They are more suitable for tracking student progress within a curriculum and diagnosing specific deficiencies. 3.  (i) passing (ii)    classroom teachers (iii)    learning Self-Check-Exercise-2 1.    b) Scores are often reported as percentile ranks. 2.    b) They often cause teachers to overemphasize memorization. 3.  (i) normal (ii)    percentile (iii)    memorisation 13.8    REFERENCES/SUGGESTED READING •    Taiwo, Adediran A. (2005). Fundamentals of Classroom Testing. New Delhi: Vikas Publishing House Pvt. Ltd. •    Walter W. Cook (1958). Educational Measurement. Washington D.C.: American Council on Education. •    Withers, G. (1997).Item writing for Tests and Examinations. Paris: International Institute for Educational Planning. 13.9    TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    Discuss the primary differences between Norm-Referenced Tests (NRTs) and Criterion-Referenced Tests (CRTs). 2.    Evaluate the advantages and disadvantages of Criterion-Referenced Tests (CRTs). 3.    Explain the feature of criterion referenced test. 4.    What are the features of norm referenced test? 5.    Give the advantages and disadvantages of norm referenced test? ******* UNIT-14 QUESTIONNAIRE AND SCHEDULES Structure 14.1    Introduction 14.2    Learning Objectives 14.3    Questionnaire 14.3.1    Characteristics of a well-designed questionnaire 14.3.2    Types of questionnaires 14.3.3    Steps of Questionnaire 14.3.4    Advantages and Limitations of Questionnaire Self-Check Exercise-1 14.4    Schedule, Comparison between questionnaire and schedule Self-Check Exercise-2 14.5    Summary 14.6    Glossary 14.7    Answers to Self-Check Exercise 14.8    References/Suggested Readings 14.9    Terminal Questions 14.1    INTRODUCTION Dear Learner, A questionnaire is a research tool consisting of questions to collect data from respondents, using both close-ended and open-ended formats. A well-designed questionnaire features a limited number of questions, logical sequencing, simplicity, and clear instructions. Types include structured, unstructured, and open-ended questionnaires. Designing one involves defining goals, identifying respondents, developing questions, and conducting a pilot test. Advantages include standardization, efficiency, anonymity, costeffectiveness, objectivity, and flexibility, but limitations are limited depth, response bias, and lack of context. In contrast, a schedule is filled out by researchers based on informant responses, offering more control and depth but with limited reach and requiring researcher involvement. Understanding these tools aids in selecting the appropriate method for effective data collection. 14.2    LEARNING OBJECTIVES After studying this unit, you will be able to: •    Describe the characteristics of a good questionnaire. •    Explain various types of questionnaires. •    Explain in brief, the steps of designing a questionnaire. •    Discuss advantages and limitations of questionnaire. •    Explain Schedule. •    Compare between questionnaire and schedule. 14.3    QUESTIONNAIRE A questionnaire is a research instrument that consists of a set of questions or other types of prompts that aims to collect information from a respondent. A research questionnaire is typically a mix of close-ended questions and open-ended questions. Open-ended, long-form questions offer the respondent the ability to elaborate on their thoughts. Research questionnaires were developed in 1838 by the Statistical Society of London. 14.3.1    Characteristics of a well-designed questionnaire Here are the key characteristics of a well-designed questionnaire: 1.    Limited Number of Questions: Keep the number of questions as limited as possible, focusing only on those relevant to the inquiry. 2.    Proper Sequence of Questions: Arrange questions logically, placing simple and direct ones at the start and more complex or indirect ones toward the end. 3.    Simplicity: Use clear,  easy-to-understand  language with short questions. Avoid complexity. 4.    Clear Instructions: Provide explicit instructions for filling out the form 14.3.2    Types of questionnaires Let’s explore the different types of questionnaires: 1.    Structured Questionnaire: •    This type of questionnaire has a fixed format with predetermined questions that the respondent must answer. •    The questions are usually closed-ended, meaning that the respondent selects a response from a list of options (e.g., yes/no or multiple-choice questions). 2.    Unstructured Questionnaire: •    An unstructured questionnaire does not have a fixed format or predetermined questions. •    Instead, the interviewer or researcher can ask open-ended questions, allowing respondents to provide their own answers without predefined options. 3.    Open-ended Questionnaire: •    In an open-ended questionnaire, respondents answer questions in their own words, without any pre-determined response options. •    This format allows for more details and qualitative responses. 14.3.3    Steps of Questionnaire: Designing a questionnaire involves several key steps: 1.    Define Goals and Objectives: •  Clarify what you want to study and the specific topics or experiences you aim to explore. •  Consider the purpose of the questionnaire (e.g., research, feedback, assessment). 2.    Identify Target Respondents: •    Understand your audience: Who are the respondents? What characteristics do they have? •    Tailor the questionnaire to their background, interests, and needs. 3.    Develop Questions: •    Create valid and reliable questions aligned with your research objectives. •    Use clear and concise language. •    Consider both closed-ended (predefined options) and open-ended questions. 4.    Choose Question Types: •  Closed-ended questions: Provide response options (e.g., yes/no, Likert scale). •  Open-ended questions: Allow respondents to answer in their own words. 5.    Design Question Sequence and Layout: •    Arrange questions logically. •  Start with simple, non-threatening questions. •  Group related questions together. 6.    Run a Pilot: •    Test the questionnaire with a small group to identify any issues. •    Revise as needed based on feedback. 14.3.4    Advantages and Limitations of Questionnaire Some Advantage of Questionnaire are as follows: •    Standardization: Questionnaires allow researchers to ask the same questions to all participants in a standardized manner. This helps ensure consistency in the data collected and eliminates potential bias that might arise if questions were asked differently to different participants. •    Efficiency: Questionnaires can be administered to a large number of people at once, making them an efficient way to collect data from a large sample. •    Anonymity: Participants can remain anonymous when completing a questionnaire, which may make them more likely to answer honestly and openly. •    Cost-effective: Questionnaires can be relatively inexpensive to administer compared to other research methods, such as interviews or focus groups. •    Objectivity: Because questionnaires are typically designed to collect quantitative data, they can be analyzed objectively without the influence of the researcher’s subjective interpretation. •    Flexibility: Questionnaires can be adapted to a wide range of research questions and can be used in various settings, including online surveys, mail surveys, or in-person interviews. Limitations of Questionnaire: Limitations of Questionnaire are as follows: •    Limited depth: Questionnaires are typically designed to collect quantitative data, which may not provide a complete understanding of the topic being studied. Questionnaires may miss important details and nuances that could be captured through other research methods, such as interviews or observations. •    Response bias: Participants may not always answer questions truthfully or accurately, either because they do not remember or because they want to present themselves in a particular way. This can lead to response bias, which can affect the validity and reliability of the data collected. •    Limited flexibility: While questionnaires can be adapted to a wide range of research questions, they may not be suitable for all types of research. For example, they may not be appropriate for studying complex phenomena or for exploring participants’ experiences and perceptions in-depth. •    Limited context: Questionnaires typically do not provide a rich contextual understanding of the topic being studied. They may not capture the broader social, cultural, or historical factors that may influence participants’ responses. •    Limited control: Researchers may not have control over how participants complete the questionnaire, which can lead to variations in response quality or consistency. Self-Check Exercise-1 1.    Which of the following is NOT a characteristic of a well-designed questionnaire? a)    Limited number of questions b)    Proper sequence of questions c)    Use of complex and lengthy language d)    Clear instructions 2.    What type of questionnaire includes a fixed format with predetermined questions? a)    Structured questionnaire b)    Unstructured questionnaire c)    Open-ended questionnaire d)    None of the above Fill in the blanks; (i)    Questionnaires are typically a mix of __________ and __________ questions, allowing for a range of responses. (ii)    One key advantage of questionnaires is their ability to ensure __________, asking the same questions to all participants consistently. (iii)    One limitation of questionnaires is their potential for __________ bias, where participants may not answer questions truthfully or accurately. 14.4 SCHEDULE, COMPARISON BETWEEN QUESTIONNAIRE AND SCHEDULE Like the questionnaire, schedule is technique of data collection, which contains a list of questions. The difference between a questionnaire and schedule, however, is that while the former is filled by respondents, the latter is filled by the researcher. The researcher goes to the informants with the schedule, and asks them the questions. Researcher plays an important role in the collection of data, through schedules. They explain the aims and objects of the research to the respondents and interpret the questions to them when required. Most common example of data collection through schedule is population census. The main advantage of schedule is the presence of the researcher. In simple terms, the researcher could explain the question in detail, seek additional information (i.e., information beyond the questions listed in the schedule), obtain clarification on the response, may change the sequence, language and style of questions. While framing a schedule, the researcher has to take many aspects into consideration. In fact, it is appropriate to identify the aspects on which the schedule needs to be prepared. These aspects are logically arranged and relevant questions are framed. It is likely that more than one question is asked on an aspect with the purpose of obtaining complete information. Comparison between Questionnaire and Schedule There are certain differences between questionnaire and schedule as listed below. A questionnaire is filled by the respondents, while the researcher fills the schedule. Questionnaire is more rigid in structure than schedule. Researchers have no control over response rate in case of questionnaires as many people do not respond and/or often return them without answering all the questions. On the contrary, researchers have control over the response rate of schedules since they collect data themselves. •    While questionnaire has a larger reach since it can be distributed to a large number of people at the same time, schedule has a limited reach. •    Identity of respondents is protected when data is collected using questionnaire technique while identity of informants is revealed when data is collected using schedule technique of data collection. •    The success of the questionnaire depends much on the quality of the questionnaire while the research acumen and experience of the researcher determines the success of a schedule. •    The questionnaire can be employed only when the respondents are literate while schedule can be used for data collection from both literate and illiterate informants. Possibility of obtaining incomplete and imprecise information is relatively more when data is collected through questionnaires than through schedules since the researcher is present in the field situation to verify and corroborate data there and-then. Self-Check Exercise-2 1.    Which of the following statements is true about schedules? a)    They are filled out by the respondents themselves. b)    They allow the researcher to seek additional information and obtain clarifications on responses. c)    They have a larger reach compared to questionnaires. d)    The identity of the informants is protected. 2.    One of the main advantages of using a schedule over a questionnaire is that: a)    It ensures the anonymity of the respondents. b)    It can be distributed to a large number of people simultaneously. c)    It allows the researcher to control the response rate by collecting data themselves. d)    It requires respondents to be literate. 14.5    Summary The questionnaire method presents researchers with a versatile and potent tool for data collection across diverse research domains. Its structured format enables standardized data collection, organization, and analysis, particularly beneficial for quantitative research endeavors. With advantages such as costeffectiveness, accessibility, and the capacity to reach a broad and diverse population, questionnaires offer researchers an efficient means of gathering comprehensive insights. However, it is essential to acknowledge the disadvantages associated with this method, including low response rates, potential bias due to non-response, and challenges in ensuring the representativeness   of respondents. Despite these drawbacks, the questionnaire remains a valuable instrument in the research arsenal, providing a structured approach to gathering insights and contributing to the advancement of knowledge in various fields. In contrast, schedules have limited reach, reveal informant identity, and depend on the researcher's expertise. Questionnaires require literate respondents and may yield incomplete information, while schedules can be used with both literate and illiterate informants and ensure more accurate data collection. 14.6    Glossary Survey: The process of studying a phenomenon. A questionnaire is designed to generate data that measures various attributes related to the phenomenon. Administration: The process of managing the survey, including selecting the audience, extending invitations, collecting responses, and loading data. 14.7    Answers to Self-Check Exercise 1 & 2. Self-Check Exercise-1 1.    c) Use of complex and lengthy language 2.    a) Structured questionnaire 3.    (i) close-ended, open-ended (ii)    standardization (iii)    response Self-Check Exercise-2 1.    b) They allow the researcher to seek additional information and obtain clarifications on responses. 2.    c ) It allows the researcher to control the response rate by collecting data themselves. 14.8    REFERENCES/SUGGESTED READINGS •    Izard, J. (1997). Content Analysis and Test Blueprints. Parisinternational Institute for Educational Planning. •    Mehrens, WA. and Lehmann, I.J. (1984). Measurement and evaluation in education and psychology.(3rd Ed.) New York: Holt, Rinehart and Winston. •    Nandra, I.D.S.(2011). Learning Resources and Assessment of Learning.Patiala, 21s* Century Publications. •    Ebel, Robert L.(1966) “Measuring Educational Achievement, Prentice Hall of India Pvt. Ltd. Pp. 481 14.9    TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    Explain the characteristics of a good questionnaire. 2.    What things will you keep in mind while designing a questionnaire? 3.    Discuss about advantages and disadvantages of questionnaire. 4.    Explain the main differences between questionnaires and schedules as techniques of data collection. Provide examples to illustrate these differences. ******* UNIT-15 RATING SCALE, ATTITUDE SCALE AND PERFORMANCE TEST Structure 15.1   Introduction 15.2   Learning Objectives 15.3    Rating Scale Self-Check Exercise-1 15.4    Attitude Scale Self-Check Exercise-2 15.5   Performance Test Self-Check Exercise-2 15.6  Summary 15.7    Glossary 15.8  Answers to Self-Check Exercise 15.9  References/Suggested Readings 15.10    Terminal Questions 15.1    INTRODUCTION Dear Learner, Rating scales quantify subjective judgments in research and assessment, using a continuum of values for traits, performances, or phenomena. Major approaches include paired comparison, ranking, and rating scales. Attitude scales, such as Thurstone and Likert, convert qualitative attitudes into quantitative data, offering insights in educational and social research. Performance tests require individuals to demonstrate skills through tasks, emphasizing problem-solving, critical thinking, and real-world application, making them essential for evaluating and guiding students' proficiency in educational settings. 15.2    LEARNING OBJECTIVES After studying this unit, you will be able to: •    Explain meaning and purpose of a rating scale. •    Understand precautions while constructing a rating scale. •    Discuss the uses and limitations of a rating scale. •    Describe Likert method and Thurston method of measuring attitudes. •    Explain purpose and limitations of attitude scales. 15.3    RATING SCALE Rating scale is one of the enquiry form. Form is a term applied to expression or judgment regarding some situation, object or character. Opinions are usually expressed on a scale of values. Rating techniques are devices by which such judgments may be quantified. Rating scale is a very useful device in assessing quality, especially when quality is difficult to measure objectively. For Example, —How good was the performance? It is a question which can hardly be answered objectively. Rating scales record judgment or opinions and indicates the’ degree or amount of different degrees of qualify which are arranged along a line is the scale. For example: How good was the performance? Excellent Very good Good Average Below average Poor Very poor This is the most commonly used instrument for making appraisals. It has a large variety of forms and uses. Typically, they direct attention to a number of aspects or traits of the thing to be rated and provide a scale for assigning values to each of the aspects selected. They try to measure the nature or degree of certain aspects or characteristics of a person or phenomenon through the use of a series of numbers, qualitative terms or verbal descriptions. Ratings can be obtained through one of three major approaches: •    Paired comparison •    Ranking and •    Rating scales The first attempt at rating personality characteristics was the man to man technique devised during World-war-1. This technique calls for a panel of raters to rate every individual in comparison to a standard person. This is known as the paired comparison approach. In the ranking approach every single individual in a group is compared with every other individual and to arrange the judgment in the form of a scale. In the rating scale approach which is the more common and practical method rating is based on the rating scales, a procedure which consists of assigning to each trait being rated a scale value giving a valid estimate of its status and then comparing the separate ratings into an overall score. Purpose of Rating Scale: Rating scales have been successfully utilized for measuring the following: •    Teacher Performance/Effectiveness •    Personality, anxiety, stress, emotional intelligence etc. •    School appraisal including appraisal of courses, practices and programmes. Useful Hints on Construction of Rating Scale: A rating scale includes three factors like: i) The subjects or the phenomena to be rated, ii). The continuum along which they will be rated and iii) The judges who will do the rating. All taken three factors should be carefully taken care by you when you construct the rating scale. 1) The subjects or phenomena to be rated are usually a limited number of aspects of a thing or of a traits of a person. Only the most significant aspects for the purpose of the study should be chosen. The usual may to get judgment is on five to seven point scales as we have already discussed. 2) The rating scale is always composed of two parts: i) An instruction which names the subject and defines the continuum and ii) A scale which defines the points to be used in rating. 3) Any one can serve as a rater where non-technical opinions, likes and dislikes and matters of easy observation are to be rated. But only well informed and experienced persons should be selected for rating where technical competence is required. Therefore, you should select experts in the field as rater or a person who form a sample of the population in which the scale will subsequently be applied. Pooled judgments increase the reliability of any rating scale. So employ several judges, depending on the rating situation to obtain desirable reliability. Use of Rating Scale: Rating scales are used for testing the validity of many objective instruments like paper pencil inventories of personality. They are also advantages in the following fields like: •    Helpful in writing reports to parents •    Helpful in filling out admission blanks for colleges •    Helpful in finding out student needs •    Making recommendations to employers. •    Supplementing other sources of understanding about the child •    Stimulating effect upon the individuals who are rated. Limitations of Rating Scale: The rating scales suffer from many errors and limitations like the following: As you know that the raters would not like to run down their own people by giving them low ratings. So in that case they give high ratings to almost all cases. Sometimes also the raters are included to be unduly generous in rating aspects which they had to opportunity to observe. If the raters rate in higher side due to those factors, then it is called as the generosity error of rating. The Errors of Central Tendency: Some observes wants to keep them in safe position. Therefore, they rate near the midpoint of the scale. They rate almost all as average. •    Stringency Error: Stringency error is just the opposite of generosity of error. These types of raters are very strict, cautions and hesitant in rating in average and higher side. They have a tendency to rate all individuals low. •    The Hallo Error: When a rater rates one aspect influenced by other is called hallo effect. For if a person will be rated in higher side on his achievement because of his punctually or sincerely irrespective of his perfect answer it called as hallo effect. The biasedness of the rater affects from one quality to other. •    The Logical Error: It is difficult to convey to the rater just what quality one wishes him to evaluate. An adjective or Adverb may have no universal meaning. It the terms are not properly understood by the rater and he rates, then it is called as the logical error. Therefore, brief behavioural statements having clear objectives should be used. Self-Check Exercise-1 1.    What is a rating scale? a)    A method of qualitative data collection b)    A method of quantitative data collection c)    A method of data analysis d)    A method of data visualization 2.    Rating Scale are used to record judgements about a)    oneself b)    objects c)    others d)    All of the above 15.4 ATTITUDE SCALE Attitude scale is a form of appraisal procedure and it is also one of the enquiry term. Attitude scales have been designed to measure attitude of a subject of group of subjects towards issues, institutions and group of peoples. The term attitude is defined in various ways: The behaviour which we define as attitudinal or attitude is a certain observable set II organism or relative tendency preparatory to and indicative of more complete adjustment. -L. L. Bernard An attitude may be defined as a learned emotional response set for or against something.                                               -     Barr David Johnson An attitude is spoken of as a tendency of an individual to read in a certain way towards a Phenomenon. It is y/hat a person feels or believes in. It is the inner feeling of an individual. It may be positive, negative or neutral. Opinion and attitude are used sometimes in a synonymous manner but there is a difference between two. You will be able to know when we will discuss about opinionnaire. An opinion may not lead to any kind of activity in a particular direction. But an attitude compels one to act either favourably or unfavourably according to what they perceive to bo correct. We can evaluate attitude through questionnaire. But it is ill adapted for scaling accurately the intensity of an attitude. Therefore, Attitude scale is essential as it attempts to minimise the difficulty of opinionrtaire and questionnaire by defining the attitude in terms of a single attitude object. All items, therefore, may be constructed with graduations of favour or disfavour. Purpose of Attitude Scale: In educational research, these scales are used especially for finding the attitudes of persons on different issues like: •  Co-education •    Religious education •    Corporal punishment •    Democracy in schools •    Linguistic prejudices •    International co-operation etc.' Characteristics of Attitude Scale: Attitude scale should have the following characteristics. •    It provides for quantitative measure on a uni-dimensional scale of continuum. •  It uses statements from the extreme positive to extreme negative position. •  It generally uses a five point scale as we have discussed in rating scale. •    It could be standardized and norms are worked out. •    It disguises the attitude object rather than directly asking about the attitude on the subject. Examples of Some Attitude Scale: Two popular and useful methods of measuring attitudes indirectly, commonly used for research purposes are: 1.    Thurstone Techniques of scaled values. 2.    Likert's method of summated ratings. . 1.    Thurstone Technique: Thurstone Technique is used when attitude is accepted as a uni-dimensional linear Continuum. The procedure is simple. A large number of statements of various shades of favourable and unfavourable opinion on slips of paper, which a large number of judges exercising complete detachment sort out into eleven plies ranging from the most hostile statements to the most favourable ones. The opinions are carefully worded so as to be clear and unequivocal. The judges are asked not express tier opinion but to sort them at their face value. The items which bring out a marked disagreement between the judges un assigning a position are discarded. Tabulations are made, which indicate the number of judges who placed each item in each category. The next step consists of calculating cumulated proportions for each item and ogive are constructed. Scale values of each item are read from the Ogive, the values of each item being that point along the baseline in terms of scale value units above and below which 50% of the judges placed the item. It we'll be the median of the frequency distribution in which the score ranges from 0 to 11. The respondent is to give his reaction to each statement by endorsing or rejecting it. The median values of the statements that he checks establishes his score, or quantifies his opinion. He wins a score as an average of the sum of the values of the statements he endorses. Thurstone technique is also known as the technique equal appearing intervals. 2.    The Likert Scale: The Likert scale uses items worded for or against the proposition, with five point rating response indicating the strength of the respondent's approval or disapproval of the statement. This method removes the necessity of submitting items to the judges for working out scaled values for each item. It yields scores very similar to those obtained from the Thurstone scale. It is an important over the Thurstone method. The first step is the collection of a member of statements about the subject in. question. Statements may or may not be correct but they must be representative of opinion held by a substantial number of people. They must express definite favourableness or unfavourableness to a particular point of view. The number of favourable and unfavourable statements should be approximately equal. A trial test maybe administered to a number of subjects. Only those items that correlate with the total test should be retained. The Likert’s calling techniques assigns a scale value to each of the five responses. All favourable statements are scored from maximum to minimum i. e. from a score of 5 to a score of one or 5 for strongly agree and so on 1 for strongly disagree. The negative statement or statement opposing the proposition would be scored in the opposite order. e. from a score of 1 to a score of 5 or 1 for strongly agree and so on 5 for strongly disagree. The total of these scores on all the items measures a respondent's favourableness towards the subject in question. It a scale consists of 30 items, Say, the following score values will be of interest. Limitations of Attitude Scale: In the attitude scale, the following limitations may occur; •    An individual may express socially acceptable opinion conceal his real attitude. •    An individual may not be a good judge of himself and may not be clearly aware of his real attitude. •    He may not have been controlled with a real situation to discover what his real attitude towards a specific phenomenon was. •    There is no basis for believing that the five positions indicated in the Likert's scale are equally spaced. •    It is unlikely that the statements are of equal value in‘for’and‘against’. •    It is doubtful whether equal scores obtained by several individuals would indicate equal favourableness towards again position. •    It is unlikely that respondent can validity react to a short statement on a printed form in the absence of real like qualifying Situation. •  In spite of anonymity of response, individuals tend to respond according to what they should feel rather than what they really feel. •  However, until more precise measures are developed, attitude scale remains the best device for the purpose of measuring attitudes and beliefs in social research. Self-Check Exercise-2 1)    Another name for a Likert Scale is: a)    Interview protocol b)    Event sampling c)    Summated rating scale d)    Ranking 2)    The easiest attitudinal scale, which is a summated rating scale, is the: a)    Guttman scale b)    Likert Scale c)    Thurstone scale d)    MLA scale 1 5.5 PERFORMANCE TEST It is an assessment that requires an examinee to actually perform a task or activity, rather than simply answering questions referring to specific parts.it requires students to demonstrate that they have mastered specific skills and competencies by performing or producing something. Performance test- is non-standardized test. (Non-standardized test is usually flexible in scope and format, variable in difficulty and significance) • It can be used to determine the proficiency level of students, to motivate students to study and to provide feedbacks to the students • It also allows teachers to observe achievements, habits of mind, ways of working and behavior of value in the real world. In many cases, these are outcomes that conventional tests may miss. The following characteristics should be remembered when designing a performance task: • It has various outcomes; it does not require one right answer. • It is integrative, combining different skills. • It encourages problem-solving and critical thinking skills. • It encourages divergent thinking. • It focuses on both product and process • It promotes independent learning, involving planning, revising and summation. • It builds on pupils' prior experience. • It can include opportunities for peer interaction and collaborative learning. • It enables self-assessment and reflection. • It is interesting, challenging, meaningful and authentic. • It requires time to complete. How to Design and Assess a Performance Task Step 1. List the specific skills and knowledge you wish pupils to demonstrate. Step 2. Design a performance task that requires pupils to demonstrate these skills and this knowledge. Step 3. Develop explicit performance criteria and expected performance levels measuring pupils' mastery of skills and knowledge (rubrics). Example of Performance Test •Oral presentation •Dance/movement •Science lab demonstration •Athletic competition •Dramatic reading •Enactment •Debate •Musical recital •Tableau Self-Check Exercise-3 1.    What is a primary characteristic of a performance test? a)    It focuses on answering multiple-choice questions. b)    It requires examinees to perform tasks or activities. c)    It is highly standardized with a fixed format. d)    It provides only written feedback. 2.    What is the first step in designing a performance task? a)    Develop explicit performance criteria and expected performance levels. b)    List the specific skills and knowledge you wish pupils to demonstrate. c)    Design a performance task that requires pupils to demonstrate these skills and knowledge. d)    Administer the task to a trial group of students. 15.6    SUMMARY After going through this lesson, you must have understood about the meaning of a questionnaire, its characteristics and designing process. You were also acquainted with rating scales, its purpose, uses and limitations. In the last part of the lesson, we discussed about attitude scales, its purpose, limitations and techniques of measuring attitudes. You also understood the concept of performance test and their examples. 15.7    GLOSSARY Rating Scale: A rating scale is a method that requires the rater to assign a value, sometimes numeric, to the rated object, as a measure of some rated attribute. Likert Scale: A Likert scale is a rating scale used to measure opinions, attitudes, or behaviors. Performance test: A performance test, is an educational assessment that requires students to demonstrate their knowledge, skills, and abilities through tasks, projects, or presentations. 15.8    ANSWER TO SELF-CHECK EXERCISE-1, 2 & 3. Self-Check Exercise-1 1.    b) A method of quantitative data collection 2.    d) All of the above Self-Check Exercise-2 1.    c) Summated rating scale 2.    d) MLA scale Self-Check Exercise-3 1 .     b) It requires examinees to perform tasks or activities 2 .    b) List the specific skills and knowledge you wish pupils to demonstrate. 1 5.8 REFERENCES/SUGGESTED READINGS •    Mehrens, W A. and Lehmann, I.J. (1984).Measurement and evaluation in education and psychology.(3rd Ed.) New York: Holt, Rinehart and Winston. •    Nandra, I.D.S.(2011). Learning Resources and Assessment of Learning.Patiala, 21st Century Publications. •    Taiwo, Adediran A. (2005). Fundamentals of Classroom Testing. New Delhi: Vikas Publishing House Pvt. Ltd. •    Walter W.-Cook. (1958). Educational Measurement. Washington D.C.: American Council on Education. •    Withers, G. (1997).Item writing for Tests and Examinations. Paris: International Institute for Educational Planning. 1 5.9 TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    What do you mean by rating scale? 2.    Discuss the uses and limitations of a rating scale. 3.    What do you mean by attitude scale? 4.    Describe Likert method of measuring attitudes. 5.    Write down limitations of attitude scales. 6.    Describe Performance test with examples. 156