--- title: "Vol 4 3" book: "SLM EME EDUCC 111" category: "General" publisher: "Ratan Prakashan Mandir Pvt. Ltd." type: "Educational Material" --- According to Latest Syllabus Read For Sure Success In University Examination RATAN TEXT BOOK EDUCATION MEASUREMENT AND EVALUATION Vol-4 M.A.Education (Sem-III) Dr. Nisha Tiwari Published by Ratan Prakashan Mandir Pvt. Ltd. 2nd Floor, Centre Plaza, Parinay Kunj, Lajpat Kunj Marg, Agra-282002 Copyright Authors & Publishers Revised Edition ISBN :978-93-0970-897-8 Price 85.00 only Printed at : KIDS INTERNATIONAL PVT. LTD. C-60, 61, 62, 63, EPIP, Shastripuram, Agra - 282007 Ph. : +91 9719004921 UNIT-16 ACHIEVEMENT TESTS Structure 16.1    Introduction 16.2  Learning objectives 16.3  Meaning and Characteristics of achievement test Self- check Exercise-1 16.4    Advantage and Disadvantage of achievement test Self- check Exercise-2 16.5    Types of achievement test Self- check Exercise-3 16.6    Summary 16.7    Glossary 16.8    Answer to self-check Exercise 16.9    References/Suggestive Readings 16.10    Terminal Questions 16.1    INTRODUCTION Dear learner, Achievement tests are designed to measure the knowledge, skills, and abilities that individuals have acquired in a specific area of study. They are crucial tools in educational and training settings, providing valuable information on how well students have understood and can apply what they have learned. This unit explores the meaning, characteristics, advantages, disadvantages, and types of achievement tests. 16.2    LEARNING OBJECTIVES After studying this unit, you will be able to: •    Define achievement tests. •    Describe the characteristics of achievement tests. •    Discuss the advantages and disadvantages of achievement tests. •    Identify and explain different types of achievement tests. 16.3    MEANING AND CHARACTERISTICS OF ACHIEVEMENT TESTS A test is an instrument or a tool. It follows a systematic procedure for measuring a sample of behaviour by posing a set of questions in a uniform manner. It is an attempt to measure what a person knows or can do at a particular point in time. Furthermore, a test answers the question ‘how well’ does the individual perform either in comparison with others or in comparison with a domain of performance tasks? A test designed to appraise what the individual has learned as a result of planned previous experience or training is an Achievement Test. Since it relates to what has been learnt already its frame of reference is on the present or past. Achievement tests attempt to measure what a person knows or can do at a particular point in time. Furthermore, our reference is usually to the past; that is we are interested in what has been learned as a result of a particular course or experience or a series of experiences. An achievement test is designed to evaluate a person's knowledge, skills, and proficiency in a specific area of study or subject matter. It aims to measure how well individuals have learned and can apply what they have been taught. Achievement tests are used teachers to measure or test the achievements and success achieved in any particular field by the students. Whatever the student learns in school is called his achievement and examinations conducted to test that achievement are called achievement tests. Characteristics of Achievement tests Achievement tests have several key characteristics that distinguish them from other types of assessments: 1.    Standardization: Achievement tests are standardized, meaning they are administered and scored in a consistent manner to ensure fairness and comparability of results. 2.    Objective Measurement: These tests aim to provide an objective measurement of students' knowledge and skills, reducing the influence of subjective factors. 3.    Content Specificity: The content of achievement tests is aligned with the curriculum or training program, focusing on specific knowledge areas and skills. 4.    Reliability: Achievement tests are designed to be reliable, providing consistent results over repeated administrations. 5.    Validity: These tests measure what they are intended to measure, ensuring that the test content accurately reflects the learning objectives. Self-Check Exercise –1 1.    What is an achievement test designed to measure? a)    Future potential b)    Innate intelligence c)    Past learning and proficiency d)    Social skills 2.    Which of the following is NOT a characteristic of achievement tests? a)    Standardization b)    Subjectivity c)    Content specificity d)    Reliability 16.4   ADVANTAGES   AND   DISADVANTAGES   OFACHIEVEMENT     TESTS Achievement tests offer several advantages in educational and professional settings: 1.    Objective Assessment: Provides an objective evaluation of students' knowledge and skills, reducing biases in assessment. 2.    Standardized Comparison: Enables comparison of performance across different groups of students or individuals, providing benchmarks for achievement. 3.    Feedback for Improvement: Offers valuable feedback to students, teachers, and administrators on areas of strength and areas needing improvement. 4.    Motivation: Encourages students to study and perform well, as they know their achievements will be measured. 5.    Accountability: Holds educators and institutions accountable for students' learning outcomes. Disadvantages of Achievement tests Despite their advantages, achievement tests also have some limitations: 1.    Limited Scope: May not capture all aspects of a student's abilities or knowledge, particularly skills that are difficult to measure objectively. 2.    Test Anxiety: Can cause anxiety and stress among students, which may affect their performance. 3.    Teaching to the Test: May encourage teaching practices that focus primarily on test content, neglecting broader educational goals. 4.    Cultural Bias: Standardized tests may contain cultural biases that disadvantage certain groups of students. 5.    Resource Intensive:  Developing, administering, and scoring achievement tests can be resource-intensive in terms of time and cost. Self-Check Exercise-2 1.    How do achievement tests benefit educators and institutions? a)    By increasing the overall number of tests administered annually. b)    By offering financial rewards for high test scores. c)    By holding educators and institutions accountable for students' learning outcomes. d)    By reducing the need for curriculum development. 2.    What is a major disadvantage of achievement tests? a)    Objective assessment b)    Test anxiety c)    Feedback for improvement d)    Standardized comparison 16.5 TYPES OF ACHIEVEMENT TESTS There are several types of achievement tests, each serving different purposes and assessing different aspects of learning: 1.    Diagnostic Tests:  Used to identify students' strengths and weaknesses in specific areas before instruction begins. They help teachers plan targeted interventions and support. 2.    Formative Tests: Administered during the instructional process to monitor students' progress and provide ongoing feedback. They help teachers adjust their teaching strategies to improve learning outcomes. 3.    Summative Tests: Given at the end of an instructional period to evaluate students' overall learning and achievement of course objectives. Examples include final exams and standardized tests. 4.    Criterion-Referenced Tests: Measure students' performance against a fixed set of criteria or learning standards. They determine whether students have mastered specific skills or knowledge areas. 5.    Norm-Referenced Tests: Compare a student's performance to that of a larger group (norm group). These tests rank students and provide information on how they perform relative to others. 6.    Performance-Based Tests: Require students to perform tasks or produce work that demonstrates their knowledge and skills. Examples include projects, presentations, and portfolios. Self-Check Exercise-3 1.    Which type of test is used to identify students' strengths and weaknesses before instruction begins? a)   Summative Tests b)   Formative Tests c)   Diagnostic Tests d)   Performance-Based Tests 2.    Which of the following tests is designed to measure students' performance against a fixed set of criteria or learning standards? a) b) c) d) Norm-Referenced Tests Criterion-Referenced Tests Diagnostic Tests Formative Tests 16.6    SUMMARY Achievement tests are standardized assessments designed to measure the knowledge, skills, and proficiency individuals have acquired in specific areas of study. They are characterized by their standardization, objectivity, content specificity, reliability, and validity. While they offer several advantages, such as objective assessment and feedback for improvement, they also have limitations, including limited scope and potential for test anxiety. Understanding the different types of achievement tests helps in selecting the appropriate test for specific assessment needs. 16.7    GLOSSARY •    Achievement Test: A standardized test designed to measure the knowledge, skills, and proficiency that an individual has acquired in a specific subject area. •    Standardization: The process of administering and scoring a test in a consistent manner to ensure fairness and comparability of results. •  Reliability: The extent to which a test provides consistent results over repeated administrations. •  Validity: The degree to which a test measures what it is intended to measure. 16.8    ANSWERS TO SELF-CHECK EXERCISE Self-Check Exercise-1 1.    c) Past learning and proficiency 2.    b) Subjectivity Self-Check Exercise-2 1.    c) By holding educators and institutions accountable for students' learning outcomes. 2.    b) Test anxiety Self-Check Exercise-3 1.    c) Diagnostic Tests 2.    b) Criterion-Referenced Tests 16.9 REFERENCES/SUGGESTIVE READINGS •    Aggarwal, J. C. (2007). Educational Measurement and Evaluation. Delhi: Vikas Publishing House. •    Dandekar, W. N. (2004). Measurement, Evaluation, and Assessment in Education. Pune: Vidya Prakashan. •  Verma, J. P. (2013). Essentials of Educational Measurement. New Delhi: Sage Publications. •  Kamte, V. (2015). Educational Testing and Measurement. Mumbai: Himalaya Publishing House. 16.9 TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    Define achievement tests and explain their significance in educational settings. 2.    Discuss the key characteristics that distinguish achievement tests from other types of assessments. 3.    Describe the advantages and disadvantages of using achievement tests. 4.    Identify and explain different types of achievement tests and their purposes. ******* UNIT-17 ACHIEVEMENT TEST CONSTRUCTION Structure 17.1    Introduction 17.2    Learning Objectives 17.3    Steps, Planning and Preparation of an Achievement Test Self- check Exercise-1 17.4    Preparation of the Test Blueprint and Writing of Test Items Self- check Exercise-2 17.5    Assembling and arranging items in the test Self- check Exercise -3 17.6    Writing Instructions or guidelines for Test Administration and Scoring Self- check Exercise-4 17.7    Performing Item Analysis Self- check Exercise-5 17.8    Summary 17.9    Glossary 17.10    Answers to Self-check Exercise 17.11    References/Suggestive Readings 17.12    Terminal Questions 17.1    INTRODUCTION Dear learner, Test construction is based upon practical and scientific rules that are applied before, during, and after each item until it finally becomes a part of the test. Construction of tests is an important part of assessing students' understanding of course content and their level of competency in applying what they are learning. In this lesson, we will learn about the procedure and principles of achievement test construction. We will also learn about item analysis and desirable attributes of a good achievement test. 17.2    LEARNING OBJECTIVES After studying this unit, you will be able to: •    Discuss the steps of constructing an achievement test. •    Prepare a blueprint of an achievement test. •    Perform item analysis to evaluate test items. •    Write clear instructions for test administration and scoring. •    Identify desirable attributes of a good achievement test. 17.3    STEPS, PLANNING AND PREPARATION OF ANACHIEVEMENT TEST Any test designed to assess the achievement in any subject with regard to a set of predetermined objectives involves several major steps: •    Planning and Preparation of a Design for the Test •    Determine the objective of the test. •    Determine the maximum time and maximum marks. •    Test Length: The number of items that should constitute the final form of a test is determined by the purpose of the test or its proposed uses, and by the statistical characteristics of the items. •    Preparation of a Design for the Test Important factors to be considered in design for the test are as follows. >    Weightage to objectives >    Weightage to content >    Weightage to Type of questions >    Weightage to difficulty level. ■    Weightage to objectives Here's the data presented in a table format:- | Sr. No. | Objectives | Marks | Percentage | |---|---|---|---| | 1. | Knowledge | 6 | 24 | | 2. | Understanding | 8 | 32 | | 3. | Application | 11 | 44 | | Total | | 25 | 100 | ■    Weightage to content | Sr. NO. | Contents | Marks | Percentage | |---|---|---|---| | 1. | Sub topic | 10 | 40 | | 2. | Sub topic | 15 | 60 | | Total | | 25 | 100 | ■    Weightage to Type of questions | Sr.no. | Types    of questions | No.       of questions | Marks | Percentage | |---|---|---|---|---| | 1. | Objective type | 13 | 13 | 52 | | 2. | Short   answer type | 1 | 2 | 8 | |---|---|---|---|---| | 3. | Essay type | 1 | 10 | 40 | | 4. | Total | 15 | 25 | 100 | Weightage to difficulty level | Sr. No. | Forms       of questions | Marks | Percentage | |---|---|---|---| | 1. | Easy | 5 | 20 | | 2. | Average | 15 | 60 | | 3. | Difficult | 5 | 20 | | Total | | 25 | 100 | Self-Check Exercise-1 1.    Which objective is given the highest percentage of marks? a)    Knowledge b)    Understanding c)    Application d)    Analysis 2.    What percentage of the total marks is allocated to "Average" difficulty level questions? a)    20% b)    60% c)    40% d)    52% 17.4    PREPARATION OF THE TEST BLUEPRINT AND WRITING OF TEST ITEMS Test Blueprint: A test blueprint is a crucial tool in the test development process. It serves as a detailed plan or a three-dimensional chart that outlines the structure and content of an assessment. The blueprint ensures that the test is comprehensive and aligned with the learning objectives, content areas, and question formats intended to be covered. Here’s a detailed breakdown of the process and components involved in preparing a test blueprint and writing test items: | Objectives Form    of Questions Content | Knowledge | Understanding | Application | Grand Total | |---|---|---|---|---| | O | SA | E | O | SA | E | O | SA | O | | Sub- Topic -1 | 1(3) | | | 1(6) | | | 1(1) | | | 10 | | Sub- Topic-2 | 1(3) | | | | 2(1) | | | 10(1) | | | | Total Marks | 6 | 0 | 0 | 5 | 3 | 0 | 2 | | 10 | | | Grand Total | 6 | 8 | 11 | 25 | Table of Specifications: A table of specifications is a two-way table that represents along one axis the content area/topics that the teacher has taught during the specified period and the cognitive level at which it is to be measured, along the other axis. In other words, the table of specifications highlights how much emphasis is to be given to each objective or topic. After preparation of test blue print, the items are written. The details regarding writing of test items are given in next lesson. Assembling and Arranging Items in the Test: Items after having written and selected they are organized in the form of a test. Item of the same format may be placed together. Each item type requires specific set of directions and a somewhat different mental set on the part of the examinee. So far as possible, within item type, items dealing with the same content may be grouped together. The examinee will be able to concentrate on a single domain at a time rather than having to shift back and forth among areas of content. Furthermore, the examiner will have an easier job of analysing the results, as it will be easier to see at a glance whether the errors are more frequent in one content area than the other. Items may be so arranged that difficulty progress from easy to hard Items should be arranged in the test booklet so that answers follow no set pattern. Self-Check Exercise-2 1.    What is the primary purpose of a test blueprint in the test development process? a)    To increase the difficulty of the test b)    To outline the structure and content of an assessment c)    To reduce the number of questions in a test d)    To provide financial incentives to teachers 2.    In the test blueprint provided, how many marks are allocated to the "Knowledge" objective for "Sub-Topic 1"? a) b) c) d) 3 6 5 10 17.5    ASSEMBLING AND ARRANGING ITEMS IN THE TEST Items after having written and selected they are organized in the form of a test. Items of the same format may be placed together. Each item type requires specific set of directions and a somewhat different mental set on the part of the examinee. So far as possible, within item type, items dealing with the same content may be grouped together. The examinee will be able to concentrate on a single domain at a time rather than having to shift back and forth among areas of content. Furthermore, the examiner will have an easier job of analyzing the results, as it will be easier to see at a glance whether the errors are more frequent in one content area than the other. Items may be so arranged that difficulty progress from easy to hard Items should be arranged in the test booklet so that answers follow no set pattern. Self-Check Exercise-3 1.    Items should be arranged in the test booklet so that difficulty progresses from __________ to __________, and answers follow no set pattern. Answer: easy, hard 17.6    WRITING INSTRUCTIONS OR GUIDELINES FOR TEST ADMINISTRATION AND SCORING The directions should be simple but complete. They should indicate the purpose of the test, the time limits and the score value of each question. Write a set of directions for each item type that is used on the test specifying what the respondent is expected to do and how one is required to record the responses. All pupils must be given a fair chance to demonstrate their achievement. Physical and psychological environment be conducive to their best efforts. Control all factors that might interfere with valid measurement: Adequate workspace, quiet, proper light and ventilation are important. Pupils must be put at ease, tension and anxiety should be reduced to the minimum. Separate answer sheets, which are easier to score, can be used at high school level and beyond. If the pupils’ answers are recorded on the test paper, the teacher may make a scoring key by marking the correct answers on a blank copy of the test. When separate answer sheets are used, a scoring stencil is a blank answer sheet with holes punched where correct answer should appear. Before scoring procedure is used, each test paper should also be scanned to make sure that only one answer was marked for each item. Any item containing more than one answer should be eliminated from scoring. In scoring objective tests, each correct answer is usually counted as one point. When pupils are told to answer every item on the test, a pupil's score is simply the number of items answered correctly. Short answer questions may sometime require awarding partial credit and may pose some problem in scoring. However, a detailed key may be prepared in advance to avoid confusion. For each question and for the test as a whole, the examiner may make a tally for each kind error that the examinees make. A summary of these errors could then be used to plan instructional activities. Self-Check Exercise-4 1.    The directions for a test should indicate the __________ of the test, the time limits, and the score value of each question. 2.    When separate answer sheets are used, a scoring __________ is a blank answer sheet with holes punched where the correct answers should appear. 17.7    PERFORMING ITEM ANALYSIS Often students judge, after taking the exam, whether the test was fair and good. Teacher is also usually interested about how the test worked for the students. One way to ascertain this is to undertake item analysis. It provides objective, external and empirical evidence for the quality of the items we have pre-tested. The objective of item analysis is to identify problematic or poor items which might be either confusing the respondents or do not have a clearly correct response or a distracter might well be competing with the keyed answer. Good test making requires careful attention to the principles of item evaluation. The basic methods involve are assessment of item difficulty and item discrimination. These measures comprise item analysis. Item analysis is about how difficult an item is and how well it can discriminate between the good and the poor students. (i)    Item Difficulty Index/Facility Index Item difficulty is determined from the proportion (p) of students who answered each item correctly. Item difficulty can range from zero (none could solve it) to hundred (all persons solved it correctly). The goal is usually to have items of all difficulty levels in the test so that test could identify poor, average as well as good students. However, most of the items are designed to be average in difficulty levels for they are more useful. Item analysis exercise provides us the difficulty level of each item. •    Optimally difficult items are those that 50%-75% of students answer correctly. •    Items are considered low to moderately difficult if (p) is between 70% and 85% •    Items that only 30% or below solve correctly are considered difficult ones. Item Difficulty Percentage can also be denoted as Item Difficulty Index by expressing it in decimals e.g. .40 for items which could be solved by 40 % of the test-takers. Thus index can range from 0 to 1. Items should fall in a variety of difficulty levels in order to differentiate between good and average as well as average and poor students. Easy items are usually placed in the initial part of the test to motivate students in taking the test and alleviating test-anxiety. The optimal item difficulty depends on the question type and number of possible distracters as well. (ii)    Item Discrimination Another way to evaluate items is to ask “Who gets this item correct”- the good, average and the weak students? Assessment of item discrimination answers this query. Item discrimination refers to the percentage difference in correct responses between the poor and the high scoring students. The discrimination index is a basic measure of the validity of an item. It is a measure of an items ability to discriminate between those who scored high on the total test and those who scored low. Though there are several steps in its calculation, once computed, this index can be interpreted as an indication of the extent to which overall knowledge of the content area or mastery of the skills is related to the response on an item. Perhaps the most crucial validity standard for a test item is that whether a student got an item correct or not is due to their level of knowledge or ability and not due to something else such as chance or test bias. In a small class of 30 students, one can administer the test items, score them and then rank. Next, we separate the upper 15 students and the-low 15 into two groups: The UPPER and the LOW groups. Finally, we find how well each item was solved correctly (p) by each group. In other words, percentage of students passing (p) each item in each of the two groups is worked out. Discrimination (D) power of the item is then known by finding difference between the percentage of upper group and the low group. The higher the difference, the greater the discrimination power of an item. D = (p of upper group - p of lower group) In a large class of 100 or more students, we take the top 25% and the lower 25% students to form upper and lower groups, to cut short the labor or amount of work. The discrimination ratio for an item falls between -1.0 and +1.0. The closer the ratio is to +1.0, the more effectively that item distinguishes students who know the material (the top group) from those who don’t (the bottom group). An item with a discrimination of 60% or greater is considered a very good item, whereas a discrimination of less than 20% indicates a low discrimination and the item needs to be revised. An item with a negative index of discrimination indicates that the poor students answer correctly more often than do the good students. Strange! Such items should be dropped from the test. For example, ten students in a class have taken a ten items quiz. The students’ responses are shown below from high to low. The top five students can be called the high score group and the bottom half as the low scoring group. The number “1” indicates a correct answer; a “0’’ indicates an incorrect answer. | Student | Total score % | Items No. | |---|---|---| | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | | 1. | 100 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | | 2. | 90 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 0 | 1 | | 3. | 80 | 1 | 1 | 0 | 1 | 1 | 1 | 1 | 1 | 0 | 0 | | 4. | 70 | 0 | 1 | 1 | 1 | 1 | 1 | 0 | 1 | 0 | 1 | | 5. | 70 | 1 | 1 | 1 | 0 | 1 | 1 | 1 | 0 | 0 | 1 | | 6. | 60 | 1 | 1 | 1 | 0 | 1 | 1 | 0 | 1 | 0 | 0 | | 7. | 60 | 0 | 1 | 1 | 0 | 1 | 1 | 0 | 1 | 0 | 1 | | 8. | 50 | 0 | 1 | 1 | 1 | 0 | 0 | 1 | 0 | 1 | 0 | | 9. | 40 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 1 | 1 | | 10. | 30 | 0 | 1 | 0 | 0 | 0 | 1 | 0 | 0 | 1 | 0 | Difficulty index and Discrimination Index are calculated below: | Items No. | Correct  High Group | Low   Correct Group | Difficulty % | Discrimination % | |---|---|---|---|---| | 1 | 4 | 2 | 60 | 40 | | 2 | 5 | 5 | 100 | 0 | | 3 | 4 | 4 | 80 | 0 | | 4 | 4 | 1 | 50 | 60 | | 5 | 5 | 2 | 80 | 60 | | 6 | 5 | 3 | 80 | 40 | | 7 | 4 | 1 | 50 | 60 | | 8 | 4 | 2 | 60 | 40 | | 9 | 1 | 3 | 30 | 40 | | 10 | 4 | 2 | 60 | 40 | • Question no 2 was the easiest; no 9 was most difficult. •    Question 9 also had negative discrimination and should be removed from the test. •    100% discrimination would occur if all those in the upper group answered correctly and all those in the lower group answered incorrectly. •    Zero discrimination occurs when equal numbers in both groups answer correctly. •    Negative discrimination, a highly undesirable condition, occurs when more students in the lower group than the upper group answer correctly. •    Items with 25% and above discrimination are considered good. (iii)    Analysis of Response Options (Distracter Analysis): In addition to examining the performance of an entire test item, teachers are often interested in examining the performance of individual distracters (incorrect answer options) on multiple-choice items. By calculating the proportion of students who chose each answer option, teachers can identify which distracters are “working” and appear attractive to students who do not know the correct answer, and which distracters are simply taking up space and not being chosen by many students. To eliminate blind guessing which results in a correct answer purely by chance (which hurts the validity of a test item), teachers want as many plausible distracters as is feasible. Analyses of response options allow teachers to fine tune and improve items they may wish to use again with future classes. Interpreting Distracter Values: Distracters should be ideally equally attractive, but not more than the answer. Minimum, it must be opted by at least 5% of the examinees. Weak or nonfunctional distracters may be substituted with new ones and make sure that they align with the stem as well as the objective of the item, well connected with the rest, and are grammatically correct. Effectiveness of Distracters: Difficulty and discrimination index are estimates about an item which overall comprises a stem and a set of distracters or options. The item analysis statistics reflects on the goodness of both distracters and the stem. Let us look at the guidelines which can help us improve them. 1.    Most MCQs have 2-4 distracters; 3 is better, 4 is best at the college level Where it is difficult to think of more than one distracter, frame it as true/false item 2.    Distracters that have less than 5 percent response rate are weak and may be changed / improved. Distracters which attracted no response are not working at all. 3.    No distracter should be chosen more than the keyed response in the upper group. 4.    Similarly, no one distracter should pull more than about half the students. 5.    If students have respond about equally to all the options, they might be marking randomly or wildly guessing. Critically check contents of such items. They might have been written badly and the students seem to have no idea what you are asking. It could be very difficult items and students might be completely baffled. 6.    If the low group gets the keyed answer as often as the upper group, all the distracters might be looked into again. Or drop the item if you have a large pool of items. Self-Check Exercise-5 1.    What is the primary objective of item analysis in test development? a)    To increase the number of questions b)    To identify problematic or poor items c)    To reduce the test duration d)    To improve student attendance 2.    What should teachers do with distracters that attract less than 5% of the examinees? a)     Keep them as they are. b)    Substitute them with new ones. c)    Increase their difficulty. d)    Eliminate the correct answer. 17.8    SUMMARY Test construction involves several key steps to ensure the reliability and validity of an achievement test. These steps include planning the test by determining objectives, time, and length; designing the test by assigning weightage to objectives, content, question types, and difficulty levels; preparing a test blueprint and writing test items; and arranging the test items logically. Additionally, clear instructions for test administration and scoring are essential. Item analysis, including evaluating item difficulty and discrimination, helps improve test quality by identifying problematic items and ensuring the test distinguishes between different levels of student performance. Analysis of response options also helps refine multiple-choice questions by examining the effectiveness of distracters. 17.9    GLOSSARY Blueprint :- A test blueprint is a rubric, document, or a table that lists the learning outcomes to be tested , the level of complexity, and the weight for the learning outcomes in rubric. Item Analysis: A process used to evaluate the effectiveness of individual test items by measuring their difficulty and discrimination indices. Item Difficulty Index: A measure that indicates the proportion of students who answered a test item correctly, ranging from 0 to 1. Item Discrimination: A measure of how well a test item differentiates between students who perform well overall and those who perform poorly. 17.10    ANSWER TO SELF-CHECK EXERCISE Self-Check Exercise-1 1.    c) Application 2.    b) 60% Self-Check Exercise-2 1.    b) To outline the structure and content of an assessment 2.    a) 3 Self-Check Exercise-3 1.    easy, hard Self-Check Exercise-4 1.    purpose 2.    stencil Self-Check Exercise-5 1.    b) To identify problematic or poor items 2.    b) Substitute them with new ones. 17.11    REFERENCES/SUGGESTIVE READINGS •    Education Measurement and Evaluation : J.Swarupa Rani, Discovery Publishing house. •    Measurement and Evaluation in Teaching : Norman Edward Gronlund Macmillian •    Measurement and Assessment in Teaching : Robert L. Linn Pearson Education India. •    Program Evaluation and performance measurement : James C. Me. David, Laura 17.12 TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    Discuss the major steps involved in constructing an achievement test. 2.    Explain the significance of preparing a test blueprint in the test construction process. 3.    What is item analysis, and why is it important in test construction? 4.    How does item discrimination contribute to the validity of a test item? ******* UNIT-18 CONSTRUCTION OF NORM-REFERENCED TEST Structure 18.1    Introduction 18.2    Learning Objectives 18.3    Purpose and Characteristics of Norm-Referenced Tests Self-check Exercise-1 18.4    Constructing Norm-Referenced Tests Self-check Exercise-2 18.5    Summary 18.6    Glossary 18.7    Answers to self-check Exercise 18.8    References/Suggestive Reading 18.9    Terminal Questions 18.1    INTRODUCTION Dear Learner, Norm-referenced tests (NRTs) are designed to compare a student's performance to that of a group. These tests play a crucial role in educational assessment by identifying the relative standing of students. This unit will delve into the process of constructing NRTs, ensuring that participants gain a thorough understanding of each step involved. 18.2    LEARNING OBJECTIVES After studying this unit, you will be able to: •    Explain the purpose and importance of norm-referenced tests. •    Outline the steps involved in constructing norm-referenced tests. •    Develop test items that align with curriculum content. •    Conduct and analyze item trials. •    Assemble and finalize a norm-referenced test 18.3    PURPOSE AND CHARACTERISTICS OF NORM-REFERENCED TESTS The primary purpose of NRTs is to rank students and compare their performance to a norm group. This helps in: •    Identifying high and low achievers. •    Placing students in appropriate educational tracks or programs. •    Informing instruction and curriculum development. •    Evaluating the effectiveness of educational programs. Characteristics of Norm-Referenced Tests 1.    Comparative Evaluation: NRTs rank students based on their performance relative to a norm group. 2.    Standardization: Tests are administered and scored in a consistent manner to ensure comparability. 3.    Statistical Analysis: Data from NRTs are analyzed using statistical methods to establish norms and interpret scores. Self-check Exercise-1 1.    Which of the following is NOT a characteristic of Norm-Referenced Tests (NRTs)? a)    Comparative Evaluation b)    Standardization c)    Formative Assessment d)    Statistical Analysis 2.    How do Norm-Referenced Tests (NRTs) assist in curriculum development? a)    By providing detailed individual feedback for each student b)    By ranking students and comparing their performance to identify high and low achievers c)    By offering personalized learning paths for each student d)    By assessing students' growth over the academic year 18.4    CONSTRUCTING NORM-REFERENCED TESTS The steps for constructing norm-referenced tests are briefly discussed below: 1.    Content analysis and test blueprints A content analysis provides a summary of the intentions of the curriculum expressed in content terms. Which content is supposed to be covered in the     curriculum? Are there significant sections of this content? Are there significant sub-divisions within any of the sections? Which of these content areas should a representative test include? 2.    Item writing Item writing is the preparation of assessment tasks which can reveal the knowledge and skill of students when their responses to these tasks are inspected. Tasks which confuse, which do not engage the students, or which offend, always obscure important evidence by either failing to gather appropriate information or by distracting the student from the intended task. 3.    Item review Writing assessment tasks for use in tests requires skill. Sometimes the item seems clear to the person who wrote it but may not necessarily be clear to others. Before empirical trial, assessment tasks need to be reviewed by a review panel (with a number of people) with questions like: > Is the task clear in each item? Is it likely that the person attempting an item will know what is expected? > Are the items expressed in the simplest possible language? > Is each item a fair item for assessment at this level of education? > Is the wording appropriate to the level of education where the item will be used? > Are there unintended clues to the correct answer? > Is the format reasonably consistent so that students know what is required from item to item? > Is there a single clearly correct (or best) answer for each item? > Is the type of item appropriate to the information required? > Are there statements in the items which are likely to offend? > Is there content which reflects bias on gender, racial, or other grounds? > Are the items representative of the behaviours to be assessed? > Are there enough representative items to provide an adequate sample of the behaviours to be assessed? This review before the items are tried should ensure that we avoid tasks which are expressed in language too complex for the idea being tested, avoid redundant words, multiple negatives, and distracters which are not plausible. The review should also identify items with no correct (or best) answer and items with multiple correct answers. Such items may be discarded or rewritten. 4.    Trial of the Items Item trial is sometimes called pilot testing - but in this context it does not mean testing those who fly aeroplane. As well as considering the best efforts of item writers and item reviewers as a means of eliminating faulty items and improving the quality of items, it is necessary to subject the proposed items to empirical trial with students similar to those who are going to use the final form of the test. It is usual to allocate the trial forms on a random basis within each trial examination room so that (on the average) each trial test is attempted by candidates of comparable ability. The same form of a test should not be given to candidates sitting in adjacent seats so as to ensure that candidates do not improve their scores by looking at another candidate’s paper. 5.    Processing Test Responses after Trial Testing If the test needs to be scored before analysis; this scoring is done next. If there are essays to be scored, it is good practice to mark the first essay all the way through the stack of test papers. Then start the stack again to score the next essay. When all items have been marked, the scores on each item are entered into a computer file. If the test is multiple-choice in format, the responses may be entered into a computer file directly. 6.    Item Analysis Empirical trial can identify instances of confused meaning, alternative explanations not already considered by the test constructors, and (for multiple-choice questions) options which are popular amongst those lacking knowledge, and incorrect options which are chosen for some reason by very able students. The item analysis also provides an opportunity to collect information about how each item performs relative to other items in the same test, and to judge the consistency of the whole test. 7.    Amending the Test by Discarding/Revising/Replacing items Items which do not perform as expected can be discarded or revised. However, discarding questions when there is a shortage of replacement questions can lead to distortions of the achieved test specification. If the original specification represents the best sampling of content, skills, and item formats, in the judgments of those preparing and reviewing the test, then leaving some cells of the grid vacant will indicate a less than adequate test To avoid this possibility, test constructors may prepare three or four times as many questions that they think they will need for each cell in the grid. 8.    Assembling the Final Test (or a further trial test) and the Corresponding Score Key After trial, tasks may be re-ordered to take account of their difficulty. Usually the easiest questions are presented first. This is to encourage candidates to proceed through the test and to ensure that the weaker candidates do not become discouraged before providing adequate evidence of their achievements and skills. Minor changes to items may have to be made for layout reasons (for example, to keep all of an item on one page of the test, or to avoid obvious patterns in the list of correct answers). 9.    Other Practical Concerns in Preparing the Test y How much time will students have to do the actual test? What time will be set aside to give instructions to those students attempting the test? Will the final number of items be too large for the test to be given in a single session? Will there be a break between testing sessions when there is more than one session? •/ Will the students be told how the items are to be scored? Will they be told the relative importance of each item? Will they be given advice on how to do their best on the test? •/ What test administration information will be given to those who are giving the trial test to students? Will the students be told that the results will be returned to them? Are the tests to be treated as secure tests (with no copies left behind in the venue where the test is administered)? •/ Do students need advice on how they are to record their responses? If practice items are to be used for this purpose, what types of response should they cover? How many practice items will be necessary? •/ Will the answers be recorded on a separate answer sheet (perhaps so that a test booklet can be used again)? Will this use of a separate sheet add to the time given for the trial test? What information should be requested in addition to the actual responses to the items? (This might include student name, school, year level, sex, age, etc.) •/ Has the layout of the test (and answer sheet if appropriate) been arranged for efficient scoring of responses? Are distracters for multiple-choice tests shown as capital letters (easier to score than lower case letters)? •/ Have the options in multiple-choice items been arranged in some logical order (for example, from smallest to largest)? Have the items been placed in order from easiest to most difficult (to encourage candidates to continue through the test)? Has the layout of items avoided patterns in the correct answers such as 3 or more of the same letter in a row, or other patterns like ABCD or ABABAB (which might lead to ‘correct’ responses for the ‘wrong’ reasons)? Developing Norms for Interpretation of Test Scores Norm is average score of sample population. These are the level obtained by a particular group of persons on a test. There are many types of norms like age norms, grade norms, percentile norms and standard scores. Self-check Exercise-2 1.    Why is it important to conduct an empirical trial of test items? a)    To ensure the test is administered consistently b)    To identify instances of confused meaning and alternative explanations c)    To summarize the curriculum content d)    To prepare the final test score key 2.    In preparing the final test, why are the easiest questions usually presented first? a)    To prevent students from becoming discouraged early in the test b)    To ensure that the test is scored efficiently c)    To avoid patterns in the list of correct answers d)    To cover the most important content areas first 18.5    SUMMARY Norm-referenced tests (NRTs) compare a student's performance against a predefined group, highlighting their relative standing. These tests follow a systematic process: conducting content analysis, item writing, and review; piloting items; scoring and analyzing responses; and refining the test based on item performance. Norms, such as age or grade norms, are established using data from a representative sample to interpret scores. While NRTs are useful for ranking students and informing educational decisions, they can encourage teaching to the test and may not reflect individual progress. Ensuring fairness, validity, and security is crucial in constructing effective NRTs. 18.6    GLOSSARY •    Norm-Referenced Test (NRT): A test designed to compare a student's performance to that of a group. •    Content Analysis: The process of summarizing curriculum intentions in content terms. •    Norms: Average scores of a sample population, used for interpreting test scores. 18.7    ANSWERS TO SELF-CHECK EXERCISE Self-check Exercise-1 1.    c) Formative Assessment 2.    b) By ranking students and comparing their performance to identify high and low achievers Self-check Exercise-2 1.    b) To identify instances of confused meaning and alternative explanations 2.    a) To prevent students from becoming discouraged early in the test 18.8    REFERENCES/SUGGESTIVE READINGS •    Aggarwal, J. C. (2019). Essentials of Educational Evaluation. Vikas Publishing House. •    Sachdeva, K. (2014). Educational Measurement and Evaluation. Sterling Publishers Pvt. Ltd. •    Mishra, P. (2017). Educational Measurement and Evaluation. Himalaya Publishing House 18.9    TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    Explain the process of constructing a norm-referenced test from content analysis to final test assembly. 2.    Describe the role of item analysis in the development of norm-referenced tests. What statistical methods are used, and how do they contribute to the overall quality of the test? ******* UNIT-19 WRITING TEST ITEMS-1 Structure 19.1   Introduction 19.2  Learning Objectives 19.3    Writing Test Items 19.4  Constructing Objective Type Test Items 19.5  Alternative Response Type Items 19.6    Short Answer/ Completion Type Items 19.7    Summary 19.8    Glossary 19.9    Answers to Self-Check Exercise 19.10    References/Suggestive Readings 19.11    Terminal Questions 19.1    INTRODUCTION Dear Learner, Educational assessments play a pivotal role in evaluating students' understanding and mastery of learning objectives. At the heart of any assessment are the test items—questions or prompts designed to gauge students' knowledge, skills, and abilities. Effective test item construction is not merely about drafting questions but ensuring that these questions are valid, reliable, and aligned with educational goals. This chapter delves into the principles and practices of constructing test items across various formats, such as multiple-choice questions, true-false statements, and short-answer items. Each type of item serves a distinct purpose in assessing different cognitive levels—from recalling facts to analyzing complex scenarios. By understanding the nuances of item construction, educators can design assessments that accurately measure student achievement and inform instructional decisions. 19.2    LEARNING OBJECTIVES After studying this unit, you will be able to: •    Understand the principles of constructing different types of test items. •    Identify strategies to enhance the validity and reliability of test items. •    Demonstrate the ability to develop test items aligned with learning objectives. •    Evaluate the appropriateness of test items for assessing different cognitive levels. 19.3    WRITING TEST ITEMS The next step after planning the test is preparing it in accordance with the plan. This step mainly deals with development of items and organizing them in the form of a test. The initially developed test draft or pool of items is termed as preliminary draft or rough draft of the test. Different types of questions can be devised for an achievement test, for instance, multiple choice, fill-in-theblank, true-false, matching, short answer and essay. Although each type of question is constructed differently, the following principles apply to constructing questions and tests in general: 1.    Instructions for each type of question must be simple and brief. 2.    Questions must be written in simple language. If the language is difficult or ambiguous, even a student with strong language skills and good vocabulary may answer incorrectly if his/her interpretation of the question is different from the author’s intended meaning. 3.    Test items must assess specific ability or comprehension of content developed during the course of study. 4.    Write the questions as you teach or even before you teach, so that your teaching may be aimed at significant learning outcomes. 5.    Devise questions that call for comprehension and application of knowledge skills. 6.    Some of the questions must aim at appraisal of examinees’ ability to analyze, synthesize, and evaluate novel instances of the concepts. If the instances are the same as used in instruction, students are only being asked to recall (knowledge level), 7.    Questions should be written in different formats, e.g., multiple-choice, completion, true-false, short answer etc. to maintain interest and motivation of the students. 8.    Prepare alternate forms of the test to deter cheating and to provide for make-up testing (if needed). 9.    The items should be phrased so that the content rather than the format of the statements will determine the answer. Sometimes the item contains “specific determiners” which provide an irrelevant cue to the correct answer. For example, statements that contain terms like always, never, entirely, absolutely, and exclusively are much more likely to be false than to be true. On the other hand, such terms as may, sometimes, as a rule, and in general are much more likely to be true. Besides, care should be taken to avoid double negatives, complicated sentence structures, and unusual words. 10.    The difficulty level of the items should be appropriate for the ability level of the group. Optimal difficulty for true-false items is about 75 percent, for five-option multiple choice questions about 60 percent, and for completion items approximately 50 percent. However, difficulty in itself is not an end, the item content should be determined by the importance of the subject matter. It is desirable to place a few easy items in the beginning to motivate students, particularly those who are of below average ability. 11.    The items should be devised in such a manner that different taxonomy levels are evaluated. Besides, achievement tests should be power test, not speed test. 12.    Items pertaining to a specific topic or of a particular type should be placed together in the test. Such a grouping facilitates scoring and evaluation. It will also be helpful for the examinees to think and answer the items, similar in content and format, in a better manner without fluctuation of attention and changing the mind-set. 13.    Directions to the examinees should be as simple, clear, and precise as possible, so that even those students who are of below average ability can clearly understand what they are expected to do. 14.    Scoring procedures must be clearly defined before the test is administered. 15. The test constructor, must clearly state optimal testing conditions for test administration. 16.    Item analysis should be carried out to make necessary changes, if any ambiguity is found in the items. Before we discuss preparing the test, it seems quite reasonable that we talk about different types of test items, their characteristics, use and limitations. Items commonly used for Tests of Achievement Two major types of items have been identified: 1.    Constructed Response / Supply items 2.    Structured Response / Select items 1.    Constructed Response / Supply items: In the supply type items the question is so framed that the examinee has to supply or construct the answer on his own in his own words. They generally include the following type: Essay type, Short answer type, Completion type items. 2.    Structured Response / Select items: In the select type items, as the name suggests the examinee is required to select the correct answer from amongst the given or structured options. They are often called objective items. They include: Alternate Response type, Multiple-choice type, Matching type. Self-check Exercise-1 1.    Which principle is NOT generally applicable to constructing questions and tests? a)    Instructions for each type of question must be simple and brief b)    Questions should require complex sentence structures to challenge students c)    Test items must assess specific ability or comprehension of content developed during the course of study d)    Write questions that call for comprehension and application of knowledge skills 2.    Which principle is important for determining the difficulty level of test items? a)    Test items should be extremely difficult to challenge students b)    The difficulty level should be appropriate for the ability level of the group c)    Items should always be at an optimal difficulty of 75% for all types d)    Difficulty should be determined without considering the importance of the subject matter 19.4    CONSTRUCTING OBJECTIVE TYPE TEST ITEMS Construction of test items is a crucial step for the validity of a classroom test is determined by the extent to which performance to be measured is called forth by the test items. It is not enough to have knowledge of subject matter, defined learning outcomes, or a psychological understanding of the students’ mental processes, although all of these are prerequisites. The ability to construct high-quality test items requires knowledge of the principles and techniques of test construction and skill in their application. Objective test forms typically measure relatively simple learning outcome. Self-check Exercise-2 1.    Constructing high-quality test items requires knowledge of the principles and techniques of test construction and ________ in their application. 19.5    ALTERNATIVE RESPONSE TYPE ITEMS Alternative response item is the one that offers two options to choose from. They often consist of a declarative statement that the examinee is asked to mark true or false, right or wrong, correct or incorrect, yes or no, agree or disagree, or the like. Incomplete sentences providing two options to choose from to fill in the blank also fall in this category. The most common form it takes is True - False questions. Most common use of the true- false item is in measuring the examinee’s ability to identify the correctness of statements of fact, definitions of terms, statements of principles, and the like, also to distinguish fact from opinion. Another aspect of understanding that can be measured by the true-false item is the ability to recognize cause-and-effect relationships. This type of item usually contains two true propositions in one statement, and the examinee is to judge whether the relationship between them is true or false. The true-false item also can be used to measure some simple aspects of logic. A common criticism of the true-false item is that an examinee may be able to recognize a false statement as incorrect but still not know what is correct. Suggestions for Constructing True-False Items: •    Avoid trivial statements. •    Avoid broad general statements. •    Avoid the use of negative statements, especially double negatives. •    when a negative word must be used, it should be underlined or put in italics so that students do not overtook it. •    Avoid complex sentences. Avoid including two ideas in one statement, unless cause-effect relationships are being measured. •    Avoid using opinion that is not attributed to some sources, unless the ability to identify opinion is being specifically measured. •    Avoid using true statements and false statements that are unequal in length. •    Avoid using disproportionate numbers of true statements and false statements. Self-check Exercise-3 1 . Alternative response items often consist of a declarative statement that the examinee is asked to mark ________ or ________. 2 .When constructing true-false items, it is recommended to avoid the use of negative statements, especially _______________. 1 9.6 SHORT ANSWER/ COMPLETION TYPE ITEMS The short answer item and the completion item both are supply-type test items. Yet, they are included here for their simplicity. They can be answered by a word, phrase, number, or symbol. The short-answer item uses a direct question whereas the completion item consists of an incomplete statement. Short-answer item is especially useful for measuring problem-solving ability in science and mathematics. Complex interpretations can be made When the short- answer item is used to measure the ability to interpret diagrams, charts, graphs, and pictorial data. When short-answer items are used the question must be stated clearly and concisely. It should be free from irrelevant clues, and require an answer that is both brief and definite. Suggestions for Constructing Short Answer Items •    Word the item so that the required answer is both brief and specific. A direct question is generally more desirable than an incomplete statement. •    Do not take statements directly from textbooks to use as a basis for shortanswer items. •    If the answer is to be expressed in numerical units, indicate the type of answer wanted. •    Blanks for answers should be equal in length and in a column to the right of the question. •    Do not include too many blanks. Self-Check Exercise-4 1 .The short-answer item uses a direct question, whereas the completion item consists of an ________. Question 3: 2 . The short-answer item is especially useful for measuring problem-solving ability in ________ and ________. 19.7    SUMMARY This chapter emphasizes the importance of constructing effective test items to ensure educational assessments are valid, reliable, and aligned with learning objectives. Various types of test items, including true-false statements, and short-answer items are discussed along with their specific purposes in assessing different cognitive levels. The principles of test item construction, such as clarity, simplicity, alignment with learning outcomes, and consideration of students' abilities, are highlighted. The chapter also provides specific guidelines for writing true-false and short-answer items, ensuring they accurately measure students' knowledge and skills. 19.8    GLOSSARY •    Educational Assessment: A systematic process of documenting and using empirical data on the knowledge, skill, attitudes, and beliefs to refine programs and improve student learning. •    Test Item: A question or prompt designed to assess students' understanding, skills, and abilities in a given subject area. •    True-False Statements: Test items that offer two options (true or false) for the examinee to choose from. •    Short-Answer Items: Test items that require the examinee to provide a brief, specific response. 19.9    ANSWERS TO SELF-CHECK EXERCISE Self-Check Exercise-1 1.    b) Questions should require complex sentence structures to challenge students 2.    b) The difficulty level should be appropriate for the ability level of the group Self-Check Exercise-2 1.    skill Self-Check Exercise-3 1.true, false 2.    double negatives Self-Check Exercise-4 1.    incomplete statement 2.    science, mathematics 19.10    REFERENCES/SUGGESTIVE READINGS •    Nitko, A. J., & Brookhart, S. M. (2011). Educational Assessment of Students (6th ed.). Pearson. •    Palomba, C. A., & Banta, T. W. (1999). Assessment Essentials: Planning, Implementing, and Improving Assessment in Higher Education. Jossey-Bass. •    Lane, S., Raymond, M. R., & Haladyna, T. M. (Eds.). (2015). Handbook of Test Development (2nd ed.). Routledge. •    Haladyna, T. M. (2004). Developing and Validating Test Items (3rd ed.). Routledge. 19.11    TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    What are the key principles to consider when constructing test items to ensure they are effective in assessing students' knowledge? 2.    Why is it important to include questions that assess higher-order thinking skills in a test? 3.    What are the guidelines for writing short-answer items to ensure they are clear and concise? ******* UNIT- 20 WRITING TEST ITEMS -2 Structure 20.1    Introduction 20.2    Learning Objectives 20.3    Multiple choice questions, their advantages and disadvantages, Guidelines for Constructing Multiple-Choice Items Self-check Exercise-1 20.4    Matching Type Questions and Suggestions for Constructing Matching Type Questions Self-check Exercise-2 20.5    Essay Type Questions and Guidelines for Constructing Essay type Questions Self-check Exercise-3 20.6    Summary 20.7    Glossary 20.8    Answers to Self -Check Exercises 20.9    References/Suggestive Reading 20.10    Terminal Questions 20.1    INTRODUCTION Dear learner, In educational assessments, the design and construction of test items are crucial for accurately measuring students' understanding and abilities. Two prevalent types of test items are multiple-choice questions (MCQs) and essay-type questions. Each type has its unique structure, advantages, and challenges in construction, and serves distinct purposes in evaluating different cognitive levels. Multiple-choice questions are widely used due to their efficiency in assessing a broad range of content and their ease of scoring. They consist of a stem that presents a problem and a set of options, including one correct answer and several distractors. Essay-type questions, on the other hand, are designed to assess higher-order thinking skills and the ability to organize and express thoughts coherently. This unit focuses on the principles and guidelines for constructing multiple- choice and essay-type questions. By understanding the intricacies of these item types, educators can create assessments that are both valid and reliable, providing meaningful insights into student learning. 20.2    LEARNING OBJECTIVES After studying this unit, you will be able to: •    Understand the structure and components of multiple-choice questions, including the stem and distractors. •    Identify the principles of constructing effective multiple-choice questions to minimize guessing and maximize validity. •    Develop multiple-choice questions that are clear, concise, and aligned with learning objective. •    Understand the purpose and advantages of essay-type questions in assessing higher-order thinking skills. •    Identify the key elements of well-constructed essay prompts, including clarity and specificity. •    Develop essay questions that effectively evaluate students' ability to organize and express their thoughts. 20.3    MULTIPLE CHOICE QUESTIONS, THEIR ADVANTAGES AND DISADVANTAGES, GUIDELINES FOR CONSTRUCTING MULTIPLE-CHOICE ITEMS The multiple-choice item (MCQ) consists of two distinct parts: The first part that contains task or problem is called stem of the item. The stem of the item may be presented either as a question or as an incomplete statement. The form makes no difference as long as it presents a dear and a specific problem to the examinee. Second part presents a series of options or alternatives. Each option represents possible answer to the question. In a standard form one option is the correct or the best answer called the keyed response and the others are mis-leads or foils called distracters. The number of options used differs from one test to the other. An item must have at least three answer choices to be classified as a multiple-choice item. The typical pattern is to have four or five choices to reduce the probability of guessing the answer. A good item should have all the presented options look like probable answers at least to those examinees who do not know the answer. The multiple-choice items, despite having advantages over other items, have some serious limitations as well. It takes time to construct MCQ. They are susceptible to guessing and do not provide any diagnostic information. Multiple-choice items (MCI): is usually divided into three groups or parts: 1-    The Stem 2-    The correct choice or correct answer 3-    The distracters. The Stem, the initial part of each multiple - choice items is known as the stem. It can be complete statement, an incomplete statement and question. The answer can be a word or a group of words. The distracters can be two or three or four options. They are the options which surrounded the answer so that the students with inadequate knowledge cannot find the answer. Objective items require students to select the correct response from several alternatives or supply a word or short phrase to answer a question. It includes multiple - choice, True / false, matching and completion items. Advantages of Multiple-Choice Questions (MCQs) 1.    Broad Content Coverage: MCQs allow for a wide sampling of subject content, enabling educators to assess a broad range of knowledge within a limited testing period. 2.    Objective Scoring: Scoring of MCQs is straightforward and objective, reducing potential biases and ensuring consistency. 3.    Efficiency: MCQs can be scored quickly, especially with automated systems, making them suitable for large classes. 4.    Diagnostic Information: Well-constructed MCQs can provide diagnostic information about students' understanding of specific concepts. 5.    Versatility: MCQs can assess various cognitive levels, from basic recall of facts to higher-order thinking skills like analysis and application. Disadvantages of Multiple-Choice Questions (MCQs) 1.    Time-Consuming Construction: Developing high-quality MCQs is time-consuming and requires careful crafting to ensure clarity and avoid ambiguity. 2.    Guessing: Students may guess the correct answer, which can undermine the validity of the assessment if not properly mitigated with plausible distractors. 3.    Limited Depth of Understanding: MCQs may not adequately assess complex thinking or the ability to synthesize and evaluate information. 4.    Susceptibility to Test-Wiseness: Students with test-taking skills (testwise students) may perform better than their actual knowledge warrants. 5.    Lack of Diagnostic Depth: While MCQs can identify what students know, they may not provide insights into how or why students think as they do, limiting the depth of diagnostic feedback. Guidelines for Constructing Multiple-Choice Items >    Be sure that the stem clearly formulates a problem. The stem should be worded so that the examinee clearly understands the question being asked before he reads the answer choices. >    Stem should be written either in direct question form or in an incomplete statement form. >    The stem of the item should present only one problem. Two concepts must not be combined together to form a single stem. >    Include as much of the item in the stem and keep options as short as possible. This leads to economy of space, economy of reading time and clear statement of the problem. >    Unnecessary words or phrases should not be included in the stem. Such words add to the length and complexity of the stem but do not enhance meaningfulness of the stem. The stem should be written in simple, concise and clear form. >    Avoid the use of negative words in the stem of the item. There are times when it is important for the examinee to detect errors or to know exceptions. For these purposes, sometimes the use of ‘not’ or ‘except’ is justified in the stem. When a negative word is used in a stem it should be highlighted. > Use novel material in formulating problems to measure understanding or ability to apply principles. Do not focus too closely on rote memory of the text that neglects measurement of the ability to use information. > Use plausible distracters as alternatives. If an examinee who does not know the correct answer is not distracted by a given alternative, that alternative is not plausible and it will add nothing to the functioning of the item. > Be sure that no unintentional clues > The correct answer should appear at each position in almost equal numbers. While constructing multiple-choice item, some examiners have a tendency to place correct alternative at the first position. Some place it in the middle and others at the end. Such tendencies should be consciously controlled. > Avoid using ‘none of the above’. ‘all of the above’. both a and b etc. as options for an MCQ. > Alternatives should be grammatically consistent with the stem. Grammatical Inconsistency provides irrelevant clues. Self-check exercise 1 1.    Which of the following describes the stem in a multiple-choice question? a)  The correct answer b)  The question or problem c)    The distractors d)    The scoring guidelines 2.    What is the purpose of including distractors in a multiple-choice question? a)To confuse students b)To increase the difficulty level c) To ensure the correct answer is easily identifiable d)To challenge students' understanding and knowledge 20.4 MATCHING TYPE QUESTIONS AND SUGGESTIONS FOR CONSTRUCTING MATCHING TYPE QUESTIONS Matching type questions consists of two parallel columns with each word, number, or symbol in one column being matched to a word, sentences, or phrase in the other column. Items in the column for which a match is sought are called premises, and the items in the column from which the selection is made are called responses. Suggestions for Constructing Matching Type Questions > Use only homogeneous material in a single matching exercise. > Include an unequal number of responses and premises and instruct the student that responses may be used once, more than once, or not at all. > Keep the list of items to be matched brief, and place the shorter responses on the right. > Arrange the list of responses in logical order. Place words in alphabetical order and numbers in sequence. > Indicate in the directions the basis for matching the responses and premises. > Ambiguity and confusion will be avoided. And testing time will be saved. > Place all of the items for one matching exercise on they same page. > Self- check Exercise-2 1.    What is the main characteristic of matching type questions? a)    They require short answers b)    They involve selecting from columns c)    They assess higher-order thinking d)    They are subjective in nature 2.     What is a potential benefit of using matching type questions? a)    They allow for creative responses b)    They measure deep understanding c)    They are quick to grade d)    They are easy to construct 20.5 ESSAY TYPE QUESTIONS AND GUIDELINES FOR CONSTRUCTING ESSAY TYPE QUESTIONS There are two major purposes for using essay questions that address different learning outcomes. One purpose is to assess students understanding of subject-matter content. The other purpose is to assess students writing abilities: These two purposes are so different in nature that it is best to treat them separately. An essay question is “a test item which requires a response composed by the examinee, usually in the form of one or more sentences, of a nature that no single response or pattern of responses can be listed as correct, and the accuracy and quality of which can be judged subjectively only by one skilled or informed in the subject.” An essay question should meet the following criteria: 1.    Requires examinees to compose rather than select their response. Multiple-choice questions, matching exercises, and true-false items are all examples of selected response test items because they require students to select an answer from a list of possibilities provided by the test maker, whereas essay questions require students to construct their own answer. 2.    Elicits student responses that must consist of one or more sentences. 3.    No single response or single response pattern is correct. 4.    The accuracy and quality of students’ responses to essays must be judged be subjectively by a competent specialist in the subject. Guidelines for Constructing Essay Questions: > Clearly define the intended learning outcome to be assessed by the item. > Avoid using essay questions for intended learning outcomes that are better assessed with other kinds of assessment. > Define the task and shape the problem situation. >    Helpful Instructions: Specify the relative point value and the approximate time limit in clear directions. >    Helpful Guidance: State the criteria for grading >    Use several relatively short essay questions rather than one long one. > Avoid the use of optional questions. Students should not be permitted to choose one essay question to answer from two or more optional questions. The use of optional questions should be avoided for the following reasons. Students may waste time deciding on an option. Some questions are likely to be harder which could make the comparative assessment of student abilities unfair. Self- check Exercise-3 1.    What distinguishes essay type questions from other types of questions? a)    They require short answers b)    They are subjective in nature c)    They have a single correct response d)    They measure specific knowledge 2.    Which guideline is important when constructing essay type questions? a)    Include multiple questions within one essay b)    Keep the questions vague to encourage creativity c)    Provide a clear grading rubric d)    Limit the word count for each response 20.6    SUMMARY MCQs are structured with a stem presenting a problem or question and multiple options, one of which is correct and the rest are distractors. However, constructing effective MCQs can be time-consuming, they are susceptible to guessing, and may not fully assess deep understanding or higher-order thinking skills. Essay questions require students to construct detailed responses in their own words, demonstrating understanding and application of knowledge. They are valuable for assessing complex learning outcomes, encouraging critical thinking, and evaluating higher-order cognitive skills such as analysis, synthesis, and evaluation. While MCQs are efficient for assessing basic knowledge and can cover a wide range of content, essay questions provide deeper insights into students' comprehension and analytical abilities. 20.7    GLOSSARY •    Stem: -The main part of a multiple-choice question that presents the problem or task to the examinee. •    Distractors:- Incorrect answer choices in a multiple-choice question that are designed to challenge the examinee's knowledge and understanding. 20.8 ANSWERS TO SELF-CHECK EXERCISES Self-check Exercise-1 1 .b) The question or problem 2 . d) To challenge students' understanding and knowledge Self-check Exercise-2 1 .b) They involve selecting from columns 2 .c) They are quick to grade Self-check Exercise-3 1.    b) They are subjective in nature 2.    c) Provide a clear grading rubric 20.9    REFERENCES/SUGGESTIVE READINGS •    Ebel, Robert L.(1966) “Measuring Educational Achievement, Prentice Hall of India Pvt. Ltd. •    Gronlund, N. E. (1976), Measurement and Evaluation in Teaching. McMillan, USA. •    Hopkins, C.D. and Antes, R.L. (1990). Classroom measurement and evaluation. Itasca,Illinois: Peacock. •    Mehrens, W.A. and Lehmann, I.J. (1984). Measurement and evaluation in education and psychology.(3rd Ed.) New York: Holt, Rinehart and Winston. Nandra, I.D.S.(2011). Learning Resources and Assessment of Learning.Patiala, 21st Century Publications. 20.10    TERMINAL QUESTIONS Dear learners, please check you progress by attempting the following questions: 1.    Describe the structure of a multiple-choice item (MCQ) and explain the function of each part. 2.    Discuss the advantages and disadvantages of using multiple-choice questions (MCQs) in assessments. Provide specific examples to support your points. 3.    What are some key guidelines for constructing effective multiple-choice items? Explain why each guideline is important for the validity of the assessment. 4.    Explain the characteristics and suggestions for constructing matching type questions. 5.    What are the purposes of using essay type questions in assessments? Discuss the guidelines for constructing effective essay questions. ******* 195