ERIC - Search Results

Publication Date

In 2026	0
Since 2025	0
Since 2022 (last 5 years)	0
Since 2017 (last 10 years)	1
Since 2007 (last 20 years)	2

Descriptor

Interrater Reliability	11
Test Reliability	11
Test Use	11
Test Validity	7
Scoring	6
Test Construction	5
Educational Assessment	4
Testing	4
Student Evaluation	3
Test Format	3
Elementary School Students	2
Evaluation Methods	2
Generalizability Theory	2
Language Arts	2
Performance Based Assessment	2
Performance Tests	2
Portfolio Assessment	2
Portfolios (Background…	2
Quality Control	2
Scores	2
Secondary Education	2
Test Content	2
Test Interpretation	2
Test Items	2
Testing Accommodations	2
More ▼

Source

Academic Medicine	1
Applied Measurement in…	1
Education and Training in…	1
International Journal of…	1
Journal of Consulting and…	1
New York State Education…	1

Author

Alderson, J. Charles	1
Boivin, Micheline	1
Conroy, Maureen A.	1
Dunbar, Stephen B.	1
Gearhart, Maryl	1
Herman, Joan L.	1
Meier, Augustine	1
Novak, John R.	1
Reckase, Mark D.	1
Smith, Richard Merrill	1
Wolfe, Edward W.	1
More ▼

Publication Type

Journal Articles	5
Reports - Evaluative	4
Reports - Research	4
Speeches/Meeting Papers	3
Guides - Non-Classroom	2
Guides - General	1
Numerical/Quantitative Data	1

Education Level

Early Childhood Education	1
Elementary Education	1
Grade 3	1
Grade 4	1
Grade 5	1
Grade 6	1
Grade 7	1
Grade 8	1
High Schools	1
Intermediate Grades	1
Junior High Schools	1
Middle Schools	1
Primary Education	1
Secondary Education	1
More ▼

Audience

Administrators	1
Practitioners	1
Teachers	1

Location

New York

Laws, Policies, & Programs

Individuals with Disabilities…	1
No Child Left Behind Act 2001	1

Assessments and Surveys

What Works Clearinghouse Rating

Showing all 11 results Save | Export

ITC Guidelines for the Large-Scale Assessment of Linguistically and Culturally Diverse Populations

Peer reviewed

Direct link

International Journal of Testing, 2019

These guidelines describe considerations relevant to the assessment of test takers in or across countries or regions that are linguistically or culturally diverse. The guidelines were developed by a committee of experts to help inform test developers, psychometricians, test users, and test administrators about fairness issues in support of the…

Descriptors: Test Bias, Student Diversity, Cultural Differences, Language Usage

New York State Alternate Assessment Technical Report, 2013-14

Download full text

New York State Education Department, 2014

This technical report provides an overview of the New York State Alternate Assessment (NYSAA), including a description of the purpose of the NYSAA, the processes utilized to develop and implement the NYSAA program, and Stakeholder involvement in those processes. The purpose of this report is to document the technical aspects of the 2013-14 NYSAA.…

Descriptors: Alternative Assessment, Educational Assessment, State Departments of Education, Student Evaluation

Client Verbal Response Category System: Preliminary Data.

Peer reviewed

Meier, Augustine; Boivin, Micheline – Journal of Consulting and Clinical Psychology, 1986

The Client Verbal Response Category System classifies client responses into Temporal, Directional and Experiential categories. The categories with their subcategories are defined, interjudge reliability data is presented, and the instrument's utility in psychotherapy process research is demonstrated. Initial results indicate that the instrument is…

Descriptors: Client Characteristics (Human Services), Interrater Reliability, Psychotherapy, Research Tools

An Analysis of the Reliability and Stability of the Motivation Assessment Scale in Assessing the Challenging Behaviors of Persons with Developmental Disabilities.

Peer reviewed

Conroy, Maureen A.; And Others – Education and Training in Mental Retardation and Developmental Disabilities, 1996

This study assessed the intra-rater and inter-rater reliability of the Motivation Assessment Scale as used with 20 adults with mental retardation, expanding the results of previous research by evaluating across additional time and administrations. Results from 19 raters indicated variable moderate-to-low intra-rater and inter-rater reliability.…

Descriptors: Adults, Behavior Problems, Interrater Reliability, Measures (Individuals)

Statistical Test Specifications for Performance Assessments: Is This an Oxymoron?

Download full text

Reckase, Mark D. – 1997

This paper argues that special procedures for constructing assessment tools containing performance assessment tasks are unnecessary and that current test methodology can easily be generalized to complex performance assessment tasks without destroying the desirable characteristics of those tasks. Reasonable statistical requirements for sound…

Descriptors: Educational Assessment, Generalizability Theory, High Stakes Tests, Interrater Reliability

The Triple-Jump Examination as an Assessment Tool in the Problem-Based Medical Curriculum at the University of Hawaii.

Peer reviewed

Smith, Richard Merrill – Academic Medicine, 1993

A University of Hawaii study compared objective and subjective assessments of the three-step triple jump examination which tests medical students' clinical problem-solving processes. Subjects were 58 first-year students. Results found the subjective assessments were more consistent across problems of varying difficulty level than were objective…

Descriptors: Case Studies, Difficulty Level, Higher Education, Interrater Reliability

Language Test Construction and Evaluation.

Alderson, J. Charles; And Others – 1995

The guide is intended for teachers who must construct language tests and for other professionals who may need to construct, evaluate, or use the results of language tests. Most examples are drawn from the field of English-as-a-Second-Language instruction in the United Kingdom, but the principles and practices described may be applied to the…

Descriptors: Educational Trends, English (Second Language), Interrater Reliability, Language Tests

Quality Control in the Development and Use of Performance Assessments.

Peer reviewed

Dunbar, Stephen B.; And Others – Applied Measurement in Education, 1991

Issues pertaining to the quality of performance assessments, including reliability and validity, are discussed. The relatively limited generalizability of performance across tasks is indicative of the care needed to evaluate performance assessments. Quality control is an empirical matter when measurement is intended to inform public policy. (SLD)

Descriptors: Educational Assessment, Generalization, Interrater Reliability, Measurement Techniques

A Report on the Reliability of a Large-Scale Portfolio Assessment for Language Arts, Mathematics, and Science.

Download full text

Wolfe, Edward W. – 1996

Although portfolio assessment is becoming increasingly popular, it may not survive unless portfolio scoring can meet the demands of large-scale assessment standards. The results of studies of interrater reliability with large-scale portfolio assessments have been mixed. This paper reports the scoring results of a nationwide portfolio pilot in…

Descriptors: Decision Making, Generalizability Theory, Interrater Reliability, Language Arts

Issues in Portfolio Assessment: The Scorability of Narrative Collections. Project 3.1: Studies in Improving Classroom and Local Assessments.

Download full text

Gearhart, Maryl; Novak, John R.; Herman, Joan L. – 1994

Technical questions regarding the reliability and validity of large-scale portfolio assessment were studied which focused on: (1) whether raters can score collections of writing reliably with rubrics designed for single samples; (2) whether ratings derived from different frameworks differ in their capacities to support technically sound…

Descriptors: Educational Assessment, Elementary Education, Elementary School Students, Essay Tests

Performance Testing Manual and Workbook for Vocational Education Administrators and Teachers.

Florida State Dept. of Education, Tallahassee. Div. of Vocational, Adult, and Community Education. – 1991

This packet contains a manual and a workbook for developing performance tests in vocational education. The manual gives an in-depth description of how to develop, score, and use performance tests. It includes the following sections: definitions of performance testing, steps in developing a performance test, selecting a performance development…

Descriptors: Interrater Reliability, Performance Tests, Postsecondary Education, Scoring