Papers
Topics
Authors
Recent
Search
2000 character limit reached

Assessment Self-Efficacy: MASE Instrument

Updated 9 July 2026
  • MASE is a validated instrument measuring students' self-efficacy in assessment contexts, distinguishing between cognitive execution and emotional regulation.
  • It employs parallel seven-item scales for low-stakes quizzes (MASE-Q) and high-stakes exams (MASE-E), with analyses confirming a reliable two-factor structure.
  • Robust invariance testing across time and cohorts supports the instrument’s reliability for evaluating pedagogical interventions and longitudinal research.

Searching arXiv for the specified paper and closely related assessment/self-efficacy work. The Measure of Assessment Self-Efficacy (MASE) is an instrument developed to assess students’ beliefs about their capabilities when taking assessments rather than their beliefs about domain-specific content knowledge. It was designed to capture two types of efficacy beliefs—“comprehension and execution” and “emotional regulation”—in two contrasting assessment scenarios: a low-stakes online mathematics quiz (MASE-Q) and a high-stakes invigilated final mathematics exam (MASE-E). Across three sequential studies, the instrument was developed, refined, and validated through confirmatory factor analysis, MGCFA, and cohort comparisons, with results supporting parallel two-factor measurement models and substantial evidence of invariance and validity (Riegel et al., 26 Aug 2025).

1. Conceptual basis and construct definition

MASE is grounded in Bandura’s social-cognitive theory. In Bandura’s formulation, self-efficacy consists of “beliefs in one’s capabilities to organize and execute the courses of action required to produce given attainments.” Within that framework, MASE addresses a specific omission in existing self-efficacy research: many assessment-related measures focus on beliefs about content-specific tasks, but not on beliefs about taking assessments as such (Riegel et al., 26 Aug 2025).

The construct is organized into two related but conceptually distinct dimensions. Comprehension and execution refers to the belief that one can understand an assessment’s demands and carry out the required cognitive operations within its constraints. Emotional regulation refers to the belief that one can manage the negative emotions that arise in preparing for and during the assessment, including anxiety and frustration.

This distinction is central to the instrument’s rationale. The measure does not treat assessment performance as reducible either to content mastery alone or to affective coping alone. Instead, it models assessment-taking self-efficacy as comprising both beliefs about understanding and acting under assessment conditions and beliefs about maintaining composure and positivity under those conditions. A common misconception is therefore that assessment self-efficacy is simply academic self-efficacy in another form; the MASE framework explicitly separates assessment-taking beliefs from broader or more content-bound efficacy judgments.

2. Instrument architecture and item development

In Study 1, the research team generated ten items following Bandura’s guidelines for self-efficacy scale construction. Items were scored from 0 to 100 via slider, and they were formulated to be as parallel as possible across quiz and exam contexts (Riegel et al., 26 Aug 2025).

The original item pool included, for comprehension and execution, statements such as: “I can understand the content and skills needed for the assessment,” “I can fully master the requirements I need for this assessment,” “When preparing for the assessment, I can organize my time well,” “During the assessment, I can answer the questions within the time constraints,” and “I can correctly answer the questions on the assessment.” For emotional regulation, the original pool included: “During my preparation, I am able to cope with my negative emotions toward the assessment,” “Even when I struggle while studying, I am able to stay positive about my ability to succeed,” “During the assessment, I am able to cope with my negative emotions toward the assessment,” and “Even when I struggle during the assessment, I am able to stay positive about my ability to succeed.”

After Study 1 suggested minor misfits, item-8 (“Even when I struggle while studying…”) was rephrased in Study 2, and several context-specific items were removed. The result was a seven-item structure for each form: four items for comprehension and execution and three items for emotional regulation.

Form Scenario Final composition
MASE-Q Low-stakes online mathematics quiz 4 comprehension and execution items; 3 emotional regulation items
MASE-E High-stakes invigilated final mathematics exam 4 comprehension and execution items; 3 emotional regulation items

For MASE-Q, the final comprehension and execution items were: “I can understand the content and skills needed for the assessment,” “I can fully master the requirements I need for this assessment,” “When preparing for the assessment, I can organize my time well,” and “I can correctly answer the questions on the assessment.” The final emotional regulation items were: “During my preparation, I am able to cope with my negative emotions toward the assessment,” “During the assessment, I am able to cope with my negative emotions toward the assessment,” and “Even when I struggle during the assessment, I am able to stay positive about my ability to succeed.”

For MASE-E, the final comprehension and execution items were: “I can understand the content and skills needed for the assessment,” “I can fully master the requirements I need for this assessment,” “I can understand the questions in the assessment,” and “I can correctly answer the questions on the assessment.” The final emotional regulation items were: “During my preparation, I am able to cope with my negative emotions toward the assessment,” “While studying, I am able to stay positive about my ability to succeed,” and “Even when I struggle during the assessment, I am able to stay positive about my ability to succeed.”

The parallel construction of MASE-Q and MASE-E is methodologically significant. It permits the same latent distinction to be examined across low-stakes and high-stakes contexts while preserving scenario-specific wording where necessary.

3. Factor structure, model fit, and reliability

Study 1 (N=301)(N = 301) supported the initial ten-item, two-factor model in both assessment contexts. After refinement to seven items per form, Study 2 (N=277)(N = 277) and Study 3 (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329) confirmed the two correlated factors for quiz and exam forms with very good fit (Riegel et al., 26 Aug 2025).

For MASE-Q in Study 2, the reported CFA model fit indices were as follows. At Time 1, χ2/df=2.38\chi^2/df = 2.38, TLI=.976TLI = .976, CFI=.986CFI = .986, RMSEA=.071RMSEA = .071 with 90%CI[.037–.104]90\%CI [.037–.104], and SRMR=.023SRMR = .023. At Time 2, χ2/df=2.27\chi^2/df = 2.27, (N=277)(N = 277)0, (N=277)(N = 277)1, (N=277)(N = 277)2 with (N=277)(N = 277)3, and (N=277)(N = 277)4. At Time 3, (N=277)(N = 277)5, (N=277)(N = 277)6, (N=277)(N = 277)7, (N=277)(N = 277)8 with (N=277)(N = 277)9, and (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)0.

For MASE-E in Study 2, the reported CFA model fit indices were: at Time 1, (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)1, (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)2, (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)3, (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)4 with (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)5, and (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)6. At Time 2, (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)7, (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)8, (cohorts of 277 and 329)(\text{cohorts of } 277 \text{ and } 329)9, χ2/df=2.38\chi^2/df = 2.380 with χ2/df=2.38\chi^2/df = 2.381, and χ2/df=2.38\chi^2/df = 2.382. At Time 3, χ2/df=2.38\chi^2/df = 2.383, χ2/df=2.38\chi^2/df = 2.384, χ2/df=2.38\chi^2/df = 2.385, χ2/df=2.38\chi^2/df = 2.386 with χ2/df=2.38\chi^2/df = 2.387, and χ2/df=2.38\chi^2/df = 2.388.

The standardized factor loadings reported for Study 2, Time 1, were also strong. For MASE-Q, the comprehension and execution loadings were χ2/df=2.38\chi^2/df = 2.389 with TLI=.976TLI = .9760, and the emotional regulation loadings were TLI=.976TLI = .9761 with TLI=.976TLI = .9762. For MASE-E, the comprehension and execution loadings were TLI=.976TLI = .9763 with TLI=.976TLI = .9764, and the emotional regulation loadings were TLI=.976TLI = .9765 with TLI=.976TLI = .9766.

Cronbach’s alpha for any TLI=.976TLI = .9767-item subscale may be computed as

TLI=.976TLI = .9768

where TLI=.976TLI = .9769 is the variance of item CFI=.986CFI = .9860 and CFI=.986CFI = .9861 the total test variance.

In all studies, subscale alphas ranged between .84 and .91, indicating good to excellent internal consistency. This suggests that the brief seven-item forms retained psychometric coherence despite the item reduction from the original ten-item pool.

4. Measurement invariance and comparability

Study 2 tested invariance of MASE-Q and MASE-E over three measurement occasions, and Study 3 tested invariance of MASE-E across two student cohorts. The analyses followed Cheung and Rensvold’s criterion of CFI=.986CFI = .9862 (Riegel et al., 26 Aug 2025).

For MASE-Q across time, configural, metric, scalar, and strict (residual) invariance were achieved, with CFI=.986CFI = .9863. For MASE-E across time, configural, metric, and scalar invariance were achieved, with CFI=.986CFI = .9864; residual invariance slightly exceeded the criterion at CFI=.986CFI = .9865, but strong invariance was regarded as sufficient for latent-mean comparisons. For MASE-E across cohorts, full strict invariance was obtained from configural through residual models, with CFI=.986CFI = .9866.

These invariance findings are central to the instrument’s research utility. They indicate that the latent structure is stable across repeated measurements and, in the case of the exam form, across cohorts. A plausible implication is that observed differences across time or groups are less likely to be artifacts of measurement non-equivalence. At the same time, the asymmetry between MASE-Q and MASE-E should be noted: the quiz form has temporal invariance support, whereas cohort-level invariance evidence was reported for the exam form rather than for the quiz form.

5. Scoring, interpretation, and analytical uses

Each MASE item is rated on a 0–100 slider, and subscale scores are computed as the mean of their items, yielding a range of 0–100. Higher scores indicate stronger self-efficacy (Riegel et al., 26 Aug 2025).

Because the scales are invariant over time—and, for MASE-E, across cohorts—researchers can validly compare latent means longitudinally or between groups. This makes the instrument suitable not only for cross-sectional description but also for change-sensitive designs in which assessment-related beliefs are expected to shift across preparation periods, course progress, or cohort boundaries.

The paper also identifies applied interpretations for low scores on each dimension. Low comprehension and execution scores may suggest the need for test-taking strategy training. Low emotional regulation scores may indicate the benefit of anxiety-management or positive-thinking interventions. These interpretations preserve the instrument’s two-factor logic: difficulties in understanding and acting under assessment demands are not treated as identical to difficulties in regulating affect during preparation and performance.

The reported applications include evaluating the impact of pedagogical or psychological interventions such as test-taking workshops and stress-reduction programs; investigating how self-efficacy for low-stakes quizzes relates to, or predicts, eventual exam self-efficacy and performance; and exploring the interplay between assessment self-efficacy, achievement emotions, and academic outcomes in longitudinal and cross-cultural studies.

6. Scope, limitations, and prospective extensions

Validation of MASE has been confined to mathematics contexts in undergraduate STEM courses and to two assessment formats: a low-stakes online quiz and a high-stakes invigilated final exam (Riegel et al., 26 Aug 2025).

Accordingly, generalizability to other domains, including languages and humanities, and to other assessment types, including group projects and portfolios, remains to be established. The quiz form, MASE-Q, has not yet been tested for invariance across different cohorts or institutions. These limitations constrain the current empirical scope of the instrument and delimit the populations and settings in which strong psychometric claims can presently be made.

The paper identifies several future directions: extension to secondary and graduate-level learners; adaptation and validation of parallel forms for other modalities such as oral exams and open-ended projects; combination of MASE with fine-grained task-specific self-efficacy scales to disentangle content versus context influences; and incorporation of MASE into large-scale assessment research to better understand self-efficacy trajectories and equity issues in STEM.

Taken together, these directions position MASE-Q and MASE-E as brief, reliable, and valid measures of two facets of assessment self-efficacy. Their main contribution is to move beyond broad academic self-efficacy indices toward a more differentiated account of the beliefs students hold when preparing for and taking specific forms of assessment.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Measure of Assessment Self-Efficacy (MASE).