The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by exacerbating performative anxiety and the sociological power imbalances inherent to academic hierarchies. This paper presents the theoretical foundation for the "Socratic Test," an automated, computer-mediated conversational assessment. By integrating Dynamic Assessment principles, multimodal workspaces, Bloom's Taxonomy for real-time proctoring, and the SOLO Taxonomy for structural evaluation, the Socratic Test actively maps a student's cognitive boundaries. This paper formalizes the use of graduated scaffolding to quantify the Zone of Proximal Development (ZPD) and details a non-compensatory, additive grading architecture that prioritizes mastery over penalty and human-AI alignment to ensure unprecedented measurement reliability.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Explain it Like I'm 14
Overview
This paper introduces the “Socratic Test,” a new kind of exam where you chat with an AI that asks you questions, gives you hints, and looks at your work (like drawings, math, or code) as you solve problems. The goal is to measure what you truly understand—not just what you memorized—and to do it in a way that feels fair, less stressful, and hard to game.
What Questions Is the Paper Trying to Answer?
- How can we test students in a way that is both fair and accurate?
- How can we keep the good parts of oral exams (deep understanding) without the bad parts (stress, bias, and inconsistency)?
- Can AI run a conversation-based test that adapts to each student, gives helpful hints, and still grades reliably?
- How do we prevent cheating, skipping hard questions, or AI mistakes from ruining the test?
How the Socratic Test Works (In Simple Terms)
Think of the test like playing a skill-based game with a helpful coach:
- The AI asks you questions through chat. If you get stuck, it gives you hints, starting small and getting more direct if needed.
- You can draw diagrams, write code, or do math on built-in tools. This lets you “think on paper,” not just in words.
- The AI slowly raises the difficulty, like climbing steps on a ladder, to see how far you can go.
Here are the main ideas behind it, using easy analogies:
- Dynamic Assessment: Like a coach helping during practice. The test doesn’t just check what you already know; it sees what you can do with a little help. This helps find your “Zone of Proximal Development” (ZPD): what you can do alone vs. with guidance.
- Bloom’s Taxonomy (for questions): Imagine a ladder of thinking skills—from Remember (step 1) to Create (step 6). The AI starts low and climbs higher as you show readiness.
- SOLO Taxonomy (for answers): This checks how well-structured your answer is. Are you giving one fact (basic), several facts (better), or connecting them into a clear, logical explanation (best)?
- Multimodal Workspaces: You get a whiteboard, code editor, and calculator. These tools let you offload thinking (like jotting notes during a puzzle) so you can show your logic, not just your writing ability.
The Hint System (System of Least Prompts)
The AI follows a clear hint order:
- A gentle nudge
- A specific clue
- A targeted scaffold (very focused help)
- Direct instruction (it gives the missing piece so you can move on)
Even if you forgot a fact, the AI can supply it and still let you try harder questions. You’re not “stuck” just because of one memory slip.
Scoring That Rewards Mastery (Not Point Deductions)
- You start at zero and earn points as you show what you can do.
- You can’t farm lots of easy questions to get a top score. Points live in “buckets” (for each topic and difficulty level), and each bucket has a cap. Once a bucket is full, doing more easy questions won’t raise your grade. This pushes you to try higher-level thinking when you’re ready.
- Two exam styles:
- Stair-Step Mode: More structured; teachers set clear levels per topic.
- Organic Mode: More open; the conversation flows naturally across higher-level skills.
Built-In Fairness and Safety Checks
- Evidence Buffer (Oversampling): The AI collects a little extra proof before moving on, so small grading mistakes won’t block your progress.
- Shadow Ledger: If you skip or dodge a question, the system secretly counts that attempt so you can’t skip your way to a high score. You move forward, but your potential score is limited for that area.
- Out-of-Scope Audit: If you think the AI asked something unfair or off-topic, you can flag it. The AI moves on, and the teacher reviews it later. If you were right, you get credit—this protects you from AI mistakes.
Reliable Grading: AI + Teacher, Step by Step
The platform separates the live conversation from the final grade:
- During the test, the system logs which question level you faced, how many hints you used, and your answer.
- After the test, a 7-step process aligns AI grading with the teacher’s judgment. The teacher checks tricky cases, sets examples, and can override anything. If they change one student’s score for a certain type of answer, the system fairly updates everyone else’s similar answers too.
Main Findings and Why They Matter
From early pilot use (students and teachers tried it in real classes):
- Students felt less stressed than in face-to-face oral exams and often even less than on traditional written tests. Typing and having time to think lowered anxiety.
- Students said the AI pushed them to explain their “why,” not just give final answers. It helped them reach their limits in a supportive way.
- Teachers could better tell who really understood the material—at both the pass/fail edge and among top performers.
- The system separated “math speed” or public speaking skills from true understanding. That means it measures what it’s supposed to measure.
- It helped protect against AI errors (with the audit) and against students dodging tough questions (with the shadow ledger).
What This Could Change
- Fairer, deeper exams: Tests that adapt to you, reward real thinking, and don’t punish a single memory slip.
- Less anxiety: A typed, coached conversation instead of a stressful spotlight interview.
- Strong academic integrity: If you used AI to write a paper, you still have to defend your ideas on the spot. This helps confirm real understanding.
- Better feedback: Teachers get detailed “maps” of class-wide misunderstandings and can fix them quickly in future lessons.
- Scalable and trustworthy: Because the grading is transparent and teacher-aligned, schools can use it in high-stakes settings with confidence.
In short, the Socratic Test turns exams into intelligent, guided conversations that discover what you truly know and can do—fairly, calmly, and clearly.
Knowledge Gaps
Below is a single, focused list of concrete knowledge gaps, limitations, and open questions the paper leaves unresolved. Each point is framed to enable actionable follow-up research or implementation work.
- Psychometric validation is unreported: no empirical evidence for reliability (inter/intra-rater, test–retest), construct validity, criterion validity, or concurrent validity versus traditional written and face-to-face oral exams.
- No standard-setting method is specified: how criterion-referenced cut scores are determined and equated across cohorts, courses, and years (e.g., Angoff, Bookmark) remains undefined.
- Lack of score comparability studies: no equating procedures for scores obtained via Stair-Step vs Organic modes, or across different topic mixes and conversation paths.
- ZPD quantification is only heuristic: γk and hint counts are proposed, but no formal latent-trait model (e.g., IRT, cognitive diagnostic models) links scaffolding behavior to ability estimates or reports measurement error/confidence intervals.
- Bloom’s-as-navigation lacks validation: evidence is needed that Bloom levels map to consistent difficulty across disciplines; item-bank design, difficulty calibration, and adaptivity control are unspecified.
- SOLO-by-AI content correctness risk: the evaluation centers on structural complexity; safeguards to ensure semantically incorrect but well-structured responses are not over-credited are not empirically demonstrated.
- Minimal calibration sample: the 20-item calibration and 15-item validation per cohort may be too small to ensure generalizable SOLO scoring; required sample sizes and stability across domains are unknown.
- Shadow Ledger fairness and transparency: the hidden penalty and baseline mapping (Table: expected SOLO per Bloom) are unvalidated; ethical and legal implications of undisclosed scoring rules require study; sensitivity to different baseline matrices is unexplored.
- Oversampling Factor tuning: no guidance or theory for choosing the Evidence Buffer percentage; effects on exam length, fatigue, and fairness, and statistical guarantees that buffers absorb provisional misclassifications are untested.
- Hint design fidelity: how to author and validate the “System of Least Prompts” to avoid answer leakage, ensure consistency across students, and maintain construct purity is unspecified.
- Multimodal evidence parsing: the reliability and validity of evaluating diagrams, math derivations, and code (e.g., vision/OCR accuracy, code execution/isolation) are not reported; effects of device/stylus access inequities are unexamined.
- Fairness and bias auditing are absent: no DIF/measurement invariance analyses by gender, race/ethnicity, disability status, language proficiency, or SES; impact of typed interaction on ESL or dyslexic students requires study despite multimodal affordances.
- Affective filter claims rely on self-report: controlled studies measuring anxiety, cognitive load, and performance (including subgroup analyses: neurodivergent, ESL, first-gen) are needed.
- Security and integrity controls are under-specified: identity verification, prevention of collusion/secondary devices, jailbreak attempts to elicit answers/rubrics, and logging/forensics protocols need rigorous design and evaluation.
- Privacy, consent, and governance: data retention, FERPA/GDPR compliance, consent for using student interactions as calibration data, and transparency around algorithmic decisions require formal policies and audits.
- Model drift and reproducibility: LLM versioning, freezing, and audit trails to ensure comparable behavior across terms/providers are not detailed; latency/cost impacts at scale are unknown.
- Instructor workload and scalability: time-on-task for calibration, audits, overrides, and appeals in large enrollments is not quantified; tooling needed to keep per-student cost feasible is unspecified.
- Parameter sensitivity: the grading is sensitive to V(b,s), γk, bucket caps, and Oversampling; there is no empirical basis, normative guidance, or robustness analysis for these choices.
- Progression dynamics and student agency: risk that early Shadow Ledger accrual locks students out of higher tiers; policies for recovery opportunities and transparent pacing guidance are unclear.
- Appeals process equity: reliance on written justifications may advantage more articulate students; throughput limits, SLAs, and inter-rater reliability of appeal decisions are not reported.
- Out-of-Scope audit consistency: inter-rater reliability of the 4-point audit scale among instructors/TAs and potential incentives for strategic flagging by students are untested.
- External validity across disciplines: applicability to labs, studio/performance arts, language learning/speaking tasks, proofs, and clinical skills is unverified; required modality adaptations are unspecified.
- Long-term learning impact: no longitudinal data on retention, transfer, or growth-mindset outcomes relative to alternative assessments.
- Comparisons to CAT/IRT: how the conversational adaptivity aligns or conflicts with established adaptive testing psychometrics and whether hybrid models could yield error bars is unexplored.
- Accessibility and accommodations: compliance with ADA/WCAG (screen readers, keyboard-only navigation, captioning), low-bandwidth/offline contingencies, and extended-time policies for the conversational format need evaluation.
- Cold-start problem: zero-shot SOLO anchoring quality for new courses/domains without prior calibration is unknown; strategies for bootstrapping reliable scoring are absent.
- Cross-lingual support: handling multilingual responses, code-switching, non-Latin scripts in whiteboards, and translation effects on SOLO judgments are unaddressed.
- Authoring and content governance: processes for curating high-quality prompts/hints, version control, collaborative review, and reuse across instructors while preserving privacy are unspecified.
- Writing-authorship defense use-case: the impact of post-submission Socratic defenses on actual writing quality, originality, and equity (who benefits/loses) lacks empirical assessment.
- Operational constraints: recommended exam durations, concurrency limits, compute budgets per student, and failure-handling (crashes, connectivity drops) are not defined.
- Ethical transparency trade-offs: principled justification for maintaining a hidden Shadow Ledger vs student right to understand scoring rules needs ethical analysis and experimental evidence.
- Misconception diagnostics: how the system identifies, tracks, and remediates persistent misconceptions within and across sessions—and how that affects scoring—remains unspecified.
Practical Applications
Immediate Applications
Below are actionable uses that can be deployed now (with standard institutional approvals and basic integration).
- AI-mediated summative exams for university courses — Sector: education, software
- What: Replace or augment static tests with adaptive, typed Socratic Tests that map each student’s ZPD, use Bloom for prompting and SOLO for structural scoring, and enforce non-compensatory grading via capped buckets.
- Tools/workflows: LMS integration (Canvas/Moodle/Blackboard) for rosters/grade sync; instructor setup of topics/tiers; calibration UI for the 7-step grading pipeline; multimodal workspaces (whiteboard, code, calculator).
- Assumptions/dependencies: Instructor time for initial calibration; FERPA/GDPR-compliant data handling; reliable access to LLMs; institutional approval for summative AI use.
- Oral defense to restore authorship validity of take-home writing — Sector: education
- What: Pair essays/projects with a post-submission AI oral defense; the proctor interrogates arguments and sources; Out-of-Scope audit protects students from AI errors.
- Tools/workflows: Assignment hand-in → scheduled 15–20 min Socratic defense → deterministic transcript + appeal channel.
- Assumptions/dependencies: Clear course policy on AI use; instructor-created anchor prompts; stable source access for verification.
- Scalable, fair regrade/appeals handling — Sector: education, software
- What: Use interaction-level scores (Bloom b, hints k, SOLO s) to enable transparent, item-specific appeals; overrides propagate cohort-wide through auto-recalibration.
- Tools/workflows: Appeals queue UI; auto-regeneration of gradebook on override.
- Assumptions/dependencies: Faculty adoption of SOLO rubric; logging and audit trails enabled.
- Diagnostic telemetry to steer instruction — Sector: education, analytics
- What: Use evidence buffers and per-topic capped buckets to identify cohort-wide conceptual gaps and cognitive ceilings; inform subsequent lectures and targeted review.
- Tools/workflows: Instructor dashboard with topic/tier heatmaps; exportable reports.
- Assumptions/dependencies: Sufficient sample size per topic; instructors review data promptly to act on findings.
- Lower-anxiety alternatives to traditional oral exams — Sector: education, accessibility
- What: Typed, asynchronous conversation with multimodal workspaces reduces affective filter and language load; cognitive offloading supports ESL and neurodiverse students.
- Tools/workflows: Practice mode; accessible UI; optional time guidance without hard stops.
- Assumptions/dependencies: Accessibility review; accommodations policy alignment.
- Hiring and promotion skill screens — Sector: software, robotics, data/AI, finance, energy (L&D/HR)
- What: Adaptive, role-specific “Socratic screens” test conceptual mastery and problem solving (e.g., code reasoning, incident response, risk analysis) with non-compensatory buckets.
- Tools/workflows: ATS/HRIS integration; job-family taxonomies mapped to Bloom tiers; transcripts for hiring panels.
- Assumptions/dependencies: Legal review for bias/fairness; candidate privacy disclosures; job-relevant content authored by SMEs.
- Corporate compliance and safety attestations — Sector: energy, manufacturing, healthcare, finance (GRC)
- What: Replace passive e-learning quizzes with conversational checks; Shadow Ledger prevents skipping; audit trails support audits.
- Tools/workflows: Policy-to-prompt libraries; compliance dashboards; retraining triggers for low mastery.
- Assumptions/dependencies: Regulator acceptance of AI-mediated evidence; SOP-aligned item banks.
- MOOC/bootcamp certifications with mastery proof — Sector: edtech
- What: Issue micro-credentials only when topic-level buckets are filled at required tiers; calibration pipeline scales grading across large cohorts.
- Tools/workflows: Cohort-wide calibration and mass grading; badge issuance via LTI/LRS.
- Assumptions/dependencies: Platform integrations; psychometric monitoring across intakes.
- Proctored coding and engineering practicals — Sector: software, engineering
- What: Multimodal workspace captures code, derivations, and schematics; graded structurally via SOLO, not just final outputs.
- Tools/workflows: IDE integration; code execution sandboxes; artifact-linked prompts.
- Assumptions/dependencies: Secure compute; test data isolation; plagiarism controls for code.
- Institutional assessment and accreditation evidence — Sector: education, policy
- What: Exportable, criterion-referenced evidence aligned to intended learning outcomes (Bloom mapping, SOLO levels, ZPD scaffolding) for ABET/ACEN/discipline reviews.
- Tools/workflows: Program-level rollups; outcome-to-bucket traceability; longitudinal analytics.
- Assumptions/dependencies: Alignment between course outcomes and configured tiers; privacy-safe aggregation.
- AI risk management pattern for assessments — Sector: policy, software
- What: Deploy Out-of-Scope flags + post-exam audit + Global Exclusion Rules as a governance template for AI hallucination and fairness.
- Tools/workflows: Flagging UI; instructor audit flow; policy documentation pack.
- Assumptions/dependencies: Faculty training; service-level objectives for audits.
- Interview prep and self-study practice mode — Sector: daily life, education
- What: Low-stakes Socratic practice to rehearse defenses, receive graduated prompts, and build structural explanations.
- Tools/workflows: Personal dashboards; topic playlists; self-calibration samples.
- Assumptions/dependencies: Clear separation from summative instances; user consent for data use.
Long-Term Applications
These uses require further research, scaling, regulatory acceptance, or additional development.
- AI-mediated licensure and OSCE-style assessments — Sector: healthcare, legal, finance, engineering
- What: High-stakes adaptive oral/practical exams where conversation and multimodal evidence replace or augment static items.
- Potential: Virtual standardized patients; scenario-based risk adjudication; dynamic ethics evaluations.
- Dependencies: Multi-site validity studies; measurement invariance across subgroups; regulator approval; robust identity verification.
- Cross-institution calibration banks and psychometric networks — Sector: education, edtech
- What: Shared, SOLO-calibrated interaction corpora to standardize structural scoring across institutions while preserving local content.
- Potential: Item-level statistical dashboards; difficulty/easiness drift detection; benchmarked outcomes.
- Dependencies: Data-sharing agreements; privacy-preserving methods (federated learning, differential privacy); common metadata schemas.
- Lifelong, ZPD-driven micro-credentialing and skills passports — Sector: workforce, policy
- What: Continuous, topic-level mastery records that travel with learners/employees; non-compensatory badges for specific competencies.
- Potential: HRIS/ATS-native verification; portable “skills transcripts.”
- Dependencies: Standards alignment (e.g., Open Badges, CLR/LER); employer and accreditor buy-in; identity and consent frameworks.
- Advanced multimodal and XR practicals — Sector: robotics, advanced manufacturing, energy, STEM labs
- What: AR/VR lab simulations with conversational proctoring; whiteboard/handwriting, video, and sensor data as evidence.
- Potential: Safe rehearsal of hazardous procedures; equipment troubleshooting dialogues.
- Dependencies: Reliable multimodal LLMs; on-device inference or streaming; safety and validity studies for simulated tasks.
- Fairness auditing and policy frameworks for AI assessments — Sector: policy, education
- What: Mandated audit trails (Shadow Ledger, Evidence Buffers), subgroup fairness metrics, and appeal rights embedded in regulation and accreditation standards.
- Potential: Model/measurement audits akin to financial audits; third-party certification.
- Dependencies: Legal frameworks; standardized reporting; external auditors’ capacity.
- Robust identity assurance and anti-collusion at scale — Sector: education, certification, HR
- What: Seamless identity verification (eKYC, keystroke dynamics, voice) and collusion detection that does not elevate anxiety.
- Potential: Risk-based, adaptive proctoring that respects the affective filter principle.
- Dependencies: Privacy-preserving biometrics; accessibility considerations; societal acceptance.
- Federated calibration and privacy-by-design grading — Sector: software, policy
- What: Distributed calibration loops that tune SOLO scoring without centralizing raw transcripts.
- Potential: On-premise or sovereign-cloud deployments for sensitive sectors.
- Dependencies: Federated learning infrastructure; secure aggregation protocols.
- Instructor co-pilot and real-time oversight — Sector: education, software
- What: Dashboards that surface ambiguous interactions in-session; “human-in-the-loop” pivots before evidence buffers close.
- Potential: Reduced post-hoc appeals; improved student experience.
- Dependencies: UI ergonomics; cognitive load management for instructors; latency constraints.
- Sector-specific scenario libraries — Sector: finance (compliance), energy (HSE), public sector (procurement ethics)
- What: Conversational case banks aligned to regulations, with non-compensatory pass criteria and auditable trails.
- Potential: Annual attestations that test reasoning under policy constraints, not rote recall.
- Dependencies: Ongoing regulatory updates; SME governance; legal defensibility.
- Generalizing Shadow Ledger/Evidence Buffer patterns to training and QA — Sector: customer support, sales enablement, operations
- What: Use additive mastery and skip-resistant mechanics in onboarding and QA reviews to prevent “gaming by avoidance.”
- Potential: More reliable readiness checks; targeted remediation.
- Dependencies: Domain-specific constructs and rubrics; integration with performance systems.
- National or system-wide acceptance of AI-mediated assessment — Sector: policy, education
- What: Inclusion of AI conversational assessments in accreditation and funding decisions; guidance for procurement and risk.
- Potential: Cost-effective, scalable, high-validity alternatives to standardized tests.
- Dependencies: Longitudinal outcome studies; stakeholder engagement; public trust and transparency.
- Research-grade validity instruments and benchmarks — Sector: academia, edtech
- What: Open benchmarks for Bloom-anchored prompting and SOLO scoring; standardized protocols for ZPD measurement reliability.
- Potential: Comparable results across labs/institutions; accelerated method improvements.
- Dependencies: Community governance; dataset releases with ethical review.
Notes on feasibility across applications
- Technical: Requires reliable LLMs with guardrails, multimodal inputs, and low-latency serving; robust logging for auditability.
- Organizational: Faculty/SME training in Bloom/SOLO and calibration; content authoring time upfront; change management for stakeholders.
- Legal/ethical: Data protection (FERPA/GDPR), accessibility compliance, bias/fairness auditing, and documented human oversight.
- Evidence: For high-stakes use, needs multi-cohort validity, reliability, and subgroup fairness studies; regulator or accreditor endorsement.
Glossary
- Additive grading architecture: A scoring approach where students start from zero and accrue points for demonstrated evidence, rather than losing points for mistakes. "a non-compensatory, additive grading architecture that prioritizes mastery over penalty and human-AI alignment to ensure unprecedented measurement reliability."
- Affective Filter: A psychological barrier of anxiety or fear that impedes working memory and performance during assessment. "Face-to-face interrogations raise a student's Affective Filter, a psychological barrier of anxiety and fear of judgment that blocks working memory"
- Ambiguity Calibration: A calibration step where the most uncertain interactions are reviewed to anchor consistent scoring. "Step 2: Ambiguity Calibration."
- Bloom's Taxonomy: A hierarchical framework of cognitive objectives used here to structure the difficulty of prompts. "Bloom's Taxonomy (Prompt Objective)"
- Cognitive Ceiling: The highest level of performance a student demonstrates on a specific skill within the assessment. "effectively defining their Cognitive Ceiling for that discrete skill."
- Cognitive Offloading: Using external tools or workspaces to reduce working memory demands during problem solving. "which serve a critical pedagogical function known as Cognitive Offloading"
- Computer-Mediated Communication (CMC): Communication via digital interfaces that can reduce anxiety compared to face-to-face interaction. "Decades of research into Computer-Mediated Communication (CMC) demonstrate that typed interfaces significantly reduce communication apprehension and performative anxiety"
- Computerized Adaptive Testing (CAT): An assessment method where the test adapts to the test-taker, enabling standardized measurement without identical questions. "The theoretical precedent for dismissing this fallacy is firmly established in Computerized Adaptive Testing (CAT)"
- Compensatory Grading: A grading approach where strengths in some areas can mathematically offset weaknesses in others. "rejects Compensatory Grading (where low-level skills can mathematically compensate for a lack of high-level skills) in favor of Non-Compensatory Grading."
- Construct-irrelevant variance: Score variation caused by factors unrelated to the intended construct (e.g., anxiety or public speaking ability). "The genuine limitation is the introduction of significant construct-irrelevant variance."
- Constructive Alignment: Ensuring assessments and grading criteria directly align with intended learning outcomes. "Rooted in the theory of Constructive Alignment (i.e., the pedagogical principle that assessment tasks and grading criteria must strictly align with intended learning outcomes)"
- Contextual Mass Grading: Cohort-wide, post-exam grading performed with full conversational context after calibration. "Step 4: Contextual Mass Grading."
- Criterion-Referenced Grading: Evaluating performance against fixed standards rather than relative to other students. "True equity is achieved through criterion-referenced grading, where the specific conversational path varies, but the structural threshold required to prove mastery remains immutable"
- Deterministic gradebook: A grading record computed from explicitly defined variables without opaque or subjective adjustments. "the final deterministic gradebook."
- Dynamic Assessment (DA): An approach that integrates assessment with instruction to measure learning potential and responsiveness to support. "The Socratic Test is rooted in Dynamic Assessment (DA)"
- Evidence Buffer: Additional collected evidence beyond the required cap to protect students from provisional scoring errors. "Instead, the system utilizes an Evidence Buffer parameterized by an Oversampling Factor."
- Extended Abstract: The highest SOLO level where a response generalizes principles to novel domains. "Extended Abstract: The response generalizes the integrated principle to a completely novel domain."
- Global Exclusion Rule: A rule that automatically excludes a verified AI error for all students to ensure equity. "This verification automatically generates a ``Global Exclusion Rule'' within the grading engine."
- Graceful Exit: A controlled handoff where, after multiple hints fail, the AI supplies missing knowledge and pivots while logging the interaction. "Direct Instruction (Graceful Exit and Pivot): If the student exhausts the prior three hints, the AI explicitly provides the missing foundational knowledge"
- Hidden Curriculum: Unwritten norms and power dynamics in academia that can affect student performance. "the ``Hidden Curriculum''"
- Human-AI alignment: Systematic processes ensuring AI evaluations agree with instructor judgments. "and human-AI alignment to ensure unprecedented measurement reliability."
- Inter-Rater Reliability: The consistency of scoring across different human graders. "inter-rater and intra-rater reliability issues"
- Intra-Rater Reliability: The consistency of scoring by the same grader over time. "inter-rater and intra-rater reliability issues"
- Intrinsic Cognitive Load: The inherent complexity of material that strains working memory. "reducing intrinsic cognitive load"
- LLMs: Advanced AI models trained on vast text corpora that can act as conversational tutors or proctors. "LLMs as conversational tutors for formative feedback"
- Multimodal Evidence: An evidence trail combining text, diagrams, code, and other modalities to demonstrate understanding. "a significantly richer, multimodal evidence trail."
- Non-Compensatory Grading: A grading structure where mastery in each required area is necessary; strengths cannot offset deficits elsewhere. "in favor of Non-Compensatory Grading."
- Oversampling Factor: The percentage by which evidence collection exceeds the required cap to absorb potential downgrades during auditing. "an Evidence Buffer parameterized by an Oversampling Factor."
- Out-of-Scope Protocol: A procedure allowing students to flag invalid prompts for audit while keeping progression fair and tamper-resistant. "Mitigating Hallucination and the ``Out-of-Scope'' Protocol"
- Shadow Ledger: A hidden accounting of baseline points for skipped or evaded prompts used to manage progression and deter gaming. "The platform achieves this via a hidden Shadow Ledger of ``attempted points.''"
- SOLO Taxonomy: A framework for classifying the structural quality of learning outcomes, used here to evaluate responses. "the AI utilizes the SOLO Taxonomy (Structure of the Observed Learning Outcome) to measure the structural quality of the response"
- Stair-Step Mode: A structured proctoring configuration with predefined cognitive tiers and caps per topic. "Stair-Step Mode - Structured Scaffolding"
- Standardization Fallacy: The mistaken belief that identical questions are required for standardized measurement. "Standardization Fallacy, i.e., the misconception that standardizing the student's experience (identical questions) is required to standardize the measurement"
- Summative Assessment Engine: A platform intended for high-stakes, final evaluation rather than ongoing practice. "as a summative assessment engine."
- System of Least Prompts: A hierarchy of increasingly specific hints used to scaffold student responses. "System of Least Prompts (graduated prompting)"
- Technology Acceptance Model (TAM): A theory explaining technology adoption based on perceived usefulness and ease of use. "Technology Acceptance Model (TAM) and the Unified Theory of Acceptance and Use of Technology (UTAUT)"
- Unified Theory of Acceptance and Use of Technology (UTAUT): A comprehensive model of factors influencing technology adoption and use. "Technology Acceptance Model (TAM) and the Unified Theory of Acceptance and Use of Technology (UTAUT)"
- User-Directed Pivoting: Allowing students to switch topics during the exam without penalty while preserving state. "the platform abandons forced temporal constraints in favor of user-directed pivoting."
- Vertical Gate: The condition controlling progression to higher tiers based on accumulated earned and shadow points. "The Vertical Gate (i.e., the blocker to proceed to a higher tier) for a cognitive tier only opens when the sum of the student's Earned Points plus their Shadow Ledger Points meets the oversampled tier capacity."
- Vygotskian Scaffolding: Guided support grounded in Vygotsky’s theory to extend a learner’s performance beyond independent capability. "applying a non-compensatory mathematical model to Vygotskian scaffolding"
- Zone of Proximal Development (ZPD): The range between what a learner can do independently and with guidance. "Zone of Proximal Development (ZPD), i.e., the space between what a learner can do independently and what they can achieve with guidance"
Collections
Sign up for free to add this paper to one or more collections.