Multimodal Learning Analytics (MmLA)
- Multimodal Learning Analytics (MmLA) is an interdisciplinary field that integrates sensor data, AI, and analytics methods to capture, fuse, and interpret complex learning processes.
- It employs systematic pipelines to synchronize diverse signals like speech, video, and sensor data, transforming heterogeneous inputs into coherent feedback.
- Fusion strategies, including early, mid, and late fusion, enable real-time and personalized analysis that enhances outcomes in collaborative, clinical, and online learning environments.
Multimodal Learning Analytics (MmLA), often written as MMLA, is an interdisciplinary research field combining advances in the learning sciences, artificial intelligence, and sensor technologies to collect, fuse, analyze, and interpret diverse types of data—such as speech, video, sensors, and logs—to better understand and support learning and training experiences (Cohn et al., 2024). Across collaborative healthcare simulation, standard medical procedures, MOOCs, computer-based online learning, collaborative programming, and mixed-reality science activities, MmLA uses advanced sensing technologies and artificial intelligence to capture complex learning processes, while repeatedly confronting the challenge of integrating diverse data sources into cohesive and actionable insights (Yan et al., 2024).
1. Conceptual scope and field structure
A recurrent framing in the literature describes MmLA as a pipeline with four main stages: Learning/Training Environment, Multimodal Data, Learning Analytics Methods, and Feedback (Cohn et al., 2024). In this formulation, the environment may be physical, virtual, or blended; multimodal data are collected from multiple sources and transformed into modalities; learning analytics methods fuse, analyze, and interpret those modalities; and feedback comprises both direct feedback to learners or instructors and indirect feedback for researchers or system designers. This pipeline positions MmLA not merely as multimodal sensing, but as a feedback-oriented analytic enterprise.
A second organizing principle is the taxonomy of modality groups proposed in the systematic literature review on learning and training environments (Cohn et al., 2024). The review characterizes the domain in terms of five modality groups and emphasizes that multimodality often provides a more holistic understanding of behaviors and outcomes, even when predictive accuracy is not improved.
| Modality group | Example modalities | Example sources |
|---|---|---|
| Natural Language | prosodic speech, transcribed speech, raw text, audio spectrograms | audio, textual input, speech transcripts |
| Vision (Video) | pose, affect, gesture, activity, gaze, fatigue estimation, raw pixels | video, depth cameras, eye trackers |
| Sensors | EDA, pulse, EEG, EMG, temperature, blood pressure, fatigue, body pose, gaze | wearable or embedded biometric sensors, IMUs, eye trackers |
| Human-Centered | qualitative observations, participant-produced artifacts, researcher-produced artifacts, surveys, interviews | manual annotation, surveys, interviews, observational notes |
| Environment Logs | logs, screen recordings | computer-based learning platforms, virtual environments, educational apps |
This structuring has methodological consequences. It supplies a shared vocabulary for reporting and comparing studies, and it foregrounds that MmLA includes not only physiological and behavioral sensing, but also logs, artifacts, surveys, and interviews when these are integrated into analytic workflows (Cohn et al., 2024).
2. Data capture, synchronization, and feature construction
MmLA studies are defined as much by synchronization and representation as by sensing. Representative implementations collect heterogeneous signals such as x-y positioning via UWB sensors, audio from wireless headset microphones, and heart rate from wrist-worn sensors; derive behavioral indicators from each modality; and then align them on a common temporal basis before analysis (Yan et al., 2024). In the collaborative healthcare simulation study that integrated latent class analysis, 17 monomodal indicators were extracted from positional, audio, and physiological data and synchronized into 60-second intervals, with each indicator binarized as presence or absence within the interval (Yan et al., 2024).
In standard medical procedure analysis, synchronization is carried out at higher temporal resolution. The ABCDE nursing study combined wearable accelerometers on both wrists, BLE proximity estimates, and gaze from Tobii Pro Glasses 3; all devices were time-synchronized via a Raspberry Pi hub using Lab Streaming Layer, achieving a system latency of about 50 ms (Heilala et al., 2023). Feature extraction then produced hand movement velocities, discretized proximity states, and gaze entropy over 5-second sliding windows, which were rendered together in a behaviorgram using dense pixel technique and dimensional stacking (Heilala et al., 2023). This integrated visual representation made procedural phases visible as aligned multimodal patterns rather than isolated signal traces.
Web-based infrastructures adopt analogous principles. M2LADS standardizes timestamps across biosignals, videos, and activity logs; maps data to granular learning activities; computes smoothed data using a 30-second sliding window; and stores processed multimodal data in MongoDB for dashboard-based inspection (Becerra et al., 21 Feb 2025). In the earlier open-education version of M2LADS, multimodal fusion was organized through activity matrices, variable data matrices, and a learner matrix that aligned biometric, behavioral, background, and performance data by timestamp and activity identifier (Becerra et al., 2023).
These pipelines show that temporal standardization is not ancillary. In MmLA, the operational meaning of a “learning event” often depends on how streams with different sampling rates, latencies, and semantics are synchronized, discretized, smoothed, and mapped to task structure.
3. Fusion strategies and analytic methods
A central problem in MmLA is where and how fusion occurs. Reviews distinguish feature-level / early fusion, decision-level / late fusion, and hybrid fusion, and one recent review adds mid fusion as the integration of processed but still observable features between raw input and final decision or hypothesis spaces (Chango et al., 25 Nov 2025). In formal terms, early fusion is often written as
where features from several modalities are concatenated into a single heterogeneous vector, while late fusion combines modality-specific decisions, for example
after separate modeling of each modality (Chango et al., 25 Nov 2025). The introduction of mid fusion is significant because many educational sensing pipelines operate on processed observables—such as joint positions, gaze tracks, or pose estimates—that are neither raw sensor streams nor high-level inferred constructs (Cohn et al., 2024).
MmLA employs both exploratory and predictive analytic techniques. A person-centered example is latent class analysis (LCA), used to identify homogeneous multimodal behavior patterns from binary monomodal indicators (Yan et al., 2024). In the healthcare simulation study, the standard LCA likelihood was given as
with models from 1–10 classes compared using the lowest Bayesian Information Criterion (BIC) and highest log-likelihood (Yan et al., 2024). Once a class solution is selected, each learner-interval is assigned to the most probable latent class, converting high-dimensional monomodal streams into a sequence of multimodal class labels.
Other analytic families recur across the literature. Epistemic Network Analysis (ENA) is used to compare multimodal indicators with monomodal ones and to visualize co-occurrences among communication behaviors (Yan et al., 2024). Process mining and sequential pattern analysis are proposed in mobile multimodal learning analytics for self-regulated learning (Khalil, 2020). Random Forest and Support Vector Machine models, often combined with smoothing, PCA, or SelectKBest, appear in online-learning distraction detection (Becerra et al., 20 Jun 2025). Visualization-oriented methods such as behaviorgrams, ward maps, sociograms, and timeline views are used not as mere presentation layers, but as analytic instruments for linking multimodal traces to procedural, collaborative, or reflective interpretation (Heilala et al., 2023).
4. Collaborative, clinical, and embodied applications
Collaborative healthcare simulation has become a prominent MmLA setting because it combines movement, speech, spatial coordination, physiology, and post-activity reflection. In the LCA-based study, 17 monomodal indicators were reduced to four multimodal latent classes—Collaborative Communication, Embodied Collaboration, Distant Interaction, and Solitary Engagement—each representing a distinct co-occurrence pattern across task prioritization, team communication, and physiology (Yan et al., 2024). ENA then showed that the four multimodal indicators were more parsimonious and explained almost twice the target variance along key axes, including 17.5% vs 9.3% for task satisfaction and 15.6% vs 8.5% for collaboration satisfaction (Yan et al., 2024). The result is important because it demonstrates that MmLA can move from many fine-grained monomodal variables to a smaller set of interpretable multimodal constructs without sacrificing explanatory power.
In procedural medical education, MmLA has been used to decompose expert performance into phases. The ABCDE study linked gaze entropy, hand movement velocity, and proximity measures to four main phases—Preparation, Breathing assessment, Circulation/Disability/Exposure assessment, and Review/check-listing—and argued that the fused view could distinguish expert behaviors and spot potential inefficiencies or errors (Heilala et al., 2023). Here, the contribution of MmLA is the joint interpretation of manual, visual, and spatial signals as a procedural signature.
Reflection-oriented systems extend these ideas into debriefing practice. TeamVision captures voice presence, automated transcriptions, body rotation, and positioning data, and presents ward maps, speech sociograms, communication networks, and indexed video snippets to guide debriefs immediately after simulation (Echeverria et al., 17 Jan 2025). In an in-the-wild study with 56 teams (221 students) and debriefs led by six teachers, educators reported that the system supported flexible, evidence-based discussion, while interviews with 15 students and five teachers highlighted usefulness alongside concerns about nuance, trust, and occasional mismatches between visualizations and observed behavior (Echeverria et al., 17 Jan 2025).
MmLA has also been applied to collaborative programming. CPVis integrates final and intermediate code, screen recordings, video recordings of discussions, transcribed dialogue, speaker diarization, and instructor interventions, then uses a flower-based visual encoding and time-based views to support group and individual assessment (Zhang et al., 25 Feb 2025). In a within-subject experiment with 22 participants, users gained more insights, found the visualization more intuitive, and reported increased confidence in their assessments of collaboration (Zhang et al., 25 Feb 2025).
Embodied mixed-reality learning presents a related but distinct use case. A timeline-based approach for photosynthesis learning combines movement, gaze, affect, and system logs to support interaction analysis, allowing researchers to align critical learning moments identified by machine learning with those identified by qualitative analysis (Fonteles et al., 2024). This suggests a hybrid MmLA role in which automated multimodal coding does not replace interpretation but reorients attention toward temporally salient events.
5. Dashboards, datasets, and online-learning infrastructures
A major branch of MmLA focuses on systems that integrate, store, visualize, and inspect multimodal data rather than only fitting models. M2LADS is a web-based platform with three modules—Activity Data Processing, Activity Data Management, and Activity Data Visualization—designed to integrate, synchronize, visualize, and analyze multimodal data recorded during computer-based learning sessions with biosensors (Becerra et al., 21 Feb 2025). It ingests EEG, heart rate, eye-tracking, webcam video, activity logs, demographic and medical background information, and optionally pre/posttests; organizes all signals by activity; supports synchronized replay of biosignals and videos; and provides trend analysis, correlation analysis, comparative analytics, and relabeling support when activity labels are incorrect (Becerra et al., 21 Feb 2025). The same system lineage was previously presented for MOOC and UX contexts, where it was used to capture learners’ holistic experience and to compare attention, gaze, heart rate, and interaction patterns across dashboard activities (Becerra et al., 2023).
Publicly documented datasets are comparatively scarce, which makes MUTLA notable. It provides time-synchronized learning logs, EEG brainwaves, and webcam video from 156 students working in authentic educational activities on the Squirrel AI Learning System (Xu et al., 2019). The synchronized resource includes 2170 video segments, with 61% yielding at least 50% valid facial tracking, and was explicitly positioned as a real-world alternative to controlled laboratory corpora (Xu et al., 2019). In MmLA, such datasets serve not only predictive benchmarking but also event-level alignment between educational tasks and multimodal evidence.
Online learning has also motivated targeted detection systems. The smartphone-distraction study used multimodal biometrics from the IMPROVE database—head pose from webcam video, heart rate, and EEG bands—to classify Phone Use versus No Phone Use in a balanced dataset of 132 windows from 66 participants (33 phone users, 33 non-users) (Becerra et al., 20 Jun 2025). The best unimodal physiological models achieved 61–70% accuracy, Head Pose alone achieved 87%, and the multimodal EEG+HR+Head Pose model reached 91% accuracy (Becerra et al., 20 Jun 2025). These results do not imply that physiology is unnecessary; rather, they show that non-intrusive signals may dominate certain tasks while multimodal integration can still yield the best performance.
Scalability-oriented infrastructures are beginning to address persistent deployment bottlenecks. Watch-DMLT and ViSeDOPS were introduced for real-time smartwatch acquisition and synchronized dashboard visualization in classroom presentations, with a deployment involving 65 students and up to 16 smartwatches in parallel (Becerra et al., 2 Dec 2025). BadgeX extends the infrastructure agenda by combining smart badges or smartphones with LLM-based interpretation; in a pilot with 2 individuals during a 43-minute STEM collaborative problem-solving session, action recognition matched manual annotation in 161/176 (90%) cases, and LLM interpretation required about 10 seconds per time bucket (Li et al., 5 Apr 2026). This suggests a shift from post-hoc dashboarding toward narrative, theory-aligned interpretation layers built on synchronized multimodal traces.
6. Ethics, deployment, misconceptions, and open problems
The ethical and sociotechnical literature on MmLA has increasingly emphasized that fairness, accountability, transparency, and ethics cannot be reduced to model metrics alone. In an authentic collaborative learning context, interviews with 14 undergraduate students showed that fairness was tied to accurate and comprehensive data representation, accountability to differentiated data access and responsibility for misuse, transparency to understanding what was collected and how it was processed, and ethics to informed consent as an ongoing process rather than a checkbox (Jin et al., 2024). The study recommended tiered, role-based access, stronger transparency practices, and a move from dichotomous consent models toward continuous, measurable scales of understanding (Jin et al., 2024).
In-the-wild deployment studies expose the practical stakes of these concerns. A 2-year human-centred MMLA deployment with 399 students and 17 educators in nursing simulation synthesized lessons across technological and physical deployment, multimodal data and interfaces, design process, participation and privacy, and sustainability (Martinez-Maldonado et al., 2023). Off-the-shelf sensors were easy to source but not easy to maintain or synchronize; incomplete or noisy data could still be over-trusted because teachers and students associated sensor outputs with “objectivity”; and successful use required teacher upskilling in both technical operation and interpretation (Martinez-Maldonado et al., 2023). Recommendations included modular architectures, simpler classroom integration, clearer consent procedures, and teacher agency over analytics rather than automation.
Several common misconceptions are directly addressed in the literature. One is that adding more modalities necessarily improves prediction. Systematic reviews state instead that multiple modalities often offer a more holistic understanding of learner behavior, and even when multimodality does not enhance predictive accuracy, it can contextualize and elucidate unimodal data (Cohn et al., 2024). Another misconception is that sensor-rich systems are self-evidently objective or fair; qualitative studies show that missingness, role bias, decontextualized visualizations, and misunderstood pipelines can distort interpretation unless limitations are made visible and discussed (Jin et al., 2024). A third misconception is that technical feasibility alone is sufficient for educational value; reviews emphasize limited scalability, scarcity of public datasets, data quality and synchronization challenges, insufficient theoretical grounding, and the need for real-world longitudinal validation (Becerra et al., 9 Sep 2025).
Open problems remain concentrated around standardization, generalizability, and integration. Reviews call for larger and more diverse samples, open data and protocol standardization, better handling of missing values and heterogeneous temporal structures, more advanced and hybrid fusion methods, stronger stakeholder involvement, and closer links between multimodal learning and training studies and foundational AI research (Cohn et al., 2024). The field also continues to expand from student-only sensing toward richer configurations that may incorporate teacher behavior, environmental data, and LLM-mediated interpretation, but these developments intensify rather than resolve longstanding questions of trust, privacy, and pedagogical usefulness (Chango et al., 25 Nov 2025).
Taken together, these studies suggest that MmLA is best understood as a family of feedback-oriented analytic practices built on synchronized multimodal evidence rather than as a single modeling paradigm. Its distinctive contributions lie in parsimonious multimodal representations, interpretable visual analytics, and context-sensitive integration of behavioral, physiological, spatial, linguistic, and log data. Its most persistent challenges concern not only prediction, but also alignment, transparency, fairness, sustainability, and the translation of rich traces into educationally meaningful action.