Papers
Topics
Authors
Recent
Search
2000 character limit reached

AMI Corpus Definition: Augmented Multi-party Interaction Learning Dataset

Updated 7 September 2026
  • You can use AMI as a set of corpus-specific research oriented resources offering multimodal indoors speech and audio use, observable equipment dependent data, naturalized immersion process and audio-event detecting resources.
  • A widely accepted Augmented Multi-party Interaction (AMI) corpus is a meticulously collected, annotated, and reestablished website.
  • The Arcminute Microkelvin Imager’s (AMI) astronomy corpus provides observable measures in Galactic diffuse emissions, young stellar bodies, and galaxy clusters.

The term AMI Corpus is used for several distinct research resources and corpus-oriented bodies of work. In speech and multimodal interaction research, it most commonly denotes the Augmented Multi-party Interaction (AMI) corpus, a publicly available collection of naturally occurring indoor office meetings with synchronized audio recordings, multiple microphone channels, and annotations of speech and acoustic events. In energy research, AMI denotes Advanced Metering Infrastructure, whose data comprise fine-grained electricity or gas measurements and associated operational metadata. In radio astronomy, AMI refers to the Arcminute Microkelvin Imager, and an “AMI corpus” may denote collections of calibrated interferometric observations, visibility data, maps, source catalogues, and derived astrophysical parameters. Ambient-intelligence research uses AmI for distributed, context-aware multi-agent environments, while several papers develop simulated or annotated resources relevant to such systems. These meanings are technically unrelated and should be disambiguated by domain and expansion of the acronym.

1. Terminology and domain disambiguation

The most established corpus use of AMI in language and audio research is the Augmented Multi-party Interaction corpus. It contains recordings of indoor meetings held in an office environment and includes recordings from omnidirectional microphone arrays, headset microphones, and lapel microphones. The corpus is accompanied by annotations for multiple audio events, including coughs, and is publicly available under a Creative Commons Attribution 4.0 license. A re-annotation study corrected cough-event boundaries and separated successive coughs, producing 1,369 individual cough events from 1,116 original annotations (Leamy et al., 2019).

In utility engineering, AMI means Advanced Metering Infrastructure: a bidirectional communications infrastructure connecting smart meters, utility systems, customer equipment, and data-management systems. AMI data are typically time-indexed energy-consumption measurements, often collected at 15–60-minute intervals. The corpus may include electricity or gas readings, feeder and neighborhood aggregates, customer or premises metadata, weather and calendar variables, tariff information, and program-participation attributes. Because high-resolution longitudinal consumption can reveal occupancy, appliance use, sleep schedules, work patterns, vacations, electric-vehicle ownership, and health-related equipment use, AMI data are privacy-sensitive (Westrich, 13 May 2025).

In radio astronomy, AMI is the Arcminute Microkelvin Imager, a pair of interferometric arrays operating near centimetre wavelengths. The AMI research corpus consists not of a single standardized dataset but of observational products: calibrated complex visibilities, frequency-channel data, synthesized images, source catalogues, calibration metadata, radio light curves, and Bayesian posterior products. AMI observations support studies of galaxy-cluster Sunyaev–Zel’dovich effects, Galactic diffuse emission, young stellar objects, supernovae, radio surveys, and transients (Grainge et al., 2012).

In artificial intelligence, AmI means Ambient Intelligence. AmI systems embed sensing, networking, autonomous or semi-autonomous agents, context recognition, personalization, adaptation, and proactive assistance in everyday environments. Corpus-oriented AmI research may consist of interaction traces, simulation episodes, ontological representations, sensor-context records, agent messages, goals, plans, and actions. An airport simulation in NetLogo, for example, models agents, profiles, services, context, FIPA-like messages, BDI intentions, movement, queues, and satisfaction outcomes (Carbo et al., 2024).

These meanings should not be conflated. An AMI audio corpus is a multimicrophone meeting resource; an AMI smart-meter corpus is a privacy-sensitive longitudinal utility dataset; an AMI astronomical corpus is an interferometric observational archive; and an AmI corpus is a resource for context-aware intelligent environments.

2. The Augmented Multi-party Interaction corpus

The Augmented Multi-party Interaction corpus contains naturally occurring office meetings recorded through multiple microphone configurations. The available channels include omnidirectional microphone arrays in near and far positions, headset microphones, and lapel microphones. This arrangement exposes the same acoustic event to different distances, reverberation conditions, signal-to-noise ratios, and background-noise environments.

The corpus supports research on multimodal interaction, speech processing, speaker and event analysis, and audio-event detection. Its naturally occurring meeting-room recordings differ from laboratory or deliberately elicited datasets because events occur in conversational environments with overlapping speech, room acoustics, and non-ideal microphone conditions. The corpus includes annotations for coughs and other audio events.

The original cough annotations contained 1,116 annotations. Inspection identified incorrect start and end times, implausible durations, and successive coughs grouped into single annotations. A re-annotation procedure used a purpose-built MATLAB graphical user interface to import audio, inspect recordings, and mark the onset and end of individual coughs. The revised resource contains 1,369 individual cough events, an increase of 253 annotations, or approximately 22.7% relative to the original count (Leamy et al., 2019).

The re-annotation tool and revised cough locations were made available for public use. The annotation unit is a single cough event represented by onset and offset timestamps. The study does not report inter-annotator agreement, repeated annotation trials, boundary-error statistics, precision, recall, F-score, annotator numbers, or a formal adjudication procedure. The revised labels are therefore manually improved annotations rather than a quantitatively validated clinical ground truth.

The cough annotations are intended for audio-event detection, machine-learning development, evaluation, and comparison. They may support onset and offset localization, event-level or frame-level evaluation, robustness analysis under speech and office noise, and comparisons across microphone channels. The recordings are not presented as medically diagnosed examples and should not be treated as a clinical cough-diagnosis corpus. Cough events identify acoustic phenomena, not diseases or health conditions.

3. Advanced Metering Infrastructure data

An Advanced Metering Infrastructure corpus contains high-resolution, time-indexed measurements from electricity or gas meters together with metadata and derived variables. A representative electricity series may be written as

xi=(xi,1,xi,2,…,xi,T),x_i=(x_{i,1},x_{i,2},\ldots,x_{i,T}),

where xi,tx_{i,t} is the usage of customer ii during interval tt. The representative granularity discussed in the literature is 15-minute sampling, corresponding to 96 electricity observations per day. Coarser hourly, daily, weekly, and aggregated representations are also used.

Potential corpus fields include meter or account identifiers, customer or premise attributes, generalized location such as ZIP code, customer type, tariff or time-of-use status, participation in solar, electric-vehicle, or demand-response programs, building characteristics, weather and calendar variables, feeder or substation assignment, and data-quality information. These fields increase analytical value but also create quasi-identification risks. A distinctive load pattern combined with generalized location, solar participation, or a known work schedule can identify a premise even after direct identifiers are removed (Westrich, 13 May 2025).

AMI data support short- and long-term load forecasting, feeder and system peak prediction, demand-response evaluation, renewable and distributed-energy integration, outage and anomaly detection, grid planning, billing, tariff analysis, customer segmentation, appliance or end-use disaggregation, and econometric studies. They can also support correlations among usage, weather, tariffs, solar generation, electric-vehicle charging, and demand-response interventions.

The scale of AMI data creates computational constraints. A study using 323 customers and 5,208 hourly measurements per customer contains 1,682,184 AMI records. Wavelet representations can compress individual customer profiles and further compress them by grouping customers with similar load patterns. A level-3 DB1 wavelet approximation represents each 5,208-point customer profile with 651 coefficients, producing an 87.5% reduction in stored values. A classified wavelet model based on two-dimensional wavelet coefficients, k-means clustering, and five typical load profiles produced a reported 98.41% reduction (Zhong et al., 2015).

The individual-customer wavelet model retained 90–99.998% of the original signal energy for 93.81% of profiles, while 74% of synthesized hourly loads were within 10% of the original measurements. The classified model retained total energy within 10% for 72% of profiles. In a time-series power-flow experiment, 71% of hourly power-flow estimates from the individual wavelet model and 80% from the classified model were within 10% of the original-AMI results. Aggregate phase-energy errors were at most 0.26%, although individual hourly and peak estimates could be substantially less accurate (Zhong et al., 2015).

AMI communications research addresses the transmission of large numbers of meter readings. A grouped hierarchical architecture places meters in local groups managed by data concentrators; the concentrators communicate with LTE eNodeBs, which forward data to utility control systems. Each concentrator sends an aggregate total-consumption message together with scheduled individual meter readings. Analytical evaluations report lower LTE resource consumption than a flat architecture in which every meter communicates directly over LTE. The approach trades reduced radio contention and fewer LTE modules for additional concentrator infrastructure, aggregation delay, local-network dependence, and possible concentration of failures (Elmesalawy et al., 2016).

Privacy-preserving AMI corpora require protection across the full data lifecycle. Techniques discussed in the literature include pseudonymization and aggregation, differential privacy, federated learning, synthetic-data generation, secure multiparty computation, and homomorphic encryption. Differential privacy provides a formal guarantee of the form

Pr⁡[M(D)∈O]≤eεPr⁡[M(D′)∈O],\Pr[M(D)\in O]\leq e^\varepsilon\Pr[M(D')\in O],

for neighboring datasets DD and D′D'. However, longitudinal AMI data require household-level sensitivity definitions and privacy-budget accounting because repeated releases accumulate privacy loss. A privacy-preserving architecture may separate identifiers from usage data, restrict raw-data access to an internal sandbox, expose differential-privacy APIs, generate synthetic research data, coordinate federated learning, and audit releases and privacy budgets (Westrich, 13 May 2025).

The EPIC protocol provides a different security-oriented approach. Each meter sends a masked reading, with pairwise masks designed to cancel after aggregation. Homomorphic hashes, HMACs, signatures, timestamps, and batch verification provide integrity and authenticity while limiting disclosure of individual readings. The protocol supports single-hop and multi-hop AMI networks, attacker identification, utility–relay collusion analysis, and dynamic-pricing billing. Its privacy guarantees depend on distributed proxy selection, secure key storage, correct mask synchronization, cryptographic assumptions, and the condition that colluding parties do not obtain all masks protecting a victim (Alsharif et al., 2018).

4. The Arcminute Microkelvin Imager observational corpus

The Arcminute Microkelvin Imager consists of two interferometric arrays. The Small Array has ten 3.7-m antennas, compact baselines, approximately 3-arcminute resolution, and sensitivity to extended low-surface-brightness structures. The Large Array has eight 13-m antennas, longer baselines, approximately 25–30-arcsecond resolution, and approximately ten times the collecting area and flux sensitivity of the Small Array. The arrays operate over approximately 13.5–18 GHz and are used together to measure diffuse emission and identify compact contaminating sources (Grainge et al., 2012).

An AMI observational corpus includes calibrated complex visibilities, source-subtracted and unsubtracted maps, channel spectra, calibration records, observing dates, synthesized beams, noise estimates, source positions, flux densities, spectral indices, and posterior distributions from Bayesian inference. AMI measurements are fundamentally samples of the Fourier transform of primary-beam-weighted sky brightness. Short baselines probe larger angular scales, whereas long baselines provide higher angular resolution.

For galaxy clusters, the Small Array measures arcminute-scale thermal SZ decrements while the Large Array identifies radio sources that can fill in or distort those decrements. Bayesian analyses compare cluster-plus-source models with source-only models, often using nested-sampling algorithms such as MultiNest or McAdam. The resulting corpus may contain cluster masses, pressure profiles, gas fractions, temperatures, integrated Comptonization, detection evidence ratios, source priors, and model-comparison results.

AMI observations of 15 hot XMM-Newton Cluster Survey clusters detected three clusters at high significance, two at lower significance, and produced null results for ten. The AMI-derived large-scale mean temperatures were systematically lower than X-ray core-temperature estimates: by a factor of approximately 1.4 for SZ detections and an average 68% upper-limit factor of approximately 1.9 for non-detections. The interpretation involves differences between X-ray density weighting, SZ pressure weighting, cluster mergers, profile assumptions, gas-fraction priors, redshift uncertainty, source contamination, and selection effects (Consortium et al., 2013).

The AMI–CARMA analysis of AMI-CL J0300+2613 illustrates the importance of visibility-domain analysis and complementary uvuv coverage. AMI provided a strong detection, while CARMA measured a weaker decrement than expected from simple frequency scaling. The joint analysis gave a total mass within r200r_{200} of (4.1±1.1)×1014 M⊙(4.1\pm1.1)\times10^{14}\,M_\odot and a phenomenological central decrement of xi,tx_{i,t}0. The discrepancy was attributed plausibly to extended, unusual cluster morphology and different sensitivity to angular scales, although the morphology explanation was not quantitatively proven (Consortium et al., 2013).

Joint AMI–Planck analyses address the degeneracy between integrated SZ signal and angular pressure scale. AMI supplies higher angular resolution and intermediate-scale structure, whereas Planck supplies multifrequency information and sensitivity to larger-scale integrated emission. Simulations show that a joint analysis can accurately recover xi,tx_{i,t}1 and angular size when the assumed pressure profile is correct. If the pressure profile is incorrectly fixed, the joint posterior can be precise but biased. Allowing profile-shape parameters to vary generally gives more robust integrated signals and can resolve apparent AMI–Planck discrepancies, as demonstrated for PSZ2 G063.80+11.42 (Perrott et al., 2019).

AMI also provides radio-continuum observations of young stellar objects, supernovae, and transients. In a sample of low-mass protostars with known outflows, 16-GHz measurements and radio-to-submillimetre spectral energy distributions showed that approximately 80% of sources had spectral indices consistent with free-free emission. The study recovered a radio-luminosity/envelope-mass relation with a slope approximately 1.0 and found evidence for radio variability in L1551 IRS 5, Serpens MMS 1, and HH 1 (Consortium et al., 2012).

The AMI Large Array monitoring of SN 2014C produced an 82-observation, 15.7-GHz light curve spanning approximately 17–567 days after first light. It showed two radio peaks, with the second approximately four times more luminous than the first. The morphology was interpreted as evidence for an initially lower-density circumstellar environment followed by interaction with a dense hydrogen-rich shell. Single-frequency data could not distinguish synchrotron self-absorption from free-free absorption, and the physical parameters inferred from those models remained model-dependent (Anderson et al., 2016).

5. Ambient Intelligence corpora and simulated interaction resources

Ambient Intelligence environments combine embedded computing, pervasive sensing and networking, autonomous or semi-autonomous agents, context awareness, personalization, adaptation, anticipation, and minimal explicit user intervention. A typical agent contains a knowledge base, goals, planning capabilities, and communication mechanisms. The reasoning cycle is commonly represented as

xi,tx_{i,t}2

An AmI corpus can therefore represent sensor observations, contextual interpretations, domain and background knowledge, user and agent goals, plans, actions, messages, and outcomes.

Conflict-oriented AmI research distinguishes five knowledge-based conflict types: sensory-input conflicts, contextual conflicts, domain and background knowledge conflicts, goal conflicts, and action conflicts. The proposed ordering is

xi,tx_{i,t}3

Sensory conflicts concern incompatible observations. Contextual conflicts concern incompatible interpretations of a situation. Domain and background conflicts concern inconsistent relatively stable knowledge. Goal conflicts arise when agents share a compatible world model but pursue incompatible objectives. Action conflicts arise when agents share compatible world models and goals but choose mutually interfering actions. The paper is conceptual and does not provide an implementation, benchmark dataset, formal proof, or quantitative evaluation (Homola et al., 2014).

An airport simulation implemented in NetLogo provides a structured synthetic AmI scenario. User agents move through entrance areas, flight-information panels, check-in counters, passport controls, shops, boarding-information panels, boarding gates, baggage belts, and exits. The architecture includes User, Provider, Facilitator, Positioning, and Evaluator Agents. It uses FIPA-like inform and request messages, a lightweight BDI representation, an ontology involving positions, places, services, products, features, contexts, and profiles, and a 12-step service protocol (Carbo et al., 2024).

The simulation compares AmI and non-AmI behavior. AmI users receive contextual information about locations, providers, gates, counters, baggage, and shops, while non-AmI users physically seek information and may spend more time moving or queuing. The reported reference experiments used 30 NetLogo runs and a 50%/50% AmI versus non-AmI population split. The authors report approximately 18% time improvement and 40% satisfaction improvement for AmI users. The exact parameterization, satisfaction weights, random seeds, and raw traces are not fully documented, so the results are simulation outcomes rather than validated measurements of real airports.

A corpus generated from this model could encode each interaction episode as a sequence containing user profile, context, beliefs, desires, intentions, messages, movement, service results, queue times, and satisfaction outcomes. Such a corpus must distinguish simulated records from observed human data. The paper does not itself release a conventional human-interaction corpus.

6. Corpus construction, annotation, privacy, and reproducibility

An AMI corpus should preserve domain-specific provenance. For the Augmented Multi-party Interaction corpus, essential metadata include meeting identifiers, microphone type and position, recording channels, event timestamps, annotation version, and event category. For cough annotations, onset and offset conventions, microphone dependencies, and the relationship between multiple recordings of the same meeting are particularly important. Train–test separation should avoid leakage across meetings, speakers, and correlated microphone channels.

For Advanced Metering Infrastructure, provenance should include measurement units, sampling interval, time zone, meter and feeder identifiers, phase association, customer or premise linkage, raw-versus-cleaned status, missingness, meter resets, imputation, outlier handling, tariff information, and aggregation level. The wavelet-load study illustrates the importance of documenting the exact wavelet family, decomposition level, retained coefficients, weekly alignment, normalization, clustering settings, and reconstruction formula (Zhong et al., 2015).

For Arcminute Microkelvin Imager observations, corpus records should distinguish raw visibilities, calibrated visibilities, flagging states, frequency channels, array configuration, baseline coverage, primary-beam corrections, synthesized beams, calibration sources, imaging weights, source-subtraction models, observing epochs, and astrophysical model assumptions. AMI’s digital correlator upgrade replaced an analogue lag-based correlator with real-time digital FX correlators using 1.2–1.22-MHz spectral channels. The upgrade reduced low-declination RFI flagging from approximately 85–90% with the old system to approximately 20% for the Large Array and 25–35% for the Small Array near zero declination, while improving typical imaging dynamic range from approximately 100 to approximately 1,000 (Hickish et al., 2017).

For Ambient Intelligence, corpus schemas should encode the distinction between raw sensory input, inferred context, stable knowledge, goals, plans, actions, communication acts, and outcomes. Source, intervenients, detection time, solvability, agent role, user role, and conflict type are useful annotation dimensions. A corpus should identify whether an episode is observed, manually annotated, simulated, or generated by a protocol.

Privacy and governance are central for AMI smart-meter corpora and relevant to audio and AmI resources as well. Direct identifiers should be separated from measurements, access should be role-controlled, and releases should be purpose-specific. Differential privacy is appropriate for many public aggregate statistics; synthetic data can provide a research interface but must be tested for memorization and membership inference; federated learning supports collaborative model training but does not itself guarantee privacy; secure multiparty computation and homomorphic encryption support exact or encrypted computations but introduce cryptographic and operational overhead (Westrich, 13 May 2025).

No single protection mechanism addresses all risks. Advanced Metering Infrastructure privacy depends on household-level sensitivity, repeated-release composition, linkage across electricity, gas, weather, tariff, and program data, and the possibility that rare load patterns act as behavioral fingerprints. Protocols such as EPIC provide integrity and authenticity in addition to privacy, but remain dependent on key management, proxy distribution, timestamp synchronization, collusion thresholds, and secure initialization (Alsharif et al., 2018).

The term AMI Corpus therefore denotes a family of domain-specific resources rather than one universal collection. Its correct interpretation requires identifying whether AMI refers to augmented meeting interaction, advanced metering infrastructure, the Arcminute Microkelvin Imager, or ambient intelligence. Across these domains, the common corpus requirements are explicit provenance, precise annotation or measurement semantics, separation of observed and simulated data, reproducible processing metadata, and privacy or security controls appropriate to the sensitivity of the underlying records.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AMI Corpus.