---
title: 'SIMORD: Open Dataset for Medical Order Extraction'
url: https://www.emergentmind.com/topics/simord
type: topic
---

# SIMORD: Open Dataset for Medical Order Extraction

Searching arXiv for the specified paper and related clinical NLP datasets to ground the article.
SIMORD (SIMulated ORDer) is the first open-source dataset for extracting medical orders from doctor-patient simulated transcripts. It is introduced to support the medical order extraction task, which consists of identifying and structuring orders discussed during consultations, including medications, laboratory tests, imaging, and follow-up actions. In the source study, SIMORD is presented as a resource for benchmarking large language models on a clinically consequential documentation problem that remains underexplored because of data scarcity and sensitivity; the stated motivation is to automate documentation, improve clinical workflows, and reduce the burden on healthcare providers [2507.05517].

## 1. Definition and task formulation

SIMORD is designed around long-form doctor-patient conversation transcripts. Each instance is annotated at the order level rather than only at the span or utterance level. For every order, the dataset records four attributes: a **Description**, defined as a concise summary of the order; an optional **Reason**, defined as the motivation or diagnosis for the order; a **Type**, drawn from the categorical label set `"medication"`, `"laboratory"`, `"followup"`, and `"imaging"`; and **Provenance**, defined as the list of conversation line numbers where the order is mentioned [2507.05517].

The required output format for both human annotation and language-model prediction is a list of standardized JSON objects, one object per order. This design makes the task explicitly structured rather than purely extractive. It also embeds traceability through line-number provenance, which is intended to support linkage between the structured output and the conversational evidence.

The framing of the task is operational rather than merely semantic. Annotators are instructed to assess every medical order within the conversation “the way a doctor would create them in the EHR,” and the dataset aims to replicate the doctor’s usual end-of-encounter order-entry process. This suggests that SIMORD is not limited to mention detection; it approximates downstream clinical documentation structure as it would be entered into an electronic health record.

## 2. Corpus composition and data sources

SIMORD is constructed from high-quality, publicly available doctor-patient consultation datasets. Its source corpora are **ACI-Bench**, described as containing 207 real-world physician-patient conversations curated for clinical realism and diversity, and **PriMock57**, described as containing 57 mock primary care consultations with audio, manual transcriptions, and consultation notes [2507.05517].

The SIMORD splits are sampled from these two sources with an explicit focus on the highest-quality and most plausible dialog episodes. The released structure comprises a training set of 64 samples, a development set of 100 samples, and a test set of 100 conversations. The test set contains 255 annotated medical orders. The development set is described as being used for prompt tuning and validation, while the training split is used for few-shot prompting.

A related synthetic dataset, **Notechat**, was evaluated during curation but was not included in SIMORD because of lower observed dialog quality. That exclusion is methodologically important: the dataset is not simply an aggregation of available conversational sources, but a filtered benchmark assembled around plausibility and annotation suitability.

The long-form nature of the conversations is central to the benchmark. Unlike tasks defined over short clinical notes or isolated medication strings, SIMORD requires models to recover structured orders from multi-turn dialogue with distributed evidence, deferred decisions, and repeated mentions. A plausible implication is that the benchmark tests both clinical semantic interpretation and long-context discourse resolution.

## 3. Annotation schema and quality control

The annotations are produced by medically trained annotators. Their instructions are to identify every medical order of type medication, imaging, lab, or follow-up within the conversation and to record the associated description, type, reason, and provenance [2507.05517].

The schema can be summarized as follows:

| Field | Definition | Notes |
|---|---|---|
| Description | Concise summary of the order | Required |
| Reason | Motivation or diagnosis for the order | Optional |
| Type | `"medication"`, `"laboratory"`, `"followup"`, or `"imaging"` | Categorical |
| Provenance | List of conversation line numbers where the order is mentioned | Supports traceability |

The inclusion of provenance is a notable feature. In this dataset, provenance is not merely a convenience for error analysis; it is an explicitly scored target field. That design choice makes traceability part of the task definition itself.

Inter-annotator agreement is reported as **Cohen’s kappa = 0.768**, which the source characterizes as substantial agreement for this complex task. For a benchmark involving long dialogues, optional fields, and clinically grounded order abstraction, this figure indicates that the annotation protocol reaches a relatively stable consensus without collapsing the task into a simpler entity-tagging formulation.

The dataset’s JSON-centered representation is also consequential. Because both annotation and prediction must conform to standardized structured objects, evaluation depends not only on semantic correctness but also on output well-formedness. The study reports that some models had frequent parsing errors, making format adherence part of the empirical difficulty.

## 4. Evaluation methodology

SIMORD evaluates systems with five metrics: **Match**, **Description**, **Reason**, **Type**, and **Provenance** [2507.05517].

**Match** is an F1 score between reference and prediction based on description word overlap. It is described as order-level alignment without content details and as an upper bound on the possible scores for the remaining fields. **Description** is the F1 over bag-of-words between gold and predicted descriptions. **Reason** is the F1 over bag-of-words for the reason field. **Type** is evaluated with accuracy because the label space is small and categorical. **Provenance** is the F1 over line numbers identifying the source segments for the order.

For Match, Description, Reason, and Provenance, the paper uses the standard F1 definition
$$
F_1 = \frac{2\,TP}{2\,TP + FP + FN}.
$$

The metric suite reflects the dataset’s multi-attribute design. Match evaluates whether the model recovered the existence of the order; Description and Reason evaluate free-text content; Type evaluates categorical classification; and Provenance evaluates traceability to the dialogue. This decomposition is technically important because a model can succeed at order detection while failing at reason attribution or evidence localization.

The benchmark also distinguishes prompting configurations. Closed-weight models are evaluated in zero-shot and one-shot settings, while open-weight models are evaluated in zero-shot and two-shot settings. The study further examines “reasoning” or chain-of-thought prompt variants. In this setup, in-context examples are not incidental prompt engineering details; they are treated as a systematic component of the benchmark.

## 5. Benchmark results and model behavior

The paper evaluates both closed-weight and open-weight language models on SIMORD. The closed-weight systems include GPT-4o, GPT-4.1, o1-mini, o1-prev, and o3-mini. The open-weight systems include Phi3.5-mini-instruct (3.8B), Mediphi-Instruct (3.8B, medical fine-tuned), Llama3-8B-instruct, and Llama3-Med42-8B (medical-finetuned) [2507.05517].

Among closed-weight models, the best **Match F1** is reported as approximately **68%** for GPT-4o and o3-mini with one-shot prompting. **Description F1** reaches **38.5%** for GPT-4o in zero-shot and **42.8%** for GPT-4o with an example. The highest **Reason F1** among the closed models is **26.6%** for o1-mini with an example. **Type Accuracy** reaches **66.8%** for o3-mini with an example. The highest **Provenance F1** is **43.2%** for o1-mini in zero-shot, and the study reports large provenance gains for reasoning models.

The study attributes specific gains to adding one example in the prompt: **Description** improves by **+6.6%**, **Reason** by **+3.8%**, and **Type** by **+2.5%**. For Provenance, reasoning models show a further improvement of **+25.7%**. The reported pattern indicates that in-context exemplars and reasoning-oriented prompting are especially useful when the task requires evidence attribution rather than only coarse order detection.

For open-weight models, zero-shot performance is described as somewhat lower on Match and Description than that of closed models. However, with two-shot prompting, **Mediphi-Instruct** reaches **51.9% Description**, surpassing GPT-4o’s one-shot performance. The best **Reason F1** among the open models is **37.9%** for **Med42** in the two-shot setting. The paper also reports that parsing errors decrease as in-context examples are included.

A key result is stated explicitly: “On SIMORD, we show that the 3.8B-parameter MediPhi-Instruct attains parity with GPT-4o (two-shot vs. one-shot) and surpasses it on the description and reason metrics, demonstrating the viability of lightweight open-weight models for this task” [2507.05517]. Within the benchmark’s scope, this establishes that prompt-conditioned small medical models can approach or exceed proprietary systems on selected attributes.

The study also identifies recurring failure modes. Models frequently fabricate or omit orders and often aggregate sequential orders, especially in laboratory orders. Because the output must be structured JSON, formatting violations are an additional error class. These observations indicate that the task difficulty lies not only in medical semantics but also in boundary determination, decomposition of compound orders, and faithful serialization.

## 6. Relation to prior datasets and clinical NLP tasks

SIMORD is positioned as the first publicly available dataset for medical order extraction from conversations and, more specifically, as the first open dataset for multi-attribute, realistic medical order extraction from dialogue [2507.05517]. The paper contrasts it with prior resources such as **MedEx** and **n2c2 ADE**, which are associated with medication or adverse drug event extraction from clinical notes rather than spoken consultations.

The distinction is not only one of modality. Prior datasets cited in the comparison do not combine full dialogue context with the same attribute set of free-text description, optional reason, categorical order type, and provenance. SIMORD therefore extends beyond medication mention extraction or relation extraction into structured clinical action representation grounded in conversational evidence.

The benchmark’s emphasis on long, multi-turn dialogue also differentiates it from note-centric tasks. In notes, relevant content is already partially normalized by the authoring clinician. In doctor-patient conversations, the intended order can emerge indirectly, through negotiation, clarification, or temporal progression across turns. This suggests that SIMORD occupies an intermediate space between spoken clinical NLP and EHR action generation.

The exclusion of Notechat because of lower observed dialog quality further clarifies the benchmark’s intended scope. The dataset is not merely synthetic dialogue repurposed for extraction; it is a curated resource meant to preserve conversational plausibility while remaining publicly usable.

## 7. Availability, significance, and research directions

The paper states that SIMORD will be made publicly available upon acceptance to the target venue and describes it as an open-source dataset intended to support community benchmarking and further research on EHR order extraction from real and simulated clinical conversations [2507.05517]. Its companion role in the broader study is to provide a public benchmark for one of two underexplored clinical documentation tasks, the other being nurse observation extraction.

SIMORD’s significance lies in the combination of public accessibility, realistic conversational source material, order-level structured annotation, and explicit provenance supervision. These elements make it suitable for evaluating systems that must transform clinician-patient dialogue into EHR-like artifacts rather than merely detect isolated concepts.

At the same time, the reported benchmark results show that the task remains difficult. Match scores are materially higher than Description, Reason, and Provenance scores, indicating that identifying that an order exists is easier than producing detailed structured content and evidence alignment. This suggests that conversational order extraction should not be conflated with generic information extraction from text. It is a structured generation problem with clinical grounding, traceability requirements, and nontrivial failure modes.

A plausible implication is that SIMORD can function as a benchmark for several adjacent research directions: long-context clinical reasoning, structured JSON-constrained decoding, provenance-aware extraction, and the comparison of medical fine-tuning against proprietary general-purpose models. Within the bounds of the reported study, the dataset provides a public test bed for these questions while addressing a previously missing component of clinical NLP evaluation.

Source: https://www.emergentmind.com/topics/simord