SIMORD: Open Dataset for Medical Order Extraction
- SIMORD is a dataset for extracting structured medical orders from simulated doctor-patient dialogues, covering orders for medications, laboratory tests, imaging, and follow-up actions.
- The dataset features detailed annotations with attributes such as description, reason, categorical type, and provenance to support traceability and structured output.
- Evaluations reveal that in-context examples and reasoning prompts improve performance, with models showing up to 68% Match F1 and significant gains in detailed extraction metrics.
Searching arXiv for the specified paper and related clinical NLP datasets to ground the article. SIMORD (SIMulated ORDer) is the first open-source dataset for extracting medical orders from doctor-patient simulated transcripts. It is introduced to support the medical order extraction task, which consists of identifying and structuring orders discussed during consultations, including medications, laboratory tests, imaging, and follow-up actions. In the source study, SIMORD is presented as a resource for benchmarking LLMs on a clinically consequential documentation problem that remains underexplored because of data scarcity and sensitivity; the stated motivation is to automate documentation, improve clinical workflows, and reduce the burden on healthcare providers (Corbeil et al., 7 Jul 2025).
1. Definition and task formulation
SIMORD is designed around long-form doctor-patient conversation transcripts. Each instance is annotated at the order level rather than only at the span or utterance level. For every order, the dataset records four attributes: a Description, defined as a concise summary of the order; an optional Reason, defined as the motivation or diagnosis for the order; a Type, drawn from the categorical label set "medication", "laboratory", "followup", and "imaging"; and Provenance, defined as the list of conversation line numbers where the order is mentioned (Corbeil et al., 7 Jul 2025).
The required output format for both human annotation and language-model prediction is a list of standardized JSON objects, one object per order. This design makes the task explicitly structured rather than purely extractive. It also embeds traceability through line-number provenance, which is intended to support linkage between the structured output and the conversational evidence.
The framing of the task is operational rather than merely semantic. Annotators are instructed to assess every medical order within the conversation “the way a doctor would create them in the EHR,” and the dataset aims to replicate the doctor’s usual end-of-encounter order-entry process. This suggests that SIMORD is not limited to mention detection; it approximates downstream clinical documentation structure as it would be entered into an electronic health record.
2. Corpus composition and data sources
SIMORD is constructed from high-quality, publicly available doctor-patient consultation datasets. Its source corpora are ACI-Bench, described as containing 207 real-world physician-patient conversations curated for clinical realism and diversity, and PriMock57, described as containing 57 mock primary care consultations with audio, manual transcriptions, and consultation notes (Corbeil et al., 7 Jul 2025).
The SIMORD splits are sampled from these two sources with an explicit focus on the highest-quality and most plausible dialog episodes. The released structure comprises a training set of 64 samples, a development set of 100 samples, and a test set of 100 conversations. The test set contains 255 annotated medical orders. The development set is described as being used for prompt tuning and validation, while the training split is used for few-shot prompting.
A related synthetic dataset, Notechat, was evaluated during curation but was not included in SIMORD because of lower observed dialog quality. That exclusion is methodologically important: the dataset is not simply an aggregation of available conversational sources, but a filtered benchmark assembled around plausibility and annotation suitability.
The long-form nature of the conversations is central to the benchmark. Unlike tasks defined over short clinical notes or isolated medication strings, SIMORD requires models to recover structured orders from multi-turn dialogue with distributed evidence, deferred decisions, and repeated mentions. A plausible implication is that the benchmark tests both clinical semantic interpretation and long-context discourse resolution.
3. Annotation schema and quality control
The annotations are produced by medically trained annotators. Their instructions are to identify every medical order of type medication, imaging, lab, or follow-up within the conversation and to record the associated description, type, reason, and provenance (Corbeil et al., 7 Jul 2025).
The schema can be summarized as follows:
| Field | Definition | Notes |
|---|---|---|
| Description | Concise summary of the order | Required |
| Reason | Motivation or diagnosis for the order | Optional |
| Type | "medication", "laboratory", "followup", or "imaging" |
Categorical |
| Provenance | List of conversation line numbers where the order is mentioned | Supports traceability |
The inclusion of provenance is a notable feature. In this dataset, provenance is not merely a convenience for error analysis; it is an explicitly scored target field. That design choice makes traceability part of the task definition itself.
Inter-annotator agreement is reported as Cohen’s kappa = 0.768, which the source characterizes as substantial agreement for this complex task. For a benchmark involving long dialogues, optional fields, and clinically grounded order abstraction, this figure indicates that the annotation protocol reaches a relatively stable consensus without collapsing the task into a simpler entity-tagging formulation.
The dataset’s JSON-centered representation is also consequential. Because both annotation and prediction must conform to standardized structured objects, evaluation depends not only on semantic correctness but also on output well-formedness. The study reports that some models had frequent parsing errors, making format adherence part of the empirical difficulty.
4. Evaluation methodology
SIMORD evaluates systems with five metrics: Match, Description, Reason, Type, and Provenance (Corbeil et al., 7 Jul 2025).
Match is an F1 score between reference and prediction based on description word overlap. It is described as order-level alignment without content details and as an upper bound on the possible scores for the remaining fields. Description is the F1 over bag-of-words between gold and predicted descriptions. Reason is the F1 over bag-of-words for the reason field. Type is evaluated with accuracy because the label space is small and categorical. Provenance is the F1 over line numbers identifying the source segments for the order.
For Match, Description, Reason, and Provenance, the paper uses the standard F1 definition
The metric suite reflects the dataset’s multi-attribute design. Match evaluates whether the model recovered the existence of the order; Description and Reason evaluate free-text content; Type evaluates categorical classification; and Provenance evaluates traceability to the dialogue. This decomposition is technically important because a model can succeed at order detection while failing at reason attribution or evidence localization.
The benchmark also distinguishes prompting configurations. Closed-weight models are evaluated in zero-shot and one-shot settings, while open-weight models are evaluated in zero-shot and two-shot settings. The study further examines “reasoning” or chain-of-thought prompt variants. In this setup, in-context examples are not incidental prompt engineering details; they are treated as a systematic component of the benchmark.
5. Benchmark results and model behavior
The paper evaluates both closed-weight and open-weight LLMs on SIMORD. The closed-weight systems include GPT-4o, GPT-4.1, o1-mini, o1-prev, and o3-mini. The open-weight systems include Phi3.5-mini-instruct (3.8B), Mediphi-Instruct (3.8B, medical fine-tuned), Llama3-8B-instruct, and Llama3-Med42-8B (medical-finetuned) (Corbeil et al., 7 Jul 2025).
Among closed-weight models, the best Match F1 is reported as approximately 68% for GPT-4o and o3-mini with one-shot prompting. Description F1 reaches 38.5% for GPT-4o in zero-shot and 42.8% for GPT-4o with an example. The highest Reason F1 among the closed models is 26.6% for o1-mini with an example. Type Accuracy reaches 66.8% for o3-mini with an example. The highest Provenance F1 is 43.2% for o1-mini in zero-shot, and the study reports large provenance gains for reasoning models.
The study attributes specific gains to adding one example in the prompt: Description improves by +6.6%, Reason by +3.8%, and Type by +2.5%. For Provenance, reasoning models show a further improvement of +25.7%. The reported pattern indicates that in-context exemplars and reasoning-oriented prompting are especially useful when the task requires evidence attribution rather than only coarse order detection.
For open-weight models, zero-shot performance is described as somewhat lower on Match and Description than that of closed models. However, with two-shot prompting, Mediphi-Instruct reaches 51.9% Description, surpassing GPT-4o’s one-shot performance. The best Reason F1 among the open models is 37.9% for Med42 in the two-shot setting. The paper also reports that parsing errors decrease as in-context examples are included.
A key result is stated explicitly: “On SIMORD, we show that the 3.8B-parameter MediPhi-Instruct attains parity with GPT-4o (two-shot vs. one-shot) and surpasses it on the description and reason metrics, demonstrating the viability of lightweight open-weight models for this task” (Corbeil et al., 7 Jul 2025). Within the benchmark’s scope, this establishes that prompt-conditioned small medical models can approach or exceed proprietary systems on selected attributes.
The study also identifies recurring failure modes. Models frequently fabricate or omit orders and often aggregate sequential orders, especially in laboratory orders. Because the output must be structured JSON, formatting violations are an additional error class. These observations indicate that the task difficulty lies not only in medical semantics but also in boundary determination, decomposition of compound orders, and faithful serialization.
6. Relation to prior datasets and clinical NLP tasks
SIMORD is positioned as the first publicly available dataset for medical order extraction from conversations and, more specifically, as the first open dataset for multi-attribute, realistic medical order extraction from dialogue (Corbeil et al., 7 Jul 2025). The paper contrasts it with prior resources such as MedEx and n2c2 ADE, which are associated with medication or adverse drug event extraction from clinical notes rather than spoken consultations.
The distinction is not only one of modality. Prior datasets cited in the comparison do not combine full dialogue context with the same attribute set of free-text description, optional reason, categorical order type, and provenance. SIMORD therefore extends beyond medication mention extraction or relation extraction into structured clinical action representation grounded in conversational evidence.
The benchmark’s emphasis on long, multi-turn dialogue also differentiates it from note-centric tasks. In notes, relevant content is already partially normalized by the authoring clinician. In doctor-patient conversations, the intended order can emerge indirectly, through negotiation, clarification, or temporal progression across turns. This suggests that SIMORD occupies an intermediate space between spoken clinical NLP and EHR action generation.
The exclusion of Notechat because of lower observed dialog quality further clarifies the benchmark’s intended scope. The dataset is not merely synthetic dialogue repurposed for extraction; it is a curated resource meant to preserve conversational plausibility while remaining publicly usable.
7. Availability, significance, and research directions
The paper states that SIMORD will be made publicly available upon acceptance to the target venue and describes it as an open-source dataset intended to support community benchmarking and further research on EHR order extraction from real and simulated clinical conversations (Corbeil et al., 7 Jul 2025). Its companion role in the broader study is to provide a public benchmark for one of two underexplored clinical documentation tasks, the other being nurse observation extraction.
SIMORD’s significance lies in the combination of public accessibility, realistic conversational source material, order-level structured annotation, and explicit provenance supervision. These elements make it suitable for evaluating systems that must transform clinician-patient dialogue into EHR-like artifacts rather than merely detect isolated concepts.
At the same time, the reported benchmark results show that the task remains difficult. Match scores are materially higher than Description, Reason, and Provenance scores, indicating that identifying that an order exists is easier than producing detailed structured content and evidence alignment. This suggests that conversational order extraction should not be conflated with generic information extraction from text. It is a structured generation problem with clinical grounding, traceability requirements, and nontrivial failure modes.
A plausible implication is that SIMORD can function as a benchmark for several adjacent research directions: long-context clinical reasoning, structured JSON-constrained decoding, provenance-aware extraction, and the comparison of medical fine-tuning against proprietary general-purpose models. Within the bounds of the reported study, the dataset provides a public test bed for these questions while addressing a previously missing component of clinical NLP evaluation.