Papers
Topics
Authors
Recent
Search
2000 character limit reached

Post-Editing Effort in Time (PEET)

Updated 14 July 2026
  • Post-Editing Effort in Time (PEET) is defined as the human time required to transform machine-generated text into an acceptable final version, often normalized per word or segment.
  • PEET is operationalized through diverse metrics—from seconds per word measured by CAT tools to throughput-derived inverses—reflecting variations in timing protocols and active edit filtering.
  • Empirical studies show that while post-editing generally speeds up text production compared to translation from scratch, its efficiency depends on language pairs, quality estimation methods, and individual editor behavior.

Searching arXiv for recent and relevant papers on PEET and post-editing time. First, I’ll look for papers explicitly using the term PEET or closely related post-editing-time formulations. Post-Editing Effort in Time (PEET) denotes the temporal effort required to transform a system output into a final acceptable text. In grammar error correction, it is introduced explicitly as a human-centered evaluation measure based on time-to-correct (Vadehra et al., 5 Oct 2025). In earlier machine translation and post-editing research, the same construct was operationalized under other names, most notably post-editing time per word, segment-level elapsed time, and throughput-derived inverses such as words per hour and hours per 1,000 words (Scarton et al., 2019). Within Krings’ three-way distinction of temporal, technical, and cognitive effort, PEET corresponds to the temporal dimension, while remaining closely related to keystrokes, edit operations, pauses, quality judgments, and interface effects (Ahsan et al., 2021).

1. Conceptual foundations

PEET is most directly defined as the time spent by a human editor in transforming an initial text into a satisfactory final version. In the post-editing literature, this temporal dimension is repeatedly treated as the dimension most immediately linked to productivity or throughput. One canonical operationalization is post-editing time per word, denoted PETpW, where for a segment ss, with elapsed time T(s)T(s) and machine-translated length L(s)L(s), the metric is

PETpW(s)=T(s)L(s).\text{PETpW}(s) = \frac{T(s)}{L(s)}.

This ratio is described as directly usable to assess the effort of post-editing a segment (Scarton et al., 2019).

The same idea appears in several neighboring formulations. In English–Chinese MT post-editing with sentence-level quality estimation, temporal effort is operationalized strictly as editing time per segment normalized by source-text length,

PE_time_pwij=TijNij,\text{PE\_time\_pw}_{ij} = \frac{T_{ij}}{N_{ij}},

where TijT_{ij} is total post-editing time in seconds and NijN_{ij} is the number of words in the source text of segment jj for participant ii (Liu et al., 22 Jul 2025). In document-level post-editing logs, time is normalized by source length and modeled as log⁡(Tdoc/Nsrc)\log(T_{\text{doc}}/N_{\text{src}}), explicitly following prior work that uses log time per word to stabilize variance (Góis et al., 2019). In the English–Hindi direction, temporal effort is measured both as segment-level total time T(s)T(s)0 and as throughput in words per hour, with time per word recoverable as the inverse of throughput (Ahsan et al., 2021).

This family resemblance matters because the acronym PEET is recent, whereas the underlying construct is older. A plausible interpretation is that PEET functions as an umbrella label for a previously fragmented set of time-based effort variables: seconds per word, seconds per segment, time-to-correct, words per hour, and hours per 1,000 words. The unifying feature is that the target quantity is not perceived quality alone, nor edit distance alone, but the elapsed human time required to obtain the desired final text (Vadehra et al., 5 Oct 2025).

2. Operationalization and measurement

Across studies, PEET is measured with markedly different instruments and normalizations. The common structure is elapsed editing time plus a length normalization, but the logging substrate varies from CAT-tool timers to screen recordings and survey timing widgets.

Context Time variable Instrumentation
MT effort estimation T(s)T(s)1 PET tool (Scarton et al., 2019)
English–Chinese MTPE with QE T(s)T(s)2 YiCAT internal timing (Liu et al., 22 Jul 2025)
Professional En→Cs NMT PE Estimated sec/word with capped think time T(s)T(s)3 Memsource logs (Zouhar et al., 2021)
Industrial HT vs PE WPH and H/KW Customized memoQ-based CAT environment (Terribile, 2023)
Consultation-note post-editing Writing time vs post-editing time per note Heartex + manual screen-recording review (Moramarco et al., 2021)
GEC tool evaluation Sentence-level time-to-correct Qualtrics timing question (Vadehra et al., 5 Oct 2025)

Several studies make the measurement problem itself part of the analysis. In the consultation-note study, writing from scratch and post-editing are timed from screen recordings rather than platform timers, and the authors note that this is feasible but labor-intensive (Moramarco et al., 2021). In YiCAT, a default behavior fails to record time if no edits are made, so participants are instructed to type “1” at the end of accepted-as-is segments, ensuring non-zero timing for minimal verification effort (Liu et al., 22 Jul 2025). In Memsource logs, raw think time contains substantial disturbance from breaks and distractions, so the estimated think component is capped at 10 seconds per word before constructing estimated total time per word (Zouhar et al., 2021). In industrial CAT data, only active time is counted: inactivity longer than two minutes is deleted, which makes WPH values relatively high compared with studies including all working time, but still comparable across human translation and post-editing (Terribile, 2023).

These designs imply that PEET is not a single immutable metric but a measurement family with different observational assumptions. Some variants include only active CAT time, some include all within-segment elapsed time, and some approximate true editing time through filtering or capping. This suggests that comparisons of PEET across papers are valid only when the timing protocol, pause handling, and normalization unit are specified explicitly.

3. Empirical regularities across applications

The dominant empirical pattern is that post-editing is often faster than generating the target text from scratch, but the magnitude of the gain is highly contingent. In large-scale localization data from 90 million words across 11 language pairs, average post-editing speed at the translation stage is 66% higher than human translation, and average PE revision speed is 38% higher than HT revision; yet the same study also reports that PE is usually but not always faster than HT, with English→Swedish slightly slower on average and English→Polish reversing once outliers are excluded (Terribile, 2023). In DivEMT, post-editing is consistently faster than translation from scratch across Arabic, Dutch, Italian, Turkish, Ukrainian, and Vietnamese, but the magnitude of productivity gains varies widely by language and by whether the MT output comes from Google Translate or mBART-50 (Sarti et al., 2022).

Controlled experiments report the same asymmetry at smaller scale. In the English–Hindi direction, post-editing reduces translation time by 63%, with throughput rising from 359 words/hour in human translation to 979 words/hour in post-editing, while human evaluation detects no discernible quality differences between outputs produced under the two conditions (Ahsan et al., 2021). In English–Chinese MTPE with sentence-level QE, mean normalized post-editing time drops from 1.27 s/word without QE to 0.95 s/word with QE, with a significant main effect of task type (Liu et al., 22 Jul 2025). In consultation-note generation, post-editing is faster than writing from scratch in almost all cases, although the authors also report exceptions, strong inter-physician variability, and a clear familiarity effect, with the first task averaging 36 minutes and subsequent tasks 23 minutes (Moramarco et al., 2021).

Outside MT, the GEC PEET study shows the same general pattern. Editing original source sentences averages 31.16 seconds per sentence, whereas post-editing GECToR outputs averages 26.82 seconds and post-editing GEC-PD outputs 27.46 seconds, implying roughly four seconds saved per sentence by starting from a tool output (Vadehra et al., 5 Oct 2025). A plausible implication is that temporal post-editing effort is not specific to translation workflows: it behaves similarly in MT, summarization-assisted note writing, and grammar correction whenever a human refines a machine-produced first draft.

4. Determinants and predictors of PEET

A central issue in PEET research is which observable variables track time reliably. The answer is mixed. In the English–Chinese QE study, higher human-rated MT quality significantly reduces time per word, and higher translator expertise also reduces time per word, but neither factor shows a significant interaction with the time-saving effect of sentence-level QE. Within the medium/high MT quality range studied, QE produces a roughly parallel downward shift in temporal effort across quality levels and across expertise groups (Liu et al., 22 Jul 2025).

By contrast, high-level automatic MT metrics are unstable predictors of time in strong NMT regimes. In English→Czech professional post-editing, better NMT systems clearly lead to fewer changes, but BLEU is “definitely not a stable predictor of the time or final output quality.” Linear fits of system-level BLEU against seconds per word change sign depending on which subset of systems is included, and review time can even increase with better BLEU (Zouhar et al., 2021). The industrial LSP study reaches a parallel conclusion for edit distance: edit distance does not correlate strongly with speed and cannot be used as a reliable proxy for post-editing productivity in time, especially once TM matches and repetitions are included (Terribile, 2023).

Task-based metrics that compare machine outputs with actual post-edited outputs perform better. Using PETpW as gold standard, task-based PE metrics such as HTER, HBLEU, HMETEOR, and especially Keys/char track post-editing time more closely than Direct Assessment, while reference-based BLEU, TER, and METEOR perform worst (Scarton et al., 2019). In English–Hindi post-editing, H-BLEU and H-chrF correlate negatively with temporal effort, H-TER correlates positively, and the correlations with time are moderate, while average and initial pause duration show no meaningful correlation with these automatic scores (Ahsan et al., 2021).

Two later lines extend PEET prediction beyond classical MT metrics. “Translator2Vec” learns document-session and editor embeddings from action sequences that include waiting times, jumps, and mouse actions; adding dynamic editor embeddings raises Pearson correlation for time prediction from 17.58 to 47.69 on En–Fr test data and from 23.67 to 38.72 on En–De test data (Góis et al., 2019). “Assessing Human Editing Effort on LLM-Generated Texts via Compression-Based Edit Distance” introduces a directional LZ77-based compression distance T(s)T(s)4, which correlates with human edit time up to 0.81 on accounting QA edits and predicts time better than traditional edit-distance and overlap metrics in that setting (Devatine et al., 2024). Taken together, these results suggest that PEET is better predicted by task-grounded process signals and post-edit transformations than by generic reference similarity alone.

5. Human factors, interface effects, and workflow structure

PEET is strongly shaped by editor behavior. The consultation-note study reports large inter-physician variability: one physician writes shorter, terser notes and edits only substantial issues, while another edits generated notes extensively and therefore shows higher editing times and higher counts of omissions and incorrectness in scoring (Moramarco et al., 2021). The industrial HT/PE study finds the same pattern at scale, with large dispersion both by job and by linguist, and warns that a single “average PE speed” is a poor predictor of any given linguist’s performance (Terribile, 2023). DivEMT adds a cross-lingual dimension: productivity gains depend not only on system quality but also on typological relatedness to English and target morphological complexity, with Italian and Dutch benefiting much more from mBART-50 than Turkish or Arabic (Sarti et al., 2022).

Interface interventions can change PEET, but not uniformly. In QE4PE, word-level QE highlights are studied under NoHighlight, Oracle, supervised XCOMET-XXL, and uncertainty-based conditions. Domain, language, and editors’ speed are critical factors in determining highlights’ effectiveness, and the downstream differences between human-made and automated highlights are modest (Sarti et al., 4 Mar 2025). The same paper reports a marked gap between objective and subjective utility: highlights improved correction of critical errors, but many editors described them as eye distractions and not accurate enough to rely on (Sarti et al., 4 Mar 2025). A plausible implication is that PEET should not be reduced to elapsed time alone when interface design changes the distribution of cognitive attention.

Related work in interactive MT shows a neighboring measurement regime based on operation counts rather than wall-clock time. “Online Learning for Effort Reduction in Interactive Neural Machine Translation” measures TER and KSMR, not human time, and explicitly frames them as effort surrogates. Because response times are around 0.1 to 0.3 seconds and learning times around 0.1 seconds, the study argues that system latency is close to real-time usability, making human interaction costs dominant (Peris et al., 2018). This suggests that operation-based effort metrics remain relevant where direct timing is unavailable, but also that their relationship to PEET depends on interface latency being negligible.

6. Evaluation status, limitations, and recurrent misconceptions

PEET occupies an unusual position among evaluation constructs: it is expensive, highly informative, and difficult to standardize. The GEC PEET study presents time-to-correct as a human-focused evaluation scorer for ranking tools by estimated post-editing time (Vadehra et al., 5 Oct 2025), while the MT effort-estimation study argues that task-based measurements are the most reliable when estimating MT quality for a specific task, precisely because they are obtained during or after performing that task (Scarton et al., 2019). Across both lines, a recurring conclusion is that temporal effort is more directly tied to usability than reference overlap.

Several misconceptions are repeatedly challenged. One is that better automatic MT scores should translate monotonically into lower PEET. High-quality NMT studies show that this fails once systems cluster near the top: BLEU improvements do not map stably to time savings (Zouhar et al., 2021). Another is that edit distance is an adequate proxy for time. Large industrial data show weak-to-moderate correlations only, insufficient for robust productivity estimation (Terribile, 2023). A third is that faster editing necessarily means lower overall human burden. In clinical note post-editing, one physician reports that post-editing may be faster while imposing higher cognitive load because every part of the generated text must be critically evaluated (Moramarco et al., 2021).

The limitations are correspondingly broad. Timing protocols differ across tools, including active-time logging, manual screen review, and survey timers; some systems fail to record time for untouched segments unless forced edits are introduced (Liu et al., 22 Jul 2025). Many datasets remain narrow in domain, language, or participant pool, such as English–Chinese student translators, English–Hindi news, or English-only learner-text correction (Liu et al., 22 Jul 2025). Sample-size constraints, limited low-quality MT coverage, and interface-specific behaviors further complicate generalization (Vadehra et al., 5 Oct 2025). For this reason, PEET is best understood not as a single universal number but as a family of task-grounded temporal measures whose interpretation depends on editor population, text type, system quality, and logging design.

A broad synthesis follows from the literature. PEET is most informative when it is measured directly, normalized transparently, and analyzed jointly with technical and quality variables. It is least reliable when inferred from generic reference-based metrics alone. The recent explicit naming of PEET formalizes a construct that earlier post-editing research had already treated as central: the time humans actually spend making machine output usable (Vadehra et al., 5 Oct 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Post-Editing Effort in Time (PEET).