---
title: Post-Editing Effort in Time (PEET)
url: https://www.emergentmind.com/topics/post-editing-effort-in-time-peet
type: topic
---

# Post-Editing Effort in Time (PEET)

Searching arXiv for recent and relevant papers on PEET and post-editing time.
First, I’ll look for papers explicitly using the term PEET or closely related post-editing-time formulations.
Post-Editing Effort in Time (PEET) denotes the temporal effort required to transform a system output into a final acceptable text. In grammar error correction, it is introduced explicitly as a human-centered evaluation measure based on time-to-correct [2510.04394]. In earlier machine translation and post-editing research, the same construct was operationalized under other names, most notably post-editing time per word, segment-level elapsed time, and throughput-derived inverses such as words per hour and hours per 1,000 words [1910.06204]. Within Krings’ three-way distinction of temporal, technical, and cognitive effort, PEET corresponds to the temporal dimension, while remaining closely related to keystrokes, edit operations, pauses, quality judgments, and interface effects [2112.09841].

## 1. Conceptual foundations

PEET is most directly defined as the time spent by a human editor in transforming an initial text into a satisfactory final version. In the post-editing literature, this temporal dimension is repeatedly treated as the dimension most immediately linked to productivity or throughput. One canonical operationalization is post-editing time per word, denoted PETpW, where for a segment \(s\), with elapsed time \(T(s)\) and machine-translated length \(L(s)\), the metric is
\[
\text{PETpW}(s) = \frac{T(s)}{L(s)}.
\]
This ratio is described as directly usable to assess the effort of post-editing a segment [1910.06204].

The same idea appears in several neighboring formulations. In English–Chinese MT post-editing with sentence-level quality estimation, temporal effort is operationalized strictly as editing time per segment normalized by source-text length,
\[
\text{PE\_time\_pw}_{ij} = \frac{T_{ij}}{N_{ij}},
\]
where \(T_{ij}\) is total post-editing time in seconds and \(N_{ij}\) is the number of words in the source text of segment \(j\) for participant \(i\) [2507.16515]. In document-level post-editing logs, time is normalized by source length and modeled as \(\log(T_{\text{doc}}/N_{\text{src}})\), explicitly following prior work that uses log time per word to stabilize variance [1907.10362]. In the English–Hindi direction, temporal effort is measured both as segment-level total time \(T_{ij}\) and as throughput in words per hour, with time per word recoverable as the inverse of throughput [2112.09841].

This family resemblance matters because the acronym PEET is recent, whereas the underlying construct is older. A plausible interpretation is that PEET functions as an umbrella label for a previously fragmented set of time-based effort variables: seconds per word, seconds per segment, time-to-correct, words per hour, and hours per 1,000 words. The unifying feature is that the target quantity is not perceived quality alone, nor edit distance alone, but the elapsed human time required to obtain the desired final text [2510.04394].

## 2. Operationalization and measurement

Across studies, PEET is measured with markedly different instruments and normalizations. The common structure is elapsed editing time plus a length normalization, but the logging substrate varies from CAT-tool timers to screen recordings and survey timing widgets.

| Context | Time variable | Instrumentation |
|---|---|---|
| MT effort estimation | \(\text{PETpW}(s)=T(s)/L(s)\) | PET tool [1910.06204] |
| English–Chinese MTPE with QE | \(\text{PE\_time\_pw}_{ij}=T_{ij}/N_{ij}\) | YiCAT internal timing [2507.16515] |
| Professional En→Cs NMT PE | Estimated sec/word with capped think time \(\overset{*}{A}\) | Memsource logs [2109.05016] |
| Industrial HT vs PE | WPH and H/KW | Customized memoQ-based CAT environment [2312.12660] |
| Consultation-note post-editing | Writing time vs post-editing time per note | Heartex + manual screen-recording review [2104.04402] |
| GEC tool evaluation | Sentence-level time-to-correct | Qualtrics timing question [2510.04394] |

Several studies make the measurement problem itself part of the analysis. In the consultation-note study, writing from scratch and post-editing are timed from screen recordings rather than platform timers, and the authors note that this is feasible but labor-intensive [2104.04402]. In YiCAT, a default behavior fails to record time if no edits are made, so participants are instructed to type “1” at the end of accepted-as-is segments, ensuring non-zero timing for minimal verification effort [2507.16515]. In Memsource logs, raw think time contains substantial disturbance from breaks and distractions, so the estimated think component is capped at 10 seconds per word before constructing estimated total time per word [2109.05016]. In industrial CAT data, only active time is counted: inactivity longer than two minutes is deleted, which makes WPH values relatively high compared with studies including all working time, but still comparable across human translation and post-editing [2312.12660].

These designs imply that PEET is not a single immutable metric but a measurement family with different observational assumptions. Some variants include only active CAT time, some include all within-segment elapsed time, and some approximate true editing time through filtering or capping. This suggests that comparisons of PEET across papers are valid only when the timing protocol, pause handling, and normalization unit are specified explicitly.

## 3. Empirical regularities across applications

The dominant empirical pattern is that post-editing is often faster than generating the target text from scratch, but the magnitude of the gain is highly contingent. In large-scale localization data from 90 million words across 11 language pairs, average post-editing speed at the translation stage is 66% higher than human translation, and average PE revision speed is 38% higher than HT revision; yet the same study also reports that PE is usually but not always faster than HT, with English→Swedish slightly slower on average and English→Polish reversing once outliers are excluded [2312.12660]. In DivEMT, post-editing is consistently faster than translation from scratch across Arabic, Dutch, Italian, Turkish, Ukrainian, and Vietnamese, but the magnitude of productivity gains varies widely by language and by whether the MT output comes from Google Translate or mBART-50 [2205.12215].

Controlled experiments report the same asymmetry at smaller scale. In the English–Hindi direction, post-editing reduces translation time by 63%, with throughput rising from 359 words/hour in human translation to 979 words/hour in post-editing, while human evaluation detects no discernible quality differences between outputs produced under the two conditions [2112.09841]. In English–Chinese MTPE with sentence-level QE, mean normalized post-editing time drops from 1.27 s/word without QE to 0.95 s/word with QE, with a significant main effect of task type [2507.16515]. In consultation-note generation, post-editing is faster than writing from scratch in almost all cases, although the authors also report exceptions, strong inter-physician variability, and a clear familiarity effect, with the first task averaging 36 minutes and subsequent tasks 23 minutes [2104.04402].

Outside MT, the GEC PEET study shows the same general pattern. Editing original source sentences averages 31.16 seconds per sentence, whereas post-editing GECToR outputs averages 26.82 seconds and post-editing GEC-PD outputs 27.46 seconds, implying roughly four seconds saved per sentence by starting from a tool output [2510.04394]. A plausible implication is that temporal post-editing effort is not specific to translation workflows: it behaves similarly in MT, summarization-assisted note writing, and grammar correction whenever a human refines a machine-produced first draft.

## 4. Determinants and predictors of PEET

A central issue in PEET research is which observable variables track time reliably. The answer is mixed. In the English–Chinese QE study, higher human-rated MT quality significantly reduces time per word, and higher translator expertise also reduces time per word, but neither factor shows a significant interaction with the time-saving effect of sentence-level QE. Within the medium/high MT quality range studied, QE produces a roughly parallel downward shift in temporal effort across quality levels and across expertise groups [2507.16515].

By contrast, high-level automatic MT metrics are unstable predictors of time in strong NMT regimes. In English→Czech professional post-editing, better NMT systems clearly lead to fewer changes, but BLEU is “definitely not a stable predictor of the time or final output quality.” Linear fits of system-level BLEU against seconds per word change sign depending on which subset of systems is included, and review time can even increase with better BLEU [2109.05016]. The industrial LSP study reaches a parallel conclusion for edit distance: edit distance does not correlate strongly with speed and cannot be used as a reliable proxy for post-editing productivity in time, especially once TM matches and repetitions are included [2312.12660].

Task-based metrics that compare machine outputs with actual post-edited outputs perform better. Using PETpW as gold standard, task-based PE metrics such as HTER, HBLEU, HMETEOR, and especially Keys/char track post-editing time more closely than Direct Assessment, while reference-based BLEU, TER, and METEOR perform worst [1910.06204]. In English–Hindi post-editing, H-BLEU and H-chrF correlate negatively with temporal effort, H-TER correlates positively, and the correlations with time are moderate, while average and initial pause duration show no meaningful correlation with these automatic scores [2112.09841].

Two later lines extend PEET prediction beyond classical MT metrics. “Translator2Vec” learns document-session and editor embeddings from action sequences that include waiting times, jumps, and mouse actions; adding dynamic editor embeddings raises Pearson correlation for time prediction from 17.58 to 47.69 on En–Fr test data and from 23.67 to 38.72 on En–De test data [1907.10362]. “Assessing Human Editing Effort on LLM-Generated Texts via Compression-Based Edit Distance” introduces a directional LZ77-based compression distance \(d(S \to T)=\text{LZ}(S \mid T)-\text{LZ}(S)\), which correlates with human edit time up to 0.81 on accounting QA edits and predicts time better than traditional edit-distance and overlap metrics in that setting [2412.17321]. Taken together, these results suggest that PEET is better predicted by task-grounded process signals and post-edit transformations than by generic reference similarity alone.

## 5. Human factors, interface effects, and workflow structure

PEET is strongly shaped by editor behavior. The consultation-note study reports large inter-physician variability: one physician writes shorter, terser notes and edits only substantial issues, while another edits generated notes extensively and therefore shows higher editing times and higher counts of omissions and incorrectness in scoring [2104.04402]. The industrial HT/PE study finds the same pattern at scale, with large dispersion both by job and by linguist, and warns that a single “average PE speed” is a poor predictor of any given linguist’s performance [2312.12660]. DivEMT adds a cross-lingual dimension: productivity gains depend not only on system quality but also on typological relatedness to English and target morphological complexity, with Italian and Dutch benefiting much more from mBART-50 than Turkish or Arabic [2205.12215].

Interface interventions can change PEET, but not uniformly. In QE4PE, word-level QE highlights are studied under NoHighlight, Oracle, supervised XCOMET-XXL, and uncertainty-based conditions. Domain, language, and editors’ speed are critical factors in determining highlights’ effectiveness, and the downstream differences between human-made and automated highlights are modest [2503.03044]. The same paper reports a marked gap between objective and subjective utility: highlights improved correction of critical errors, but many editors described them as eye distractions and not accurate enough to rely on [2503.03044]. A plausible implication is that PEET should not be reduced to elapsed time alone when interface design changes the distribution of cognitive attention.

Related work in interactive MT shows a neighboring measurement regime based on operation counts rather than wall-clock time. “Online Learning for Effort Reduction in Interactive Neural Machine Translation” measures TER and KSMR, not human time, and explicitly frames them as effort surrogates. Because response times are around 0.1 to 0.3 seconds and learning times around 0.1 seconds, the study argues that system latency is close to real-time usability, making human interaction costs dominant [1802.03594]. This suggests that operation-based effort metrics remain relevant where direct timing is unavailable, but also that their relationship to PEET depends on interface latency being negligible.

## 6. Evaluation status, limitations, and recurrent misconceptions

PEET occupies an unusual position among evaluation constructs: it is expensive, highly informative, and difficult to standardize. The GEC PEET study presents time-to-correct as a human-focused evaluation scorer for ranking tools by estimated post-editing time [2510.04394], while the MT effort-estimation study argues that task-based measurements are the most reliable when estimating MT quality for a specific task, precisely because they are obtained during or after performing that task [1910.06204]. Across both lines, a recurring conclusion is that temporal effort is more directly tied to usability than reference overlap.

Several misconceptions are repeatedly challenged. One is that better automatic MT scores should translate monotonically into lower PEET. High-quality NMT studies show that this fails once systems cluster near the top: BLEU improvements do not map stably to time savings [2109.05016]. Another is that edit distance is an adequate proxy for time. Large industrial data show weak-to-moderate correlations only, insufficient for robust productivity estimation [2312.12660]. A third is that faster editing necessarily means lower overall human burden. In clinical note post-editing, one physician reports that post-editing may be faster while imposing higher cognitive load because every part of the generated text must be critically evaluated [2104.04402].

The limitations are correspondingly broad. Timing protocols differ across tools, including active-time logging, manual screen review, and survey timers; some systems fail to record time for untouched segments unless forced edits are introduced [2507.16515]. Many datasets remain narrow in domain, language, or participant pool, such as English–Chinese student translators, English–Hindi news, or English-only learner-text correction [2507.16515]. Sample-size constraints, limited low-quality MT coverage, and interface-specific behaviors further complicate generalization [2510.04394]. For this reason, PEET is best understood not as a single universal number but as a family of task-grounded temporal measures whose interpretation depends on editor population, text type, system quality, and logging design.

A broad synthesis follows from the literature. PEET is most informative when it is measured directly, normalized transparently, and analyzed jointly with technical and quality variables. It is least reliable when inferred from generic reference-based metrics alone. The recent explicit naming of PEET formalizes a construct that earlier post-editing research had already treated as central: the time humans actually spend making machine output usable [2510.04394].

Source: https://www.emergentmind.com/topics/post-editing-effort-in-time-peet