CapabilityCognitive automation2024-07-31
ByteDance researchers report their CLASI system scored 81.3% and 78.0% on their own quality metric for Chinese–English simultaneous interpreting of real-world speech
Interpreteroccupation page →Event date / reported
2024-07-31
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Conference and simultaneous interpreting
Interpreting speeches and meetings in real time from a booth or remotely.
Being augmented≈ Platform inference
Where this applies
A technology company's research paper on its own simultaneous speech translation system. The authors report that on real-world speeches, which are often disfluent, informal and unclear, CLASI achieves 81.3% and 78.0% on their valid information proportion metric for Chinese-to-English and English-to-Chinese. The metric and the parity threshold are the authors' own; it is not a head-to-head test in a real booth.
What this means
Machine simultaneous interpreting is getting close on its makers' own scale; whether that scale matches what listeners need is not established here.
What it does not yet show
A company scoring its own system on its own metric; not an independent evaluation.
What you can check
Open the arXiv paper 2407.21646 and find "81.3% and 78.0%".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Cheng, Huang, Ko, Li, Peng, Xu and Zhang (ByteDance) — Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent, arXiv:2407.21646 (submitted 31 Jul 2024) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
This source sells the thing it is describing
Product or template: ByteDance CLASI simultaneous speech translation
What the publisher gains from this: ByteDance built the system and defined the metric it is scored on; favourable results support its products.
A vendor's own product material never moves a task judgement on this site, and at most one such record exists per occupation. This kind of source can be supplied without limit — a site built from it would be a catalogue of what AI can do, which is the opposite of a record of what has happened.
Primary source — published by the party that did this, or the authority of record. No co-signature needed.