📣Check out our workshops at #EAMT2023! Conference attendees can deep-dive into Gender-Inclusive Translation Technologies, Automated Translation for Sign and Spoken Languages, Open Community-Driven #MT, or Multi3 Language Generation. Plus tutorials! More: https://t.co/CZcQC3M3vN
Interested in running a workshop at @MTSummitXVII (19-23 August 2019 in Dublin)? You can find all the information for submission at https://t.co/iNqXYytTYf. Deadline for workshop proposals: Friday, 18th January 2019. #mtsummit2019
Language pairs for #WMT2019 news task: English <-> Chinese, German, Gujarati and Russian. Hoping to confirm 1-2 more. Also aiming for unidirectional test sets, and document context in evaluation. #wmt2018
@ian_soboroff@yoavgo We measure this based on the human perceived quality our annotators assign to the human translations, which are part of the evaluation campaigns.
@ian_soboroff@yoavgo Same general eval setup, except for switch to source-based DA (which I compared to reference-based DA at IWSLT17). Works on same WMT17 test data, but we also added human translations to the set of systems to be evaluated. Same annotation tool.
@ian_soboroff@yoavgo Annotators were asked to judge candidate translations w.r.t. semantic transfer for given source segments. Multiple annotators worked on the same segments for redundancy. Human ability to translate was measured based on the quality annotators assigned to human translations.
@yoavgo@ian_soboroff Based on number of wins, we cluster systems to determine if there is a significant quality difference. Systems in same cluster are considered indistinguishable w.r.t. translation quality. This then denotes human parity, according to our Definition 2 in the paper (Section 2).
@yoavgo@ian_soboroff Scores are standardised for individual annotators and then averaged on segment and system level. We compare pairs of systems and count how often a system wins over other systems according to Wilcoxon rank sum test with p-level p <= 0.05, following WMT17.
@yoavgo Fair point. In the paper we refer to "news translation". The blog post mentions a "test set of news stories", but I agree that segment vs story level can easily be misunderstood. Would be interesting to look into this...
@yoavgo We compare against both post-edited and professional human translations. Context is another dimension to consider, correct. Evaluation still segment-based for most MT research papers.
Our speech recognition is moving to end to end #LSTM neural network architecture. Combined with an increase in speech data, this has improved quality up to 29%. Learn more about the @MSTranslator Speech API https://t.co/9c30pHrPsT