Evaluation datasets and reproducible scoring for long-document, meme, instruction-following and syllable-controlled translation.