Perseus is a dataset for Cross-Lingual Summarization (CLS) which collects about 94K Chinese scientific documents paired with English summaries. The average length of documents in Perseus is more than two thousand tokens.
Source: Long-Document Cross-Lingual SummarizationPaper | Code | Results | Date | Stars |
---|