Skip to content
Advertisement
MultimodalTabularText

SEC 10-K Pre-processed

SEC 10-K Pre-processed

SEC 10-K Pre-processed This dataset contains pre-processed SEC 10-K English text with sentence/chunk/paragraph files for translation-oriented use. Source Original dataset: PleIAs/SEC This release is a pre-processed derivative of that source (cleaning, filtering, and deduplication). Splits / Files sentence track.jsonl: 1,620,583 rows chunk track.jsonl: 914,217 rows paragraph track.jsonl: 5,367,690 rows translation dataset.jsonl: 7,902,490 rows (concatenated)… See the full description on the dataset page:

Source: Hugging Face Hub (alwaysgood/sec-10k-pre-processed). Metadata imported from the dataset’s Hub tags.

Advertisement