A subset of the Bangla version of the Wikipedia text

Context

A subset of the Bangla version of the Wikipedia text. To create the Wikipedia dataset, we collected the Bangla wiki-dump of 10th June, 2019. The files are then merged and each article is selected as a sample text. All HTML tags were removed and the title of the page was stripped from the beginning of the text. This dataset contains 70377 samples with a total number of words being 18229481. The entire dataset has 1289249 unique words, which is 7% of the total vocabulary.

Acknowledgements

Khatun, Aisha; Rahman, Anisur; Islam, Md Saiful (2020), “Bangla Wikipedia dataset”, Mendeley Data, V4, doi: 10.17632/3ph3n78fp7.4

Related Datasets

Large Sentiment Analysis Bangla Dataset

@kaggle
Yahoo Finance Historical Prices And Ticker Fundamentals

@yahoo
Nuclear Weapons Proliferation

@owid
SFC2014 - REACT EU Overview Allocation Vs Decided

@esifunds
Eucalyptus Growth And Environmental Data

@euremarkable
Trust Questions In The European Social Survey, Latinobarómetro And Afrobarometer

@owid

Large Sentiment Analysis Bangla Dataset

Yahoo Finance Historical Prices And Ticker Fundamentals

Nuclear Weapons Proliferation

SFC2014 - REACT EU Overview Allocation Vs Decided

Eucalyptus Growth And Environmental Data

Trust Questions In The European Social Survey, Latinobarómetro And Afrobarometer