Analyzing newspaper articles
Introduction
A couple of years ago, my close friend @indraniel and I sat down together to hack on some project. We decided to download a bunch of newspaper articles from India and analyze them using natural language processing techniques. We never got around to much analyzing, however we did manage to figure out how to download batches of articles. If you’re interested, the corpus of articles is here and the code we put together is here. The code might be helpful to anyone looking to download text from newspaper articles.
Motivation
Today, while on a walk, I thought about going back and looking at this corpus of articles. This post documents the small analysis that I’ve done with the data. This might be an iterative post where I envision adding some analysis later, time permitting.
Results
How big is this dataset?
find . -name *json | wc -l
17980
That’s 17,000 JSON files. There’s text for each article and some metadata associated with the text.
When are these articles from?
These are from January 2016 to end of March 2016, a span of three months.