2.1 Political Communities
2.1 The 114th US Senate and Clustering
Another focus today will be to introduce you to the problems of datasets, which often lack the quality to be processed easily. For instance, you will learn various strategies for dealing with missing entries by removing them completely or trying to replicate their values. The kind of social data we are dealing with is vast and unorganized, which makes organizing it for analysis no easy task. In reality, you will spend most of your time working through such data challenges.
Finally, today will be dedicated to data exploration and the insights you can gain here. Exploring data is not necessarily a very structured part of your work.
You might remember the power of Pandas functions like head(), describe(), etc., to quickly explore essential components of data. Some people consider data exploration the most important part of the data analysis process, as it is essential to understand each aspect of the data and how it is represented. In social and cultural analytics, most of our work is based on data exploration techniques rather than prediction. We will cover prediction later in the course.
In today's notebook in Google Colab, we concentrate on introducing the power of digital methodologies and data exploration using a particular method called clustering, which is closely related to the understanding of political and social communities. We will look at the basics of clustering that delivers you powerful results quickly. In particular, we will use the k-means algorithm, invented by MacQueen in the late 1960s (https://en.wikipedia.org/wiki/K-means_clustering).
In the first exercise, we will use k-means to understand voting behaviour in the US Senate. The data is a subset of the data from https://www.dataquest.io/blog/k-means-clustering-us-senators/.
The 2014 elections gave the Republicans control of the Senate (and control of both houses of Congress) for the first time since the 109th Congress. With 247 seats in the House of Representatives and 54 seats in the Senate, this Congress began with the largest Republican majority since the 71st Congress of 1929–1931. There are 23 Democrats, 1 Independent and 33 Republicans in our dataset. Please note that this represents not the entire 114th Congress but a sample.
REFERENCES
- K-means clustering. (2023). In Wikipedia. https://en.wikipedia.org/w/index.php?title=K-means_clustering&oldid=1138258715