2.2 Web scraping

2.2 Web scraping

Zarah van Hout, Analysing Communities in Digital Society (Web Scraping) (YouTube)

The web has become a unique source of data for community analysis. We can find lots of data on the web. However, a big problem with web data is that it is often inconsistent and heterogeneous. One often has to visit multiple websites and assemble their data together to access it. Finally, the data is generally published without reuse in mind, which implies that the data can be of low quality. That said, the web is so vast that it still provides an often overwhelming source of exciting data.

Websites with HTML


Staying with our example of political communities, in the notebook in Google Colab, we will scrape data on the current composition of the US Senate from Wikipedia.

US Senate from Wikipedia

Check your knowledge here: