{"id":11616,"date":"2021-01-03T12:47:08","date_gmt":"2021-01-03T11:47:08","guid":{"rendered":"http:\/\/flaven.fr\/?p=11616"},"modified":"2026-09-16T10:11:59","modified_gmt":"2026-09-16T08:11:59","slug":"web-scraping-beautifulsoup-selenium-various-explorations-in-web-scraping-with-python-and-jumping-timidly-in-surveillance-capitalism","status":"publish","type":"post","link":"https:\/\/flaven.fr\/2021\/01\/web-scraping-beautifulsoup-selenium-various-explorations-in-web-scraping-with-python-and-jumping-timidly-in-surveillance-capitalism\/","title":{"rendered":"Web scraping, Beautifulsoup, Selenium &#8211; Various explorations in web scraping with Python and jumping timidly in surveillance capitalism"},"content":{"rendered":"<p>I was in a hurry, to unleash the last post of the year&#8217;s 2020. Finally, I postponed the publication as most of the issues addressed in this post are so 2021!<br \/>\nTo make it short, last year, I was asked by a friend to apply web scrapping in a competitors\u2019 spying campaign. I should say Business Intelligence, it is more corporate and politically correct! Anyway, it does not surprise me so much that learning Testing, Automation and Python will lead me to this area partially included in Surveillance Capitalism Phenomenon. After all, this is still information. Precisely, in my case, the guy wanted solely a powerful automation that will collect, process and order a lot of public information available on webpages. [SIC]<\/p>\n<p>I won&#8217;t tell you, how, what and why but let me just say that with the help of Python, NLP, Web scraping&#8230; I can ensure you that we pinned down the crook.<\/p>\n<p>In conclusion, transparency dictatorship that prevails in surveillance capitalism has its best days ahead of it. Indeed, Digital has become a Personal Data Motherload!<\/p>\n<p>On a more personal POV \ud83d\ude42 Here is what I learnt from this experience: <\/p>\n<ol>\n<li>It has become tremendously easy to &#8220;hijack&#8221; or divert technologies from their first uses to make a Surveillance Swiss Army Knife even for an absolute beginner like me.<\/li>\n<li>Learning is rewarding but practising is even better. With this usecase, I was suddenly no longer facing theory but a complex reality. I took it as an opportunity to test my solving problem capability. That was a true exercise to practice what I have just learned (Python mostly).<\/li>\n<\/ol>\n<p><b>Hope that this quick and mundane introduction did not have a chilling effect on you! Here is some commands and scripts about Web scraping, all in Python. I have purposely excluded the things made with NLP or using JavaScript End to End Testing Framework such as CodeceptJS or Cypress that are beyond the scope of this post.<\/b><\/p>\n<p><b><a href=\"https:\/\/github.com\/bflaven\/BlogArticlesExamples\/tree\/master\/webscraping_with_python\" target=\"_blank\" rel=\"noopener noreferrer\">You can get all code and files on my github account in webscraping_with_python<\/a><\/b><\/p>\n<p><b>By the way, a great source of inspiration, all the works of Al Sweigart @<a href=\"https:\/\/inventwithpython.com\/\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/inventwithpython.com\/<\/a>. &#8220;Automate the Boring Stuff with Python&#8221; has been my recent bedside reading!<\/b><\/p>\n<h3>Some elements about Web Scraping<\/h3>\n<p>Here are various examples that automate the browser usage in order to get a map on Google Map or perform a search on Google, or download and save file. I made some attempts to grab HTML content with the help of BeautifulSoup, so all filenames with the pattern: slurpThatSoup_1.py, slurpThatSoup_2.py&#8230; are handling a BeautifulSoup Object from HTML<\/p>\n<p>FYI, BeautifulSoup (bs4) is one of the best Web Scraping library in Python is BeautifulSoup. You can find more on the official websites and documentation.<\/p>\n<ol>\n<li>Beautiful Soup: We called him Tortoise because he taught us.<a href=\"https:\/\/www.crummy.com\/software\/BeautifulSoup\/\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/www.crummy.com\/software\/BeautifulSoup\/<\/a><\/li>\n<li>Beautiful Soup Documentation: <a href=\"https:\/\/www.crummy.com\/software\/BeautifulSoup\/bs4\/doc\/\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/www.crummy.com\/software\/BeautifulSoup\/bs4\/doc\/<\/a><\/li>\n<\/ol>\n<p>As a huge supporter of testing, I made an extended usage of selenium Module to control the Browser, leveraging on Selenium\u2019s WebDriver Methods to Finding HTML Elements.<br \/>\n<b>Even so, it is sometime easier to divert testing frameworks such as CodeceptJS to perform actions in browser instead of using Selenium.<\/b><\/p>\n<p><b>Due to space disk limitations, I am using Anaconda to build my personal environment on an external disk. See the following previous post about conda install at <a href=\"https:\/\/flaven.fr\/2020\/05\/python-anaconda-worpress-json-json-schema-messy-post-with-few-practices-and-feedback-from-my-p-o-experience\/\" target=\"_blank\" rel=\"noopener noreferrer\">Python, Anaconda, WordPress, JSON, JSON-SCHEMA \u2013 Messy post with few practices and feedback from my P.O experience.<\/a><\/b><\/p>\n<p><b>Commands to install the required python libraries<\/b><\/p>\n<pre lang=\"bash\">\r\n\r\n# install requests and beautifulsoup4\r\npip install requests\r\npip install beautifulsoup4\r\n\r\n# install pyperclip\r\npip install pyperclip\r\n\r\n# check if pyperclip is working\r\n# go to a python console\r\n>>> import pyperclip\r\n>>> import requests\r\n>>> import beautifulsoup4\r\n\r\n<\/pre>\n<p><b>Check if the installed librairies are working<\/b><\/p>\n<pre lang=\"python\">\r\n# check if pyperclip is working\r\n# go to a python console\r\n>>> import pyperclip\r\n>>> import requests\r\n>>> import beautifulsoup4\r\n\r\n# if you get no errors, you are fine...\r\n<\/pre>\n<p><b>Installing the selenium Module<\/b><\/p>\n<pre lang=\"python\">\r\n# required installation\r\npip3 install selenium\r\n\r\n# better using homebrew\r\nbrew install geckodriver\r\n\r\n# check version\r\ngeckodriver --version\r\n\r\n# which \r\nwhich geckodriver\r\n\r\n# Source: https:\/\/www.dev2qa.com\/how-to-resolve-webdriverexception-geckodriver-executable-needs-to-be-in-path\/\r\n\r\n<\/pre>\n<p><b>Conclusion : Web Scraping is just an appetizer, that&#8217;s called not see the wood for the trees. It gave me the occasion to deepen a subject that matters to me : AI models generating misinformation or giving illusions of meaning. This time, I was chasing a crook but next time I&#8217;ll be the liar or the crook with the help these technologies, who knows? Ethics are melting like the polar ice! See below in Read More section advanced posts on the subject.<\/b><\/p>\n<h2>Read more<\/h2>\n<ul>\n<li>Google collects a frightening amount of data about you. You can find and delete it now<br \/><a href=\"https:\/\/www.cnet.com\/how-to\/google-collects-a-frightening-amount-of-data-about-you-you-can-find-and-delete-it-now\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.cnet.com\/how-to\/google-collects-a-frightening-amount-of-data-about-you-you-can-find-and-delete-it-now\/<\/a><\/li>\n<li>Surveillance capitalism on Wikipedia<br \/><a href=\"https:\/\/en.wikipedia.org\/wiki\/Surveillance_capitalism\" target=\"_blank\" rel=\"noopener\">https:\/\/en.wikipedia.org\/wiki\/Surveillance_capitalism<\/a><\/li>\n<li>We read the paper that forced Timnit Gebru out of Google. Here\u2019s what it says.<br \/><a href=\"https:\/\/www.technologyreview.com\/2020\/12\/04\/1013294\/google-ai-ethics-research-paper-forced-out-timnit-gebru\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.technologyreview.com\/2020\/12\/04\/1013294\/google-ai-ethics-research-paper-forced-out-timnit-gebru\/<\/a><\/li>\n<li>A college kid\u2019s fake, AI-generated blog fooled tens of thousands. This is how he made it.<br \/><a href=\"https:\/\/www.technologyreview.com\/2020\/08\/14\/1006780\/ai-gpt-3-fake-blog-reached-top-of-hacker-news\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.technologyreview.com\/2020\/08\/14\/1006780\/ai-gpt-3-fake-blog-reached-top-of-hacker-news\/<\/a><\/li>\n<li>Facebook translates &#8216;good morning&#8217; into &#8216;attack them&#8217;, leading to arrest<br \/><a href=\"https:\/\/www.theguardian.com\/technology\/2017\/oct\/24\/facebook-palestine-israel-translates-good-morning-attack-them-arrest\" target=\"_blank\" rel=\"noopener\">https:\/\/www.theguardian.com\/technology\/2017\/oct\/24\/facebook-palestine-israel-translates-good-morning-attack-them-arrest<\/a><\/li>\n<li>This could lead to the next big breakthrough in common sense AI<br \/><a href=\"https:\/\/www.technologyreview.com\/2020\/11\/06\/1011726\/ai-natural-language-processing-computer-vision\/\" target=\"_blank\" rel=\"noopener\">https:\/\/www.technologyreview.com\/2020\/11\/06\/1011726\/ai-natural-language-processing-computer-vision\/<\/a><\/li>\n<li>100 must-read classic books, as chosen by our readers<br \/><a href=\"https:\/\/www.penguin.co.uk\/articles\/2018\/100-must-read-classic-books\/\"target=\"_blank\">https:\/\/www.penguin.co.uk\/articles\/2018\/100-must-read-classic-books\/<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>I was in a hurry, to unleash the last post of the year&#8217;s 2020. Finally, I postponed the publication as most of the issues addressed&hellip; <\/p>\n<p class=\"text-center\"><a href=\"https:\/\/flaven.fr\/2021\/01\/web-scraping-beautifulsoup-selenium-various-explorations-in-web-scraping-with-python-and-jumping-timidly-in-surveillance-capitalism\/\" class=\"more-link\">Continue reading &rarr; <span class=\"screen-reader-text\">Web scraping, Beautifulsoup, Selenium &#8211; Various explorations in web scraping with Python and jumping timidly in surveillance capitalism<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":11617,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"bf_ai_meta_description":"Web scraping with Python using BeautifulSoup and Selenium. Automate data collection, processing, and ordering for competitive intelligence.","bf_ai_og_title":"Web Scraping with Python","footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[3438,3437,3444,3447,3449],"tags":[2641,2532,2638,2642,2597,2598,3493,2593,2643,1697,2355,2591,2594,2596,2592,2316,2606,2607,2595,2315,2640,2639,2599,3508,2395],"class_list":["post-11616","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-machine-learning","category-business-case-studies","category-programming-databases","category-technology-trends","category-tutorials-how-to","tag-al-sweigart","tag-anaconda","tag-beautiful-soup","tag-beautifulsoup","tag-boring-stuff","tag-bullshit-jobs","tag-clip","tag-codeceptjs","tag-cypress","tag-experience","tag-geckodriver","tag-philosophy","tag-practical","tag-productivity","tag-programming","tag-python","tag-python2","tag-python3","tag-real-world","tag-selenium","tag-shoshana-zuboff","tag-surveillance-capitalism","tag-sweigart","tag-test-automation","tag-testing"],"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p3Vuhl-31m","jetpack_featured_media_url":"https:\/\/flaven.fr\/wp-content\/uploads\/2021\/01\/webscraping_with_python_b.jpg","_links":{"self":[{"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/posts\/11616","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/comments?post=11616"}],"version-history":[{"count":13,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/posts\/11616\/revisions"}],"predecessor-version":[{"id":11630,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/posts\/11616\/revisions\/11630"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/media\/11617"}],"wp:attachment":[{"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/media?parent=11616"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/categories?post=11616"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/flaven.fr\/happy-api\/wp\/v2\/tags?post=11616"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}