Selenium Web Scraping

Redouane Niboucha 5 min read Updated nov 2020 Web Scraping Controlling a web browser from a program can be useful in many scenarios, example use cases are website text automation and web scraping, a very popular framework for this kind of automation is Selenium WebDriver. This selenium tutorial is designed for beginners to learn how to use the python selenium module to perform web scraping, web testing and create website bots. Using Selenium v3.x opening a website in a New Tab through Python is much easier now. We have to induce an WebDriverWait for numberofwindowstobe(2) and then collect the window handles every time we open a new tab/window and finally iterate through the window handles and switchTo.window(newlyopened) as required. Web Scraping JavaScript Generated Pages with Python. This project was created just for educational proposes. The code shows how to do web scraping dynamic content pages generated from Javascript using Python and Selenium. We use as data the NBA site to extract stats information from players and generate a json file with some top 10 rankings. Jan 13, 2019 In this in depth tutorial series, you will learn how to use Selenium + Python to crawl and interact with almost any websites. Selenium is a Web Browser Automation Tool originally designed to.

Imagine what would you do if you could automate all the repetitive and boring activities you perform using internet, like checking every day the first results of Google for a given keyword, or download a bunch of files from different websites.

In this post you’ll learn to use Selenium with Python, a Web Scraping tool that simulates a user surfing the Internet. For example, you can use it to automatically look for Google queries and read the results, log in to your social accounts, simulate a user to test your web application, and anything you find in your daily live that it’s repetitive. The possibilities are infinite! 🙂

*All the code in this post has been tested with Python 2.7 and Python 3.4.

Install and use Selenium

Selenium is a python package that can be installed via pip. I recommend that you install it in a virtual environment (using virtualenv and virtualenvwrapper).

To install selenium, you just need to type:

In this post we are going to initialize a Firefox driver — you can install it by visiting their website. However, if you want to work with Chrome or IE, you can find more information here.

Once you have Selenium and Firefox installed, create a python file, selenium_script.py. We are going to initialize a browser using Selenium:

</div><table><tbody><tr><td><div><div>2</div><div>4</div><div>6</div><div>8</div><div>10</div><div>12</div><div>14</div><div>16</div><div>18</div><div>20</div><div>22</div><div>24</div><div>26</div><div>28</div><div>30</div><div>32</div></div></td><td><div><div><span>from </span><span>selenium </span><span>import </span><span>webdriver</span></div><div><span>from </span><span>selenium</span><span>.</span><span>webdriver</span><span>.</span><span>support</span><span>.</span><span>ui </span><span>import </span><span>WebDriverWait</span></div><div><span>from </span><span>selenium</span><span>.</span><span>webdriver</span><span>.</span><span>support </span><span>import </span><span>expected_conditions </span><span>as</span><span>EC</span></div><div><span>from </span><span>selenium</span><span>.</span><span>common</span><span>.</span><span>exceptions </span><span>import </span><span>TimeoutException</span></div><div><span>driver</span><span>=</span><span>webdriver</span><span>.</span><span>Firefox</span><span>(</span><span>)</span></div><div><span>return</span><span>driver</span></div><div><span>driver</span><span>.</span><span>get</span><span>(</span><span>'http://www.google.com'</span><span>)</span></div><div><span>box</span><span>=</span><span>driver</span><span>.</span><span>wait</span><span>.</span><span>until</span><span>(</span><span>EC</span><span>.</span><span>presence_of_element_located</span><span>(</span></div><div><span>button</span><span>=</span><span>driver</span><span>.</span><span>wait</span><span>.</span><span>until</span><span>(</span><span>EC</span><span>.</span><span>element_to_be_clickable</span><span>(</span></div><div><span>box</span><span>.</span><span>send_keys</span><span>(</span><span>query</span><span>)</span></div><div><span>except </span><span>TimeoutException</span><span>:</span></div><div><span>if</span><span>__name__</span><span>'__main__'</span><span>:</span></div><div><span>lookup</span><span>(</span><span>driver</span><span>,</span><span>'Selenium'</span><span>)</span></div><div><span>driver</span><span>.</span><span>quit</span><span>(</span><span>)</span></div></div></td></tr></tbody></table><p>In the previous code:</p><ul><li> the function <span>init_driver</span> initializes a driver instance.<ul><li> creates the driver instance</li><li> adds the <span>WebDriverWait</span> function as an attribute to the driver, so it can be accessed more easily. This function is used to make the driver wait a certain amount of time (here 5 seconds) for an event to occur.</li></ul></li><li> the function <span>lookup</span> takes two arguments: a driver instance and a query lookup (a string).<ul><li> it loads the Google search page</li><li> it waits for the query box element to be located and for the button to be clickable. Note that we are using the <span>WebDriverWait</span> function to wait for these elements to appear.</li><li> Both elements are located by name. Other options would be to locate them by <span><span>ID</span><span>,</span><span>XPATH</span><span>,</span><span>TAG_NAME</span><span>,</span><span>CLASS_NAME</span><span>,</span><span>CSS_SELECTOR</span></span> , etc (see table below). You can find more information here.</li><li> Next, it sends the query into the box element and clicks the search button.</li><li> If either the box or button are not located during the time established in the wait function (here, 5 seconds), the <span>TimeoutException</span> is raised.</li></ul></li><li> the next statement is a conditional that is true only when the script is run directly. This prevents the next statements to run when this file is imported.<ul><li> it initializes the driver and calls the lookup function to look for “Selenium”.</li><li> it waits for 5 seconds to see the results and quits the driver</li></ul></li></ul><p>Finally, run your code with:</p><p>Did it work? If you got an <span>ElementNotVisibleException</span> , keep reading!</p><h2>How to catch an ElementNotVisibleExcpetion</h2><p>Google search has recently changed so that initially, Google shows this page:</p><p>and when you start writing your query, the search button moves into the upper part of the screen.</p><p>Well, actually it doesn’t move. The old button becomes invisible and the new one visible (and thus the exception when you click the old one: it’s not visible to click!).</p><p>We can update the lookup function in our code so that it catches this exception:</p><div><textarea wrap='soft' readonly='>from selenium.common.exceptions import ElementNotVisibleException def lookup(driver, query): driver.get('http://www.google.com') try: box = driver.wait.until(EC.presence_of_element_located( (By.NAME, 'q'))) button = driver.wait.until(EC.element_to_be_clickable( (By.NAME, 'btnK'))) box.send_keys(query) try: button.click() except ElementNotVisibleException: button = driver.wait.until(EC.visibility_of_element_located( (By.NAME, 'btnG'))) button.click() except TimeoutException: print('Box or Button not found in google.com')

from selenium.common.exceptions import ElementNotVisibleException

def lookup(driver,query):

try:

box=driver.wait.until(EC.presence_of_element_located(

button=driver.wait.until(EC.element_to_be_clickable(

box.send_keys(query)

button.click()

button=driver.wait.until(EC.visibility_of_element_located(

button.click()

print('Box or Button not found in google.com')

the element that raised the exception, button.click() is inside a try statement.
if the exception is raised, we look for the second button, using visibility_of_element_located to make sure the element is visible, and then click this button.
if at any time, some element is not found within the 5 second period, the TimeoutException is raised and caught by the two end lines of code.
Note that the initial button name is “btnK” and the new one is “btnG”.

Method list in Selenium

To sum up, I’ve created a table with the main methods used here.

Note: it’s not a python file — don’t try to run/import it 🙂

</div><table><tbody><tr><td><div><div>2</div><div>4</div><div>6</div><div>8</div><div>10</div><div>12</div><div>14</div><div>16</div><div>18</div><div>20</div><div>22</div><div>24</div><div>26</div><div>28</div><div>30</div><div>32</div></div></td><td><div><div><span>from </span><span>selenium </span><span>import </span><span>webdriver</span></div><div><span>from </span><span>selenium</span><span>.</span><span>webdriver</span><span>.</span><span>support</span><span>.</span><span>ui </span><span>import </span><span>WebDriverWait</span></div><div><span>driver</span><span>=</span><span>webdriver</span><span>.</span><span>Firefox</span><span>(</span><span>)</span></div><div><span># WAIT FOR ELEMENTS</span></div><div><span>from </span><span>selenium</span><span>.</span><span>webdriver</span><span>.</span><span>support </span><span>import </span><span>expected_conditions </span><span>as</span><span>EC</span></div><div><span>element</span><span>=</span><span>driver</span><span>.</span><span>wait</span><span>.</span><span>until</span><span>(</span></div><div><span>EC</span><span>.</span><span>element_to_be_clickable</span><span>(</span></div><div><span>(</span><span>By</span><span>.</span><span>NAME</span><span>,</span><span>'name'</span><span>)</span></div><div><span>(</span><span>By</span><span>.</span><span>LINK_TEXT</span><span>,</span><span>'link text'</span><span>)</span></div><div><span>(</span><span>By</span><span>.</span><span>TAG_NAME</span><span>,</span><span>'tag name'</span><span>)</span></div><div><span>(</span><span>By</span><span>.</span><span>CSS_SELECTOR</span><span>,</span><span>'css selector'</span><span>)</span></div><div><span>)</span></div><div><span># CATCH EXCEPTIONS</span></div><div><span>TimeoutException</span></div></div></td></tr></tbody></table><p>That’s all! Hope it was useful! 🙂</p><p>Don’t forget to share it with your friends!</p><p>Imagine what would you do if you could automate all the repetitive and boring activities you perform using internet, like checking every day the first results of Google for a given keyword, or download a bunch of files from different websites.</p><p>In this post you’ll learn to use Selenium with Python, a Web Scraping tool that simulates a user surfing the Internet. For example, you can use it to automatically look for Google queries and read the results, log in to your social accounts, simulate a user to test your web application, and anything you find in your daily live that it’s repetitive. The possibilities are infinite! 🙂</p><p>*All the code in this post has been tested with Python 2.7 and Python 3.4.</p><h2>Install and use Selenium</h2><p>Selenium is a python package that can be installed via pip. I recommend that you install it in a virtual environment (using virtualenv and virtualenvwrapper).</p><p>To install selenium, you just need to type:</p><p>In this post we are going to initialize a Firefox driver — you can install it by visiting their website. However, if you want to work with Chrome or IE, you can find more information here.</p><p>Once you have Selenium and Firefox installed, create a python file, <span>selenium_script.py</span>. We are going to initialize a browser using Selenium:</p><div><textarea wrap='soft' readonly='>import time from selenium import webdriver driver = webdriver.Firefox() time.sleep(5) driver.quit()

from selenium import webdriver

driver=webdriver.Firefox()

driver.quit()

This just initializes a Firefox instance, waits for 5 seconds, and closes it.

Well, that was not very useful…

How about if we go to Google and search for something?

Web Scraping Google with Selenium

Let’s make a script that loads the main Google search page and makes a query to look for “Selenium”:

</div><table><tbody><tr><td><div><div>2</div><div>4</div><div>6</div><div>8</div><div>10</div><div>12</div><div>14</div><div>16</div><div>18</div><div>20</div><div>22</div><div>24</div><div>26</div><div>28</div><div>30</div><div>32</div></div></td><td><div><div><span>from </span><span>selenium </span><span>import </span><span>webdriver</span></div><div><span>from </span><span>selenium</span><span>.</span><span>webdriver</span><span>.</span><span>support</span><span>.</span><span>ui </span><span>import </span><span>WebDriverWait</span></div><div><span>from </span><span>selenium</span><span>.</span><span>webdriver</span><span>.</span><span>support </span><span>import </span><span>expected_conditions </span><span>as</span><span>EC</span></div><div><span>from </span><span>selenium</span><span>.</span><span>common</span><span>.</span><span>exceptions </span><span>import </span><span>TimeoutException</span></div><div><span>driver</span><span>=</span><span>webdriver</span><span>.</span><span>Firefox</span><span>(</span><span>)</span></div><div><span>return</span><span>driver</span></div><div><span>driver</span><span>.</span><span>get</span><span>(</span><span>'http://www.google.com'</span><span>)</span></div><div><span>box</span><span>=</span><span>driver</span><span>.</span><span>wait</span><span>.</span><span>until</span><span>(</span><span>EC</span><span>.</span><span>presence_of_element_located</span><span>(</span></div><div><span>button</span><span>=</span><span>driver</span><span>.</span><span>wait</span><span>.</span><span>until</span><span>(</span><span>EC</span><span>.</span><span>element_to_be_clickable</span><span>(</span></div><div><span>box</span><span>.</span><span>send_keys</span><span>(</span><span>query</span><span>)</span></div><div><span>except </span><span>TimeoutException</span><span>:</span></div><div><span>if</span><span>__name__</span><span>'__main__'</span><span>:</span></div><div><span>lookup</span><span>(</span><span>driver</span><span>,</span><span>'Selenium'</span><span>)</span></div><div><span>driver</span><span>.</span><span>quit</span><span>(</span><span>)</span></div></div></td></tr></tbody></table><p>In the previous code:</p><ul><li> the function <span>init_driver</span> initializes a driver instance.<ul><li> creates the driver instance</li><li> adds the <span>WebDriverWait</span> function as an attribute to the driver, so it can be accessed more easily. This function is used to make the driver wait a certain amount of time (here 5 seconds) for an event to occur.</li></ul></li><li> the function <span>lookup</span> takes two arguments: a driver instance and a query lookup (a string).<ul><li> it loads the Google search page</li><li> it waits for the query box element to be located and for the button to be clickable. Note that we are using the <span>WebDriverWait</span> function to wait for these elements to appear.</li><li> Both elements are located by name. Other options would be to locate them by <span><span>ID</span><span>,</span><span>XPATH</span><span>,</span><span>TAG_NAME</span><span>,</span><span>CLASS_NAME</span><span>,</span><span>CSS_SELECTOR</span></span> , etc (see table below). You can find more information here.</li><li> Next, it sends the query into the box element and clicks the search button.</li><li> If either the box or button are not located during the time established in the wait function (here, 5 seconds), the <span>TimeoutException</span> is raised.</li></ul></li><li> the next statement is a conditional that is true only when the script is run directly. This prevents the next statements to run when this file is imported.<ul><li> it initializes the driver and calls the lookup function to look for “Selenium”.</li><li> it waits for 5 seconds to see the results and quits the driver</li></ul></li></ul><p>Finally, run your code with:</p><p>Did it work? If you got an <span>ElementNotVisibleException</span> , keep reading!</p><h2>How to catch an ElementNotVisibleExcpetion</h2><p>Google search has recently changed so that initially, Google shows this page:</p><p>and when you start writing your query, the search button moves into the upper part of the screen.</p><p>Well, actually it doesn’t move. The old button becomes invisible and the new one visible (and thus the exception when you click the old one: it’s not visible to click!).</p><p>We can update the lookup function in our code so that it catches this exception:</p><img src='https://i.ytimg.com/vi/8aedS8sQ9lg/maxresdefault.jpg' alt='Selenium web scraping example' title='Selenium web scraping example' /><div><textarea wrap='soft' readonly='>from selenium.common.exceptions import ElementNotVisibleException def lookup(driver, query): driver.get('http://www.google.com') try: box = driver.wait.until(EC.presence_of_element_located( (By.NAME, 'q'))) button = driver.wait.until(EC.element_to_be_clickable( (By.NAME, 'btnK'))) box.send_keys(query) try: button.click() except ElementNotVisibleException: button = driver.wait.until(EC.visibility_of_element_located( (By.NAME, 'btnG'))) button.click() except TimeoutException: print('Box or Button not found in google.com')

from selenium.common.exceptions import ElementNotVisibleException

def lookup(driver,query):

try:

box=driver.wait.until(EC.presence_of_element_located(

button=driver.wait.until(EC.element_to_be_clickable(

box.send_keys(query)

button.click()

button=driver.wait.until(EC.visibility_of_element_located(

button.click()

print('Box or Button not found in google.com')

the element that raised the exception, button.click() is inside a try statement.
if the exception is raised, we look for the second button, using visibility_of_element_located to make sure the element is visible, and then click this button.
if at any time, some element is not found within the 5 second period, the TimeoutException is raised and caught by the two end lines of code.
Note that the initial button name is “btnK” and the new one is “btnG”.

Method list in Selenium

To sum up, I’ve created a table with the main methods used here.

Note: it’s not a python file — don’t try to run/import it 🙂

</div><table><tbody><tr><td><div><div>2</div><div>4</div><div>6</div><div>8</div><div>10</div><div>12</div><div>14</div><div>16</div><div>18</div><div>20</div><div>22</div><div>24</div><div>26</div><div>28</div><div>30</div><div>32</div></div></td><td><div><div><span>from </span><span>selenium </span><span>import </span><span>webdriver</span></div><div><span>from </span><span>selenium</span><span>.</span><span>webdriver</span><span>.</span><span>support</span><span>.</span><span>ui </span><span>import </span><span>WebDriverWait</span></div><div><span>driver</span><span>=</span><span>webdriver</span><span>.</span><span>Firefox</span><span>(</span><span>)</span></div><div><span># WAIT FOR ELEMENTS</span></div><div><span>from </span><span>selenium</span><span>.</span><span>webdriver</span><span>.</span><span>support </span><span>import </span><span>expected_conditions </span><span>as</span><span>EC</span></div><div><span>element</span><span>=</span><span>driver</span><span>.</span><span>wait</span><span>.</span><span>until</span><span>(</span></div><div><span>EC</span><span>.</span><span>element_to_be_clickable</span><span>(</span></div><div><span>(</span><span>By</span><span>.</span><span>NAME</span><span>,</span><span>'name'</span><span>)</span></div><div><span>(</span><span>By</span><span>.</span><span>LINK_TEXT</span><span>,</span><span>'link text'</span><span>)</span></div><div><span>(</span><span>By</span><span>.</span><span>TAG_NAME</span><span>,</span><span>'tag name'</span><span>)</span></div><div><span>(</span><span>By</span><span>.</span><span>CSS_SELECTOR</span><span>,</span><span>'css selector'</span><span>)</span></div><div><span>)</span></div><div><span># CATCH EXCEPTIONS</span></div><div><span>TimeoutException</span></div></div></td></tr></tbody></table><p>That’s all! Hope it was useful! 🙂</p><h3>Selenium Web Scraping Example</h3><p>Don’t forget to share it with your friends!</p><br><br><a href='https://netlify.mix-goapp.com/hard-drive-cleaner-for-mac.html#CEKYbMrdc=UAVBSlNSDQMGFAxRAFUHUVVfBQ1KQgBZXV5dFloSTlJWH0RbQlUUDwoGFFUBVB4GFUdAWgFDBFgTXVUAHAoVGwEaBQgCBUhTSFMUAV5IZ2UVDgQDSF8AQVRZUhoZWElHGBhDXUhAF0NXAB1XUTY=' target='_blank'><img src='https://cdn-ak.f.st-hatena.com/images/fotolife/r/ruriatunifoefec/20200910/20200910011341.png' style='cursor:pointer;display:block;margin-left:auto;margin-right:auto;'></a><br><br></p>
			</div>
	<footer class="entry-meta">
		
			</footer>
</article>
				<nav role="navigation" id="nav-below" class="post-navigation">
		<h1 class="screen-reader-text">Post navigation</h1>
	
		<div class="nav-previous"><a href='/evernote-without-account'>Evernote Without Account</a></div>		<div class="nav-next"><a href='/nearest-royal-mail-post-box-to-me'>Nearest Royal Mail Post Box To Me</a></div>
	
	</nav>
	
			
		
		</main>
	</div>
	<div id="secondary" class="widget-area col-md-3" role="complementary">
				<aside id="search-3" class="widget widget_search"><form role="search" method="get" class="row search-form" action="#">
	<label>
		<span class="screen-reader-text">Search for:</span>
		<input type="search" class="search-field" placeholder="Search …" value="" name="s">
	</label>
	<input type="submit" class="search-submit" value="Search">
</form>
</aside>		<aside id="recent-posts-5" class="widget widget_recent_entries">		<h1 class="widget-title">MOST POPULAR ARTICLES</h1>		
					<li><a href='/nintendo-lite'>: Nintendo Lite</a></li>
<li><a href='/fiesta-crossover'>: Fiesta Crossover</a></li>
<li><a href='/zeplin-for-jira'>: Zeplin For Jira</a></li>
<li><a href='/rstudio'>: RStudio</a></li>
<li><a href='/grammarly-microsoft-edge'>: Grammarly Microsoft Edge</a></li>
				
		</aside>			</div>
</div>
	<footer id="colophon" class="site-footer container row" role="contentinfo">
		<div id="footertext" class="col-md-7">
        	        </div> 
		<div class="site-info col-md-5">
			 
<a href="/" title="Foundryfox508">Foundryfox508</a>
		</div>
		  
	</footer>
</div>
</body>
</html>