使用Selenium网页爬取巴塞罗那公证人邮箱时遭遇NoSuchElementException异常的求助
Fixing Your NoSuchElementException & Web Scraping Script for Barcelona Notaries
First, let's tackle the immediate error you're facing: that NoSuchElementException is happening because you're applying the .format(i) method to the result of driver.find_element() instead of the XPath string itself. It's a simple syntax mix-up, but there are a few other issues in your script that are keeping it from working as intended. Let's fix them one by one:
Key Issues in Your Current Code
- Incorrect XPath Formatting: You tried to format the XPath after passing it to
find_element(), which doesn't work. The.format()needs to be called on the string literal before passing it as the XPath argument. - Mixing Selenium and Requests: Selenium handles dynamic browser interactions (clicking buttons, navigating pages), but you're using
requests.get()to fetch the page again—this won't capture the dynamic changes from your clicks (like revealed email addresses). - Buggy Pagination Logic: Your item loop and pagination button selection have inconsistent indexing, and the way you're building
next_urldoesn't align with how the site's pagination works. - Unnecessary Module Imports: You imported
retwice, and several other modules (likeBeautifulSoup,pandas,googlesearch) aren't being used at all—you can clean those up. - Incorrect Loop Range: You're looping from 0 to 6 for the 5 results per page, which will try to access a non-existent
heading5element.
Corrected Code
Here's the revised script that fixes all these issues and properly scrapes the email addresses:
import re import time from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys # Initialize the browser driver = webdriver.Safari() emails_notarios = [] base_url = "https://www.notariado.org/portal/" # Navigate to the portal and search for Barcelona driver.get(base_url) time.sleep(1) # Wait for page to load # Locate the search input for location location_input = driver.find_element(By.XPATH, '//*[@id="valor4"]') location_input.clear() location_input.send_keys("Barcelona") time.sleep(1) # Click the search button search_button = driver.find_element(By.XPATH, '//*[@id="portlet_com_liferay_journal_content_web_portlet_JournalContentPortlet_INSTANCE_nWpBPvnHOnbm"]/div/div[2]/div/div[2]/div/button') search_button.click() time.sleep(2) # Give time for results to load # Handle pagination (assuming up to 26 pages, adjust if needed) for page_num in range(1, 27): # Wait for the page results to load time.sleep(2) # Iterate over the 5 results per page (heading0 to heading4) for result_idx in range(0, 5): try: # Correctly format the XPath for the reveal button reveal_button = driver.find_element(By.XPATH, f'//*[@id="heading{result_idx}"]/div/h4/span') reveal_button.click() time.sleep(1) # Wait for email to show # Extract email from the current page source page_source = driver.page_source emails = re.findall(r'\b[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}\b', page_source, re.I) for email in emails: if email not in emails_notarios: emails_notarios.append(email) except Exception as e: # Skip if the element doesn't exist (e.g., last page with fewer results) print(f"Error accessing result {result_idx} on page {page_num}: {str(e)}") continue # Move to the next page, if it exists try: # Correctly format the pagination button XPath next_page_button = driver.find_element(By.XPATH, f'//*[@id="jwpg_pagination"]/ul/li[{page_num + 1}]/a') next_page_button.click() time.sleep(2) except Exception as e: print(f"End of pagination or error moving to page {page_num + 1}: {str(e)}") break # Print all collected emails print("Collected Email Addresses:") for email in emails_notarios: print(email) # Close the browser driver.quit()
What Changed?
- Fixed XPath Formatting: Used f-strings (or you could use
.format()) to correctly insert the index into the XPath string before passing it tofind_element(). - Removed Unused Modules: Cleaned up imports to only include what's necessary.
- Scraped Directly from Selenium's Page: Instead of using
requests, we pull the page source directly from the browser after clicking the reveal button, so we get the dynamic content with emails. - Simplified Pagination: The loop now iterates over page numbers directly, and we handle the next page button with correct indexing.
- Added Error Handling: Added try-except blocks to skip over missing elements (like if a page has fewer than 5 results) and handle pagination end gracefully.
- Avoided Duplicate Emails: Check if an email is already in the list before adding it.
Additional Tips
- Adjust Sleep Times: Depending on your internet speed, you might need to increase the
time.sleep()values to ensure pages/elements load fully. Alternatively, use Selenium'sWebDriverWaitwith expected conditions for more reliable waits (instead of fixed sleeps). - Respect Robots.txt: Make sure you're complying with the site's
robots.txtand terms of service when scraping.
内容的提问来源于stack exchange,提问作者AAlbiol
相关产品推荐
相关产品推荐

