You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium网页爬取巴塞罗那公证人邮箱时遭遇NoSuchElementException异常的求助

Fixing Your NoSuchElementException & Web Scraping Script for Barcelona Notaries

First, let's tackle the immediate error you're facing: that NoSuchElementException is happening because you're applying the .format(i) method to the result of driver.find_element() instead of the XPath string itself. It's a simple syntax mix-up, but there are a few other issues in your script that are keeping it from working as intended. Let's fix them one by one:

Key Issues in Your Current Code

  • Incorrect XPath Formatting: You tried to format the XPath after passing it to find_element(), which doesn't work. The .format() needs to be called on the string literal before passing it as the XPath argument.
  • Mixing Selenium and Requests: Selenium handles dynamic browser interactions (clicking buttons, navigating pages), but you're using requests.get() to fetch the page again—this won't capture the dynamic changes from your clicks (like revealed email addresses).
  • Buggy Pagination Logic: Your item loop and pagination button selection have inconsistent indexing, and the way you're building next_url doesn't align with how the site's pagination works.
  • Unnecessary Module Imports: You imported re twice, and several other modules (like BeautifulSoup, pandas, googlesearch) aren't being used at all—you can clean those up.
  • Incorrect Loop Range: You're looping from 0 to 6 for the 5 results per page, which will try to access a non-existent heading5 element.

Corrected Code

Here's the revised script that fixes all these issues and properly scrapes the email addresses:

import re
import time
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys

# Initialize the browser
driver = webdriver.Safari()
emails_notarios = []
base_url = "https://www.notariado.org/portal/"

# Navigate to the portal and search for Barcelona
driver.get(base_url)
time.sleep(1)  # Wait for page to load

# Locate the search input for location
location_input = driver.find_element(By.XPATH, '//*[@id="valor4"]')
location_input.clear()
location_input.send_keys("Barcelona")
time.sleep(1)

# Click the search button
search_button = driver.find_element(By.XPATH, '//*[@id="portlet_com_liferay_journal_content_web_portlet_JournalContentPortlet_INSTANCE_nWpBPvnHOnbm"]/div/div[2]/div/div[2]/div/button')
search_button.click()
time.sleep(2)  # Give time for results to load

# Handle pagination (assuming up to 26 pages, adjust if needed)
for page_num in range(1, 27):
    # Wait for the page results to load
    time.sleep(2)
    
    # Iterate over the 5 results per page (heading0 to heading4)
    for result_idx in range(0, 5):
        try:
            # Correctly format the XPath for the reveal button
            reveal_button = driver.find_element(By.XPATH, f'//*[@id="heading{result_idx}"]/div/h4/span')
            reveal_button.click()
            time.sleep(1)  # Wait for email to show
            
            # Extract email from the current page source
            page_source = driver.page_source
            emails = re.findall(r'\b[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}\b', page_source, re.I)
            for email in emails:
                if email not in emails_notarios:
                    emails_notarios.append(email)
                    
        except Exception as e:
            # Skip if the element doesn't exist (e.g., last page with fewer results)
            print(f"Error accessing result {result_idx} on page {page_num}: {str(e)}")
            continue
    
    # Move to the next page, if it exists
    try:
        # Correctly format the pagination button XPath
        next_page_button = driver.find_element(By.XPATH, f'//*[@id="jwpg_pagination"]/ul/li[{page_num + 1}]/a')
        next_page_button.click()
        time.sleep(2)
    except Exception as e:
        print(f"End of pagination or error moving to page {page_num + 1}: {str(e)}")
        break

# Print all collected emails
print("Collected Email Addresses:")
for email in emails_notarios:
    print(email)

# Close the browser
driver.quit()

What Changed?

  1. Fixed XPath Formatting: Used f-strings (or you could use .format()) to correctly insert the index into the XPath string before passing it to find_element().
  2. Removed Unused Modules: Cleaned up imports to only include what's necessary.
  3. Scraped Directly from Selenium's Page: Instead of using requests, we pull the page source directly from the browser after clicking the reveal button, so we get the dynamic content with emails.
  4. Simplified Pagination: The loop now iterates over page numbers directly, and we handle the next page button with correct indexing.
  5. Added Error Handling: Added try-except blocks to skip over missing elements (like if a page has fewer than 5 results) and handle pagination end gracefully.
  6. Avoided Duplicate Emails: Check if an email is already in the list before adding it.

Additional Tips

  • Adjust Sleep Times: Depending on your internet speed, you might need to increase the time.sleep() values to ensure pages/elements load fully. Alternatively, use Selenium's WebDriverWait with expected conditions for more reliable waits (instead of fixed sleeps).
  • Respect Robots.txt: Make sure you're complying with the site's robots.txt and terms of service when scraping.

内容的提问来源于stack exchange,提问作者AAlbiol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 16:42:29