You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium动态网页爬取遇NoSuchElementException错误求助

问题:Selenium爬取动态网站时触发NoSuchElementException错误

我尝试用Selenium编写代码爬取动态网站https://www.archify.com/id/professionals的多列表数据,但卡在数据提取环节。执行代码时触发NoSuchElementException错误,提示无法定位name为href的元素。我已经把product_elements放入循环中,但问题仍未解决。以下是我的代码及报错回溯信息:

原代码

from selenium import webdriver
#to enable Wait for Page Loading
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
#to enable scrolling
from selenium.webdriver.common.keys import Keys
from bs4 import BeautifulSoup
import requests

#Initialize browser
options = webdriver.ChromeOptions()
#options.add_argument('--headless')  # Not headless because there's an error with Hotjar
driver = webdriver.Chrome(options=options)

url = "https://www.archify.com/id/professionals"
driver.get(url)

#Click Load More Button
l = driver.find_element("xpath", "//button[text()='Load More']")
l.click()

# Scroll until bottom of the page
driver.execute_script("window.scrollTo(0, document.body.scrollHeight)")

#Tell browser to wait for for loading (for elements to load / for 60 seconds)
wait = WebDriverWait(driver, 60) 

# Extract product details
product_elements = driver.find_elements(By.CLASS_NAME, 'professional-box')
product_data = []
for product_element in product_elements:
    link = product_element.find_element(By.NAME, 'href').text
    title = product_element.find_element(By.NAME, '.title').text
    subtitle = product_element.find_element(By.NAME,"subtitle").text
    product_data.append({'title': title, 'subtitle': subtitle, 'link': link})
# Print extracted data
for product in product_data:
    print(f"Title: {product['title']}, Subtitle: {product['subtitle']}, link: {product['link']}")
    
driver.quit

报错回溯信息

PS C:\Users\user\Desktop\Code> & C:/Users/user/AppData/Local/Programs/Python/Python311/python.exe c:/Users/user/Desktop/Code/Scrape_Selenium.py

DevTools listening on ws://127.0.0.1:51089/devtools/browser/581a68b5-c26f-42fd-93e0-f4970c14fd1b
Traceback (most recent call last):
  File "c:\Users\user\Desktop\Code\Scrape_Selenium.py", line 43, in <module>
    link = product_element.find_element(By.NAME, 'href').text
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\user\AppData\Local\Programs\Python\Python311\Lib\site-packages\selenium\webdriver\remote\webelement.py", line 417, in find_element
    return self._execute(Command.FIND_CHILD_ELEMENT, {"using": by, "value": value})["value"]
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\user\AppData\Local\Programs\Python\Python311\Lib\site-packages\selenium\webdriver\remote\webelement.py", line 395, in _execute
    return self._parent.execute(command, params)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\user\AppData\Local\Programs\Python\Python311\Lib\site-packages\selenium\webdriver\remote\webdriver.py", line 347, in execute
    self.error_handler.check_response(response)
  File "C:\Users\user\AppData\Local\Programs\Python\Python311\Lib\site-packages\selenium\webdriver\remote\errorhandler.py", line 229, in check_response
    raise exception_class(message, screen, stacktrace)
selenium.common.exceptions.NoSuchElementException: Message: no such element: Unable to locate element: {"method":"css selector","selector":"[name="href"]"}
  (Session info: chrome=122.0.6261.112); For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#no-such-element-exception
Stacktrace:
        GetHandleVerifier [0x00007FF6B60AAD32+56930]
        (No symbol) [0x00007FF6B601F632]
        (No symbol) [0x00007FF6B5ED42E5]
        (No symbol) [0x00007FF6B5F198ED]
        (No symbol) [0x00007FF6B5F19A2C]
        (No symbol) [0x00007FF6B5F0F13C]
        (No symbol) [0x00007FF6B5F3BCDF]
        (No symbol) [0x00007FF6B5F0F09A]
        (No symbol) [0x00007FF6B5F3BEB0]
        (No symbol) [0x00007FF6B5F581E2]
        (No symbol) [0x00007FF6B5F3BA43]
        (No symbol) [0x00007FF6B5F0D438]
        (No symbol) [0x00007FF6B5F0E4D1]
        GetHandleVerifier [0x00007FF6B6426ABD+3709933]
        GetHandleVerifier [0x00007FF6B647FFFD+4075821]
        GetHandleVerifier [0x00007FF6B647818F+4043455]
        GetHandleVerifier [0x00007FF6B6149766+706710]
        (No symbol) [0x00007FF6B602B90F]
        (No symbol) [0x00007FF6B6026AF4]
        (No symbol) [0x00007FF6B6026C4C]
        (No symbol) [0x00007FF6B6016904]
        BaseThreadInitThunk [0x00007FFEA0697344+20]
        RtlUserThreadStart [0x00007FFEA08A26B1+33]

解决方案

错误原因

  1. 元素定位逻辑错误:页面中不存在name属性为href、.title、subtitle的元素,混淆了元素属性与选择器的用法。
  2. 缺少动态加载等待:点击Load More后直接滚动、提取元素,新内容可能未加载完成,导致元素无法定位。
  3. 链接获取方式错误:href是a标签的属性,不是文本内容,不能用.text获取,需用get_attribute('href')。
  4. 语法错误:driver.quit缺少括号,无法正确关闭浏览器。

修正后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 初始化浏览器
options = webdriver.ChromeOptions()
# options.add_argument('--headless')  # 无头模式存在Hotjar错误,暂时关闭
driver = webdriver.Chrome(options=options)

url = "https://www.archify.com/id/professionals"
driver.get(url)

# 等待Load More按钮加载完成并点击
wait = WebDriverWait(driver, 60)
load_more_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[text()='Load More']")))
load_more_btn.click()

# 等待新内容加载后滚动到底部
wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'professional-box')))
driver.execute_script("window.scrollTo(0, document.body.scrollHeight)")
# 等待所有目标元素加载完成
wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'professional-box')))

# 提取数据
product_elements = driver.find_elements(By.CLASS_NAME, 'professional-box')
product_data = []
for product_element in product_elements:
    # 定位a标签获取链接
    link_tag = product_element.find_element(By.TAG_NAME, 'a')
    link = link_tag.get_attribute('href')
    # 定位标题和副标题
    title = product_element.find_element(By.CSS_SELECTOR, '.title').text.strip()
    subtitle = product_element.find_element(By.CLASS_NAME, 'subtitle').text.strip()
    product_data.append({'title': title, 'subtitle': subtitle, 'link': link})

# 打印提取的数据
for product in product_data:
    print(f"标题: {product['title']}, 副标题: {product['subtitle']}, 链接: {product['link']}")

# 关闭浏览器
driver.quit()

关键修改说明

  • 完善等待逻辑:在点击按钮、滚动页面后添加显式等待,确保元素完全加载后再操作,避免动态加载导致的元素未找到问题。
  • 修正元素定位:
    • 链接:通过a标签的get_attribute('href')获取实际地址。
    • 标题使用CSS选择器.title,副标题使用CLASS_NAMEsubtitle,匹配页面实际元素结构。
  • 优化数据格式:用strip()去除文本两端空格,提升数据整洁度。
  • 修复语法错误:补充driver.quit()的括号,确保浏览器正常关闭。

内容的提问来源于stack exchange,提问作者Jason

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 00:54:55