Selenium动态网页爬取遇NoSuchElementException错误求助
问题:Selenium爬取动态网站时触发NoSuchElementException错误
我尝试用Selenium编写代码爬取动态网站https://www.archify.com/id/professionals的多列表数据,但卡在数据提取环节。执行代码时触发NoSuchElementException错误,提示无法定位name为href的元素。我已经把product_elements放入循环中,但问题仍未解决。以下是我的代码及报错回溯信息:
原代码
from selenium import webdriver #to enable Wait for Page Loading from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC #to enable scrolling from selenium.webdriver.common.keys import Keys from bs4 import BeautifulSoup import requests #Initialize browser options = webdriver.ChromeOptions() #options.add_argument('--headless') # Not headless because there's an error with Hotjar driver = webdriver.Chrome(options=options) url = "https://www.archify.com/id/professionals" driver.get(url) #Click Load More Button l = driver.find_element("xpath", "//button[text()='Load More']") l.click() # Scroll until bottom of the page driver.execute_script("window.scrollTo(0, document.body.scrollHeight)") #Tell browser to wait for for loading (for elements to load / for 60 seconds) wait = WebDriverWait(driver, 60) # Extract product details product_elements = driver.find_elements(By.CLASS_NAME, 'professional-box') product_data = [] for product_element in product_elements: link = product_element.find_element(By.NAME, 'href').text title = product_element.find_element(By.NAME, '.title').text subtitle = product_element.find_element(By.NAME,"subtitle").text product_data.append({'title': title, 'subtitle': subtitle, 'link': link}) # Print extracted data for product in product_data: print(f"Title: {product['title']}, Subtitle: {product['subtitle']}, link: {product['link']}") driver.quit
报错回溯信息
PS C:\Users\user\Desktop\Code> & C:/Users/user/AppData/Local/Programs/Python/Python311/python.exe c:/Users/user/Desktop/Code/Scrape_Selenium.py DevTools listening on ws://127.0.0.1:51089/devtools/browser/581a68b5-c26f-42fd-93e0-f4970c14fd1b Traceback (most recent call last): File "c:\Users\user\Desktop\Code\Scrape_Selenium.py", line 43, in <module> link = product_element.find_element(By.NAME, 'href').text ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\user\AppData\Local\Programs\Python\Python311\Lib\site-packages\selenium\webdriver\remote\webelement.py", line 417, in find_element return self._execute(Command.FIND_CHILD_ELEMENT, {"using": by, "value": value})["value"] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\user\AppData\Local\Programs\Python\Python311\Lib\site-packages\selenium\webdriver\remote\webelement.py", line 395, in _execute return self._parent.execute(command, params) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\user\AppData\Local\Programs\Python\Python311\Lib\site-packages\selenium\webdriver\remote\webdriver.py", line 347, in execute self.error_handler.check_response(response) File "C:\Users\user\AppData\Local\Programs\Python\Python311\Lib\site-packages\selenium\webdriver\remote\errorhandler.py", line 229, in check_response raise exception_class(message, screen, stacktrace) selenium.common.exceptions.NoSuchElementException: Message: no such element: Unable to locate element: {"method":"css selector","selector":"[name="href"]"} (Session info: chrome=122.0.6261.112); For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#no-such-element-exception Stacktrace: GetHandleVerifier [0x00007FF6B60AAD32+56930] (No symbol) [0x00007FF6B601F632] (No symbol) [0x00007FF6B5ED42E5] (No symbol) [0x00007FF6B5F198ED] (No symbol) [0x00007FF6B5F19A2C] (No symbol) [0x00007FF6B5F0F13C] (No symbol) [0x00007FF6B5F3BCDF] (No symbol) [0x00007FF6B5F0F09A] (No symbol) [0x00007FF6B5F3BEB0] (No symbol) [0x00007FF6B5F581E2] (No symbol) [0x00007FF6B5F3BA43] (No symbol) [0x00007FF6B5F0D438] (No symbol) [0x00007FF6B5F0E4D1] GetHandleVerifier [0x00007FF6B6426ABD+3709933] GetHandleVerifier [0x00007FF6B647FFFD+4075821] GetHandleVerifier [0x00007FF6B647818F+4043455] GetHandleVerifier [0x00007FF6B6149766+706710] (No symbol) [0x00007FF6B602B90F] (No symbol) [0x00007FF6B6026AF4] (No symbol) [0x00007FF6B6026C4C] (No symbol) [0x00007FF6B6016904] BaseThreadInitThunk [0x00007FFEA0697344+20] RtlUserThreadStart [0x00007FFEA08A26B1+33]
解决方案
错误原因
- 元素定位逻辑错误:页面中不存在
name属性为href、.title、subtitle的元素,混淆了元素属性与选择器的用法。 - 缺少动态加载等待:点击Load More后直接滚动、提取元素,新内容可能未加载完成,导致元素无法定位。
- 链接获取方式错误:
href是a标签的属性,不是文本内容,不能用.text获取,需用get_attribute('href')。 - 语法错误:
driver.quit缺少括号,无法正确关闭浏览器。
修正后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 初始化浏览器 options = webdriver.ChromeOptions() # options.add_argument('--headless') # 无头模式存在Hotjar错误,暂时关闭 driver = webdriver.Chrome(options=options) url = "https://www.archify.com/id/professionals" driver.get(url) # 等待Load More按钮加载完成并点击 wait = WebDriverWait(driver, 60) load_more_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[text()='Load More']"))) load_more_btn.click() # 等待新内容加载后滚动到底部 wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'professional-box'))) driver.execute_script("window.scrollTo(0, document.body.scrollHeight)") # 等待所有目标元素加载完成 wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'professional-box'))) # 提取数据 product_elements = driver.find_elements(By.CLASS_NAME, 'professional-box') product_data = [] for product_element in product_elements: # 定位a标签获取链接 link_tag = product_element.find_element(By.TAG_NAME, 'a') link = link_tag.get_attribute('href') # 定位标题和副标题 title = product_element.find_element(By.CSS_SELECTOR, '.title').text.strip() subtitle = product_element.find_element(By.CLASS_NAME, 'subtitle').text.strip() product_data.append({'title': title, 'subtitle': subtitle, 'link': link}) # 打印提取的数据 for product in product_data: print(f"标题: {product['title']}, 副标题: {product['subtitle']}, 链接: {product['link']}") # 关闭浏览器 driver.quit()
关键修改说明
- 完善等待逻辑:在点击按钮、滚动页面后添加显式等待,确保元素完全加载后再操作,避免动态加载导致的元素未找到问题。
- 修正元素定位:
- 链接:通过a标签的
get_attribute('href')获取实际地址。 - 标题使用CSS选择器
.title,副标题使用CLASS_NAMEsubtitle,匹配页面实际元素结构。
- 链接:通过a标签的
- 优化数据格式:用
strip()去除文本两端空格,提升数据整洁度。 - 修复语法错误:补充
driver.quit()的括号,确保浏览器正常关闭。
内容的提问来源于stack exchange,提问作者Jason
相关产品推荐
相关产品推荐

