Selenium爬虫嵌套标签文本提取报错:WebElement无find/getText属性
问题:Selenium网页爬取中HTML文本提取失败
我用Selenium做网页爬取,卡在HTML文本提取环节,试了多种方案都无效。
代码片段
# Extract listing links product_elements = soup.find_all('div', class_='professional-box') # find all div element product_link = [] for product_element in product_elements: #for loop iterating a list of elements content = product_element.find('div', class_='text-box type-a') if content: link = content.find('a').get('href') #get link product_link.append({'link':link}) #append "link" dict with value from link variable in product_link list #Visit all listing links and scrape the data product_info = [] for product in product_link: #for loop iterating a list of one dict [{"link"}] driver.get(product['link']) # get from the dict [{"link"}] button = driver.find_elements(By.CLASS_NAME, "text-box left-pad-25") #click all the detail buttons for btn in button: btn.click() if product: name_parent = driver.find_element(By.CLASS_NAME,'text-box') #get name name = name_parent.find('a').text facebook_parent = driver.find_element(By.CLASS_NAME,'left-facebook phone-number').get('href') facebook = facebook_parent.find_element(By.TAG_NAME,'a').get('href')
核心问题
if product:之后的代码执行失败,尤其是name_parent.find部分。还尝试过facebook_parent.find_element以及以下代码,均无效:
driver.find_element(By.XPATH, "//div[@class='text-box']/p").getText()
目标提取的HTML结构
<div class="text-box"> <p>Lingga Studio</p> </div>
各方法对应的报错信息
name_parent.find报错:
File "c:\Users\user\Desktop\Code\Archify.py", line 51, in <module> name = name_parent.find('a').text ^^^^^^^^^^^^^^^^ AttributeError: 'WebElement' object has no attribute 'find'
facebook_parent.find_element报错:
return self._execute(Command.FIND_CHILD_ELEMENT, {"using": by, "value": value})["value"]
getText()报错:
AttributeError: 'WebElement' object has no attribute 'getText'
解决方案
文本提取错误修正
find方法误用:Selenium的WebElement对象没有find方法,需使用find_element()/find_elements();同时你要提取的是<p>标签文本,不是<a>,修正代码:
# 提取目标文本 name_parent = driver.find_element(By.CLASS_NAME, 'text-box') name = name_parent.find_element(By.TAG_NAME, 'p').text
或直接用XPATH一步定位:
name = driver.find_element(By.XPATH, "//div[@class='text-box']/p").text
getText()方法错误:Python版Selenium获取文本用.text属性,getText()是Java版用法。
Facebook链接提取错误修正
By.CLASS_NAME不支持多类名匹配(如left-facebook phone-number),需改用CSS_SELECTOR或XPATH;同时Selenium获取元素属性要用get_attribute(),而非BeautifulSoup的.get():
# 用CSS_SELECTOR定位多类名元素并获取链接 facebook_parent = driver.find_element(By.CSS_SELECTOR, '.left-facebook.phone-number') facebook = facebook_parent.find_element(By.TAG_NAME, 'a').get_attribute('href')
或用XPATH直接定位:
facebook = driver.find_element(By.XPATH, "//div[@class='left-facebook phone-number']/a").get_attribute('href')
其他细节修正
- 按钮定位时,
By.CLASS_NAME同样不支持多类名,需改为:
button = driver.find_elements(By.CSS_SELECTOR, ".text-box.left-pad-25")
- 点击按钮后建议添加等待,避免页面未加载完成就提取元素:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 等待目标元素加载完成,最长等待10秒 WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CLASS_NAME, 'text-box')))
内容的提问来源于stack exchange,提问作者Jason
相关产品推荐
相关产品推荐

