You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium+Python多网页爬虫报错:WebDriver无find_element_by_tag_name属性

解决Selenium中'WebDriver' object has no attribute 'find_element_by_tag_name'错误

错误原因

你遇到的错误是因为Selenium 4.0及以上版本已移除find_element_by_*/find_elements_by_*这类旧定位API,需要使用新的find_element()和find_elements()方法,配合By类指定定位策略。

修复后的完整代码

以下是修正后的可运行代码,同时修复了原代码中的缩进问题和URL空格问题:

from selenium import webdriver
import pandas as pd
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By

# 待爬取的URL列表(修复了第四个URL前的空格)
urls = [
    "https://www.monolithai.com/blog/4-ways-ai-is-changing-the-packaging-industry",
    "https://mitsubishisolutions.com/the-role-of-artificial-intelligence-in-smart-packaging-lines",
    "https://thedatascientist.com/how-artificial-intelligence-is-revolutionizing-the-packaging-industry/",
    "https://packagingeurope.com/comment/ai-and-the-future-of-packaging/9665.article"
]

# 初始化WebDriver
driver = webdriver.Chrome()  # 根据你的浏览器选择对应WebDriver
wait = WebDriverWait(driver, 10)

# 存储爬取数据的空列表
all_text = []
all_images = []
all_links = []

# 遍历每个URL爬取数据
for url in urls:
    driver.get(url)
    
    # 等待页面body加载完成
    wait.until(EC.presence_of_element_located((By.TAG_NAME, 'body')))
    
    # 爬取页面文本
    page_text = driver.find_element(By.TAG_NAME, 'body').text
    all_text.append(page_text)
    
    # 爬取图片URL(过滤无效的空链接)
    images = driver.find_elements(By.TAG_NAME, 'img')
    image_urls = [img.get_attribute('src') for img in images if img.get_attribute('src')]
    all_images.append(image_urls)
    
    # 爬取页面链接(过滤无效的空链接)
    links = driver.find_elements(By.TAG_NAME, 'a')
    link_urls = [link.get_attribute('href') for link in links if link.get_attribute('href')]
    all_links.append(link_urls)

# 关闭浏览器释放资源
driver.quit()

# 构建DataFrame并保存到Excel
data = {
    'URL': urls,
    'Text': all_text,
    'Images': all_images,
    'Links': all_links
}
df = pd.DataFrame(data)
df.to_excel('scraped_data.xlsx', index=False)

额外优化说明

  • 修复了原URL列表中第四个URL前的空格,避免请求错误
  • 爬取图片和链接时增加非空判断,过滤无效的src/href值
  • 启用driver.quit(),确保爬取完成后关闭浏览器释放资源
  • 调整代码缩进,确保所有爬取逻辑都在for循环内执行

内容的提问来源于stack exchange,提问作者Pradeep Birare

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 12:15:54