You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium爬取Pressreader平台Libero报纸文章报错求助

问题解决:Pressreader爬取Selenium多阶段报错处理

一、核心问题梳理

你需要爬取Pressreader网站上《Libero》报纸指定时间段、关键词的新闻并保存为CSV,尝试R包失败后改用Python Selenium,遇到以下报错:

  1. TypeError: WebDriver.__init__() got an unexpected keyword argument 'chrome_options' + NameError: name 'driver' is not defined
  2. IndentationError: unindent does not match any outer indentation level
  3. AttributeError: module 'selenium.webdriver.chrome.service' has no attribute 'path'

已安装最新版Selenium,以下是针对性解决方案:

二、分报错类型解决

1. TypeError与NameError

报错原因

  • 新版Selenium(4.x+)已弃用chrome_options参数,改用options
  • driver未定义是因为初始化失败导致后续代码无法调用

修复代码示例

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

# 初始化Chrome选项
chrome_options = Options()
chrome_options.add_argument("--headless=new")  # 无头模式,按需开启
chrome_options.add_argument("--disable-blink-features=AutomationControlled")

# 初始化驱动(新版Selenium无需手动指定ChromeDriver路径,会自动匹配)
driver = webdriver.Chrome(options=chrome_options)

# 后续爬取逻辑示例
driver.get("https://www.pressreader.com/")

2. IndentationError

报错原因

Python对缩进严格要求,代码中混合了空格和制表符,或者缩进层级不匹配

修复方式

  • 统一使用4个空格缩进,不要混用制表符
  • 检查所有代码块(如for/while循环、if判断、函数定义)的缩进是否一致
  • 可以用编辑器的"显示空格/制表符"功能排查,比如VS Code开启"Editor: Render Whitespace"

3. AttributeError: module 'selenium.webdriver.chrome.service' has no attribute 'path'

报错原因

新版Selenium中,Service类的初始化参数已变更,不再直接使用path属性手动指定驱动路径

修复方案

新版Selenium(4.6+)支持自动管理ChromeDriver,无需手动下载配置,直接简化代码:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
# 添加你的选项配置
options.add_argument("--start-maximized")

# 自动匹配ChromeDriver,无需手动指定路径
driver = webdriver.Chrome(options=options)

# 如果必须手动指定驱动路径(不推荐),用以下方式:
from selenium.webdriver.chrome.service import Service

service = Service(executable_path="/path/to/chromedriver")
driver = webdriver.Chrome(service=service, options=options)

三、爬取Pressreader的额外注意事项

  • Pressreader有反爬机制,建议添加随机延迟、更换User-Agent:
    import time
    from fake_useragent import UserAgent
    
    ua = UserAgent()
    options.add_argument(f"user-agent={ua.random}")
    
    # 操作后添加延迟
    time.sleep(2)
    
  • 处理登录(如果需要):找到用户名/密码输入框和登录按钮,用driver.find_element()定位后操作
  • 搜索关键词和时间段:模拟输入关键词,选择日期范围,点击搜索
  • 提取文章内容:定位文章标题、正文、发布时间等元素,存入列表后用pandas保存为CSV:
    import pandas as pd
    
    articles = []
    # 假设已提取到标题、正文、时间
    articles.append({"title": title_text, "content": content_text, "date": date_text})
    
    df = pd.DataFrame(articles)
    df.to_csv("libero_news.csv", index=False, encoding="utf-8-sig")
    

内容的提问来源于stack exchange,提问作者Cate16

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:31:16