You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+Selenium(RPA Framework)爬取纽约时报报错求助

解决RPA Framework Selenium爬取纽约时报时的InvalidArgumentException错误

问题根源

你遇到的错误是因为RPA.Browser.Selenium封装的find_elements/find_element方法要求显式指定定位策略(如xpath、css selector),而原生Selenium允许默认使用XPath。你直接传入XPath表达式但未声明策略,导致框架无法识别定位方式。

修复步骤及完整代码

以下是修复后的代码,同时补充了日期和描述的提取逻辑,以及相对路径定位、文本获取的细节:

from RPA.Browser.Selenium import Selenium

# Search term
search_term = "climate change"

# Open the NY Times search page and search for the term
browser = Selenium()
browser.open_available_browser(f"https://www.nytimes.com/search?query={search_term}")

# 显式指定xpath策略,查找所有文章节点
articles = browser.find_elements("xpath", "//ol[@data-testid='search-results']/li")

# Extract title, date, and description for each article and add to the list
for article in articles:
    try:
        # 用相对路径xpath(开头加.)在当前article节点内查找,避免全局查找
        title = article.find_element("xpath", ".//h4[@class='css-2fgx4k']").text
        date = article.find_element("xpath", ".//time").text
        description = article.find_element("xpath", ".//p[@class='css-16nhkrn']").text
        print(f"标题: {title}\n日期: {date}\n描述: {description}\n---")
    except Exception as e:
        print(f"提取文章信息失败: {str(e)}")
        continue

# Close the browser window
browser.close_all_browsers()

关键修改点

  • 所有find_elements/find_element调用,第一个参数必须传入定位策略(如"xpath"),第二个参数才是表达式
  • 子元素查找时使用相对路径XPath(开头加.),确保只在当前父元素范围内搜索,避免重复获取第一个元素
  • 调用.text属性获取元素的文本内容,而不是直接打印元素对象
  • 添加try-except块处理个别文章元素结构不一致的情况,避免程序直接崩溃
  • 注意:纽约时报的CSS类名可能会动态更新,若运行时仍找不到元素,建议手动检查页面的最新元素选择器

内容的提问来源于stack exchange,提问作者pedros

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 10:12:46