使用Python+Selenium(RPA Framework)爬取纽约时报报错求助
解决RPA Framework Selenium爬取纽约时报时的InvalidArgumentException错误
问题根源
你遇到的错误是因为RPA.Browser.Selenium封装的find_elements/find_element方法要求显式指定定位策略(如xpath、css selector),而原生Selenium允许默认使用XPath。你直接传入XPath表达式但未声明策略,导致框架无法识别定位方式。
修复步骤及完整代码
以下是修复后的代码,同时补充了日期和描述的提取逻辑,以及相对路径定位、文本获取的细节:
from RPA.Browser.Selenium import Selenium # Search term search_term = "climate change" # Open the NY Times search page and search for the term browser = Selenium() browser.open_available_browser(f"https://www.nytimes.com/search?query={search_term}") # 显式指定xpath策略,查找所有文章节点 articles = browser.find_elements("xpath", "//ol[@data-testid='search-results']/li") # Extract title, date, and description for each article and add to the list for article in articles: try: # 用相对路径xpath(开头加.)在当前article节点内查找,避免全局查找 title = article.find_element("xpath", ".//h4[@class='css-2fgx4k']").text date = article.find_element("xpath", ".//time").text description = article.find_element("xpath", ".//p[@class='css-16nhkrn']").text print(f"标题: {title}\n日期: {date}\n描述: {description}\n---") except Exception as e: print(f"提取文章信息失败: {str(e)}") continue # Close the browser window browser.close_all_browsers()
关键修改点
- 所有
find_elements/find_element调用,第一个参数必须传入定位策略(如"xpath"),第二个参数才是表达式 - 子元素查找时使用相对路径XPath(开头加
.),确保只在当前父元素范围内搜索,避免重复获取第一个元素 - 调用
.text属性获取元素的文本内容,而不是直接打印元素对象 - 添加
try-except块处理个别文章元素结构不一致的情况,避免程序直接崩溃 - 注意:纽约时报的CSS类名可能会动态更新,若运行时仍找不到元素,建议手动检查页面的最新元素选择器
内容的提问来源于stack exchange,提问作者pedros
相关产品推荐
相关产品推荐

