You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium无法获取页面标题问题求助

问题描述

Source Code

尝试从页面源码中获取标题,编写了多条XPath表达式,但遇到如下错误:

错误信息: DeprecationWarning: find_elements_by_* commands are deprecated. Please use find_elements() instead

try:
    #title = browser.find_element_by_xpath('/html/body/div[1]/main/div[2]/div[1]/div[2]/div/div[1]/div[1]/h1/img').text.strip()
    #title = browser.find_element(By.XPATH, "//h1[@class='cp-typography--hero-title-logo']/img").text
    #title = browser.find_elements(By.CSS_SELECTOR, "content-page__programmatic-header").text.strip()
    print("Title: "+title)
except:
    print("Missing Title")

解决方法

1. 处理弃用警告

你第一条注释的find_element_by_xpath是旧版API,确实已被弃用,改用find_element(By.XPATH, 表达式)的写法是正确的,这部分你已经在第二条尝试里做对了。

2. 修复标题获取失败的核心问题

从截图源码和你的代码来看,问题出在这几点:

  • <img>元素没有内部文本,用.text拿不到内容,它的标题通常存在alt或title属性里,得用get_attribute()获取。
  • 第三条的CSS选择器写错了:content-page__programmatic-header是类名,需要加.前缀;且find_elements返回元素列表,不能直接调用.text,取单个元素要用find_element。

修改后的可用代码:

from selenium.webdriver.common.by import By

try:
    # 获取img的alt属性作为标题(如果标题存在于title属性就换成"title")
    img_element = browser.find_element(By.XPATH, "//h1[@class='cp-typography--hero-title-logo']/img")
    title = img_element.get_attribute("alt").strip()

    # 若要获取header区域的文本(如果有),用下面这行替换上面的逻辑
    # title = browser.find_element(By.CSS_SELECTOR, ".content-page__programmatic-header").text.strip()
    
    print("Title: "+title)
except Exception as e:
    print(f"Missing Title, error: {str(e)}")

关键注意事项

  • 避免使用绝对XPath(第一条那种/html/body/...),页面结构变动就会失效,用基于类名、属性的相对XPath更稳定。
  • 针对img、input这类无内部文本的元素,必须用get_attribute()获取属性值,不能用.text。
  • find_elements返回元素列表,需遍历处理;单个元素查询用find_element。

内容的提问来源于stack exchange,提问作者karthi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 11:30:57