Python Selenium无法获取页面标题问题求助
问题描述

尝试从页面源码中获取标题,编写了多条XPath表达式,但遇到如下错误:
错误信息: DeprecationWarning: find_elements_by_* commands are deprecated. Please use find_elements() instead
try: #title = browser.find_element_by_xpath('/html/body/div[1]/main/div[2]/div[1]/div[2]/div/div[1]/div[1]/h1/img').text.strip() #title = browser.find_element(By.XPATH, "//h1[@class='cp-typography--hero-title-logo']/img").text #title = browser.find_elements(By.CSS_SELECTOR, "content-page__programmatic-header").text.strip() print("Title: "+title) except: print("Missing Title")
解决方法
1. 处理弃用警告
你第一条注释的find_element_by_xpath是旧版API,确实已被弃用,改用find_element(By.XPATH, 表达式)的写法是正确的,这部分你已经在第二条尝试里做对了。
2. 修复标题获取失败的核心问题
从截图源码和你的代码来看,问题出在这几点:
<img>元素没有内部文本,用.text拿不到内容,它的标题通常存在alt或title属性里,得用get_attribute()获取。- 第三条的CSS选择器写错了:
content-page__programmatic-header是类名,需要加.前缀;且find_elements返回元素列表,不能直接调用.text,取单个元素要用find_element。
修改后的可用代码:
from selenium.webdriver.common.by import By try: # 获取img的alt属性作为标题(如果标题存在于title属性就换成"title") img_element = browser.find_element(By.XPATH, "//h1[@class='cp-typography--hero-title-logo']/img") title = img_element.get_attribute("alt").strip() # 若要获取header区域的文本(如果有),用下面这行替换上面的逻辑 # title = browser.find_element(By.CSS_SELECTOR, ".content-page__programmatic-header").text.strip() print("Title: "+title) except Exception as e: print(f"Missing Title, error: {str(e)}")
关键注意事项
- 避免使用绝对XPath(第一条那种
/html/body/...),页面结构变动就会失效,用基于类名、属性的相对XPath更稳定。 - 针对
img、input这类无内部文本的元素,必须用get_attribute()获取属性值,不能用.text。 find_elements返回元素列表,需遍历处理;单个元素查询用find_element。
内容的提问来源于stack exchange,提问作者karthi
相关产品推荐
相关产品推荐

