You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用lxml XPath结合BeautifulSoup获取网页标题返回None问题

问题:使用BeautifulSoup+lxml XPath获取维基百科标题返回None

尝试通过以下代码获取维基百科页面标题,运行后输出None,但预期结果应为Wikipedia:About,改用完整XPath也无法解决:

from bs4 import BeautifulSoup
from lxml import etree
import requests

xpath_url = "https://en.wikipedia.org/wiki/Wikipedia:About"
xpath_headers = ({'User-Agent':
'Safari/537.36',\
'Accept-Language': 'en-US, en;q=0.5'})
xpath_wpage = requests.get (xpath_url, headers = xpath_headers)
xpath_soup = BeautifulSoup (xpath_wpage.content, "html.parser")
dom = etree.HTML (str(xpath_soup))
print (dom.xpath ('//*[@id="firstHeading"]')[0].text)

问题原因

核心问题出在BeautifulSoup转字符串再传给lxml解析的流程:
维基百科的firstHeading元素内部包含子<span>标签,标题文本实际存储在子节点中。而lxml的.text属性仅返回当前元素的直接文本内容,当元素下有子节点且自身无直接文本时,就会返回None。另外,BeautifulSoup转字符串时会调整HTML结构,进一步导致解析后的节点结构和原页面有差异。

解决方案

方案1:直接用lxml请求解析(推荐)

无需BeautifulSoup中转,直接用lxml处理响应内容,代码更简洁高效:

from lxml import etree
import requests

url = "https://en.wikipedia.org/wiki/Wikipedia:About"
headers = {
    'User-Agent': 'Safari/537.36',
    'Accept-Language': 'en-US, en;q=0.5'
}
response = requests.get(url, headers=headers)
dom = etree.HTML(response.content)
# 获取firstHeading下的所有文本并提取有效内容
print(dom.xpath('//*[@id="firstHeading"]//text()')[0].strip())

方案2:仅用BeautifulSoup获取标题

如果不需要lxml的XPath功能,直接用BeautifulSoup的API即可轻松获取:

from bs4 import BeautifulSoup
import requests

url = "https://en.wikipedia.org/wiki/Wikipedia:About"
headers = {
    'User-Agent': 'Safari/537.36',
    'Accept-Language': 'en-US, en;q=0.5'
}
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, "html.parser")
print(soup.find(id="firstHeading").get_text(strip=True))

方案3:修复原代码(不推荐)

若坚持要保留原流程,可修改XPath来获取所有后代文本:

from bs4 import BeautifulSoup
from lxml import etree
import requests

xpath_url = "https://en.wikipedia.org/wiki/Wikipedia:About"
xpath_headers = {'User-Agent': 'Safari/537.36', 'Accept-Language': 'en-US, en;q=0.5'}
xpath_wpage = requests.get(xpath_url, headers=xpath_headers)
xpath_soup = BeautifulSoup(xpath_wpage.content, "html.parser")
dom = etree.HTML(str(xpath_soup))
# 提取所有后代文本并拼接
title_text = ''.join(dom.xpath('//*[@id="firstHeading"]//text()')).strip()
print(title_text)

内容的提问来源于stack exchange,提问作者neiphu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 08:45:36