使用lxml XPath结合BeautifulSoup获取网页标题返回None问题
问题:使用BeautifulSoup+lxml XPath获取维基百科标题返回None
尝试通过以下代码获取维基百科页面标题,运行后输出None,但预期结果应为Wikipedia:About,改用完整XPath也无法解决:
from bs4 import BeautifulSoup from lxml import etree import requests xpath_url = "https://en.wikipedia.org/wiki/Wikipedia:About" xpath_headers = ({'User-Agent': 'Safari/537.36',\ 'Accept-Language': 'en-US, en;q=0.5'}) xpath_wpage = requests.get (xpath_url, headers = xpath_headers) xpath_soup = BeautifulSoup (xpath_wpage.content, "html.parser") dom = etree.HTML (str(xpath_soup)) print (dom.xpath ('//*[@id="firstHeading"]')[0].text)
问题原因
核心问题出在BeautifulSoup转字符串再传给lxml解析的流程:
维基百科的firstHeading元素内部包含子<span>标签,标题文本实际存储在子节点中。而lxml的.text属性仅返回当前元素的直接文本内容,当元素下有子节点且自身无直接文本时,就会返回None。另外,BeautifulSoup转字符串时会调整HTML结构,进一步导致解析后的节点结构和原页面有差异。
解决方案
方案1:直接用lxml请求解析(推荐)
无需BeautifulSoup中转,直接用lxml处理响应内容,代码更简洁高效:
from lxml import etree import requests url = "https://en.wikipedia.org/wiki/Wikipedia:About" headers = { 'User-Agent': 'Safari/537.36', 'Accept-Language': 'en-US, en;q=0.5' } response = requests.get(url, headers=headers) dom = etree.HTML(response.content) # 获取firstHeading下的所有文本并提取有效内容 print(dom.xpath('//*[@id="firstHeading"]//text()')[0].strip())
方案2:仅用BeautifulSoup获取标题
如果不需要lxml的XPath功能,直接用BeautifulSoup的API即可轻松获取:
from bs4 import BeautifulSoup import requests url = "https://en.wikipedia.org/wiki/Wikipedia:About" headers = { 'User-Agent': 'Safari/537.36', 'Accept-Language': 'en-US, en;q=0.5' } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, "html.parser") print(soup.find(id="firstHeading").get_text(strip=True))
方案3:修复原代码(不推荐)
若坚持要保留原流程,可修改XPath来获取所有后代文本:
from bs4 import BeautifulSoup from lxml import etree import requests xpath_url = "https://en.wikipedia.org/wiki/Wikipedia:About" xpath_headers = {'User-Agent': 'Safari/537.36', 'Accept-Language': 'en-US, en;q=0.5'} xpath_wpage = requests.get(xpath_url, headers=xpath_headers) xpath_soup = BeautifulSoup(xpath_wpage.content, "html.parser") dom = etree.HTML(str(xpath_soup)) # 提取所有后代文本并拼接 title_text = ''.join(dom.xpath('//*[@id="firstHeading"]//text()')).strip() print(title_text)
内容的提问来源于stack exchange,提问作者neiphu
相关产品推荐
相关产品推荐

