You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取嵌套多<strong>标签的<h2>标签文本?调用.string返回None

解决BeautifulSoup提取h2标签文本返回None的问题

问题原因

你的<h2>标签包含多个嵌套子节点(多层<strong>标签和分散的文本节点),而BeautifulSoup的string属性仅在标签内部只有单个文本节点时才会返回有效内容,否则返回None。

解决方案

你可以用以下两种方法提取目标文本:

方法1:使用get_text()(推荐)

get_text()会自动拼接标签下所有子节点的文本内容,还能通过参数处理空格:

firstHeader = mclarenHTML.find_all(re.compile('^h[2]'))[0]
# 提取文本、去除首尾空格、替换非断空格(对应原HTML的&nbsp;)
target_text = firstHeader.get_text(strip=True).replace('\xa0', '')
print(target_text)

执行后会输出:1950-1953:Formula 1 begins: the super-charger years

方法2:使用stripped_strings迭代器

如果需要更精细地处理每个文本片段,可以遍历stripped_strings(自动去除每个文本片段的首尾空格):

firstHeader = mclarenHTML.find_all(re.compile('^h[2]'))[0]
text_fragments = [frag.replace('\xa0', '') for frag in firstHeader.stripped_strings]
target_text = ''.join(text_fragments)
print(target_text)

补充说明

原HTML中的&nbsp;在Python中对应转义字符\xa0,所以需要用replace('\xa0', '')去除这个多余的空格,才能得到你想要的连续文本格式。

内容的提问来源于stack exchange,提问作者LearningToCode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 08:50:26