You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从HTML片段中提取'Terence Crawford'并排除span元素?

问题:提取HTML中排除子元素的目标文本

我无法从以下HTML片段中提取名称Terence Crawford,难点在于需要排除同一父元素内的<span>元素:

<td colspan="3" style="position:relative;" class="defaultTitleAlign">
<h1 style="display:inline-block;margin-right:5px;line-height:30px;">
                        <span style="font-weight:bold;"><i class="fas fa-crown" style="color:#f6b501 !important;"></i></span>
                    "Terence Crawford"
    </h1>
<div style="width:100%;position:relative;margin-top:5px;">
</div>
</td>

我尝试通过指定class属性defaultTitleAlign和style属性定位<h1>元素提取文本,但仅返回空白换行符,即使定位h1元素的全部内容也无法获取目标名称。执行代码及结果如下:

In [9]: response.xpath("//td[@class='defaultTitleAlign']/h1/text()").get()
Out[9]: '
                        '

解决方案

方法1:XPath直接筛选非空白文本并格式化

使用text()[normalize-space()]选中<h1>下的非空白文本节点,再用normalize-space()处理掉多余空白:

response.xpath("normalize-space(//td[@class='defaultTitleAlign']/h1/text()[normalize-space()])").get()

返回结果:"Terence Crawford"(若需去除引号,可再调用strip('"'))

方法2:提取所有文本节点后过滤拼接

先获取<h1>下的所有文本节点,再过滤掉空白内容并拼接:

text_nodes = response.xpath("//td[@class='defaultTitleAlign']/h1/text()").getall()
target_name = ''.join([node.strip() for node in text_nodes if node.strip()])

方法3:提取元素内全部文本后处理

使用string()获取<h1>内的所有文本(子元素<span>无有效文本,不影响),再去除前后空白:

full_text = response.xpath("string(//td[@class='defaultTitleAlign']/h1)").get()
target_name = full_text.strip()

内容的提问来源于stack exchange,提问作者Muhammad Noman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 03:42:13