如何从HTML片段中提取'Terence Crawford'并排除span元素?
问题:提取HTML中排除子元素的目标文本
我无法从以下HTML片段中提取名称Terence Crawford,难点在于需要排除同一父元素内的<span>元素:
<td colspan="3" style="position:relative;" class="defaultTitleAlign"> <h1 style="display:inline-block;margin-right:5px;line-height:30px;"> <span style="font-weight:bold;"><i class="fas fa-crown" style="color:#f6b501 !important;"></i></span> "Terence Crawford" </h1> <div style="width:100%;position:relative;margin-top:5px;"> </div> </td>
我尝试通过指定class属性defaultTitleAlign和style属性定位<h1>元素提取文本,但仅返回空白换行符,即使定位h1元素的全部内容也无法获取目标名称。执行代码及结果如下:
In [9]: response.xpath("//td[@class='defaultTitleAlign']/h1/text()").get() Out[9]: ' '
解决方案
方法1:XPath直接筛选非空白文本并格式化
使用text()[normalize-space()]选中<h1>下的非空白文本节点,再用normalize-space()处理掉多余空白:
response.xpath("normalize-space(//td[@class='defaultTitleAlign']/h1/text()[normalize-space()])").get()
返回结果:"Terence Crawford"(若需去除引号,可再调用strip('"'))
方法2:提取所有文本节点后过滤拼接
先获取<h1>下的所有文本节点,再过滤掉空白内容并拼接:
text_nodes = response.xpath("//td[@class='defaultTitleAlign']/h1/text()").getall() target_name = ''.join([node.strip() for node in text_nodes if node.strip()])
方法3:提取元素内全部文本后处理
使用string()获取<h1>内的所有文本(子元素<span>无有效文本,不影响),再去除前后空白:
full_text = response.xpath("string(//td[@class='defaultTitleAlign']/h1)").get() target_name = full_text.strip()
内容的提问来源于stack exchange,提问作者Muhammad Noman
相关产品推荐
相关产品推荐

