You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取网页description时丢失strong、br标签内容问题

问题原因

你使用的::textCSS伪选择器仅会提取匹配元素下的直接纯文本节点,会自动过滤所有子标签(包括你场景里的<strong>、<br>)和标签对应的结构关系,因此无法获取带完整HTML格式的描述内容。

调整方案

直接选中目标<p>元素后调用提取方法,不要追加::text伪选择器,Scrapy会直接返回选中元素对应的完整HTML源码:

  • 单p标签场景(和你目标页面结构匹配)
'description': response.css('#tab-description p').get()
  • 多p标签场景(如果描述块包含多个段落,需要拼接所有段落HTML)
'description': ''.join(response.css('#tab-description p').getall())

如果习惯使用XPath语法,也可以用等价写法:

'description': response.xpath('//*[@id="tab-description"]/p').get()
返回效果

上述代码执行后会直接返回你期望的带完整标签结构的内容:

<p>    <strong>Brand Name: </strong>None  <br>  <strong>Gender: </strong>Unisex  <br>  <strong>Age Range: </strong>12-15 Years  <br>  <strong>Age Range: </strong>Grownups  <br>  <strong>Material: </strong>Paper  <br>  <strong>Style: </strong>Landscape  <br>  <strong>Model Number: </strong>SMW783   </p>
补充说明
  • 只要在选择器后追加::text(CSS)或者/text()(XPath),就只会提取纯文本节点,所有HTML标签都会被忽略
  • 如果不需要外层的<p></p>标签,仅要p标签内部的HTML内容,可以用XPath提取所有子节点后拼接:
'description': ''.join(response.xpath('//*[@id="tab-description"]/p/node()').getall())

内容的提问来源于stack exchange,提问作者Hannah James

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 17:45:35