You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy与XPath提取网页文本:如何去除空格、换行及NBPS?

清理Scrapy提取的含特殊字符文本的方法

问题背景

网页结构示例:

<div class="snippet-content">
    <h2>First Child</h2>
    <p>Hello</p>
    This is large text ..........
</div>

通过以下XPath提取div下最后一段文本:

text = response.xpath('//div[contains(@class, "snippet-content")]/text()[last()]').get()

提取结果包含冗余的空格、换行符\r\n和NBSP(非断空格,Unicode编码\u00A0),示例输出:

"         \r\nRemarks byNBPS Deputy Prime Minister andNBPS Coordinating Minister for Economic Policies Heng Swee Keat at the Opening of the Bilingualism Carnival on 8 April 2023.                                "

解决方案

方法1:Python字符串方法直接处理

先提取原始文本,再通过字符串替换和清理方法去除冗余字符:

raw_text = response.xpath('//div[contains(@class, "snippet-content")]/text()[last()]').get()
if raw_text:
    # 将NBSP替换为普通空格,去除首尾空白
    clean_text = raw_text.replace('\u00A0', ' ').strip()
    # 可选:合并文本内部的多个连续空白为单个空格
    clean_text = ' '.join(clean_text.split())

方法2:XPath内置函数处理(高效简洁)

利用XPath的normalize-space()函数,该函数会自动去除首尾空白,将所有连续空白(包括换行、空格、NBSP)合并为单个空格:

clean_text = response.xpath('normalize-space(//div[contains(@class, "snippet-content")]/text()[last()])').get()

注:如果需要完全移除NBSP而非替换为空格,可以结合字符串替换使用。

方法3:正则表达式匹配清理

通过Scrapy的re_first()方法直接提取有效文本,再处理NBSP:

clean_text = response.xpath('//div[contains(@class, "snippet-content")]/text()[last()]').re_first(r'\s*(.*?)\s*$')
if clean_text:
    clean_text = clean_text.replace('\u00A0', '')

内容的提问来源于stack exchange,提问作者Ven Nilson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:15:15