You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫时如何移除列表中的\t\n\r转义空白字符

问题原因

str.strip() 仅支持移除字符串首尾的空白类字符,无法处理字符串中间夹杂的 \n、\r、\t 转义字符,这是清理后仍有残留转义符的核心原因。

可行解决方案

你可以根据需求选择以下任意一种处理方式:

方案1:用正则表达式精准移除指定转义符

无需破坏文本中的正常空格,适配表头清理场景:

import re

a = []
for i in head:
    # 只移除\n、\r、\t三类转义符,最后清理首尾残留的空白
    cleaned = re.sub(r'[\n\r\t]', '', i.text).strip()
    a.append(cleaned)
print(a)

处理后示例表头Company Type\n\r\n\t\t\t\t\t\t\t\t?会被转为Company Type?。

方案2:无正则依赖的链式替换

不想引入re库的话可以直接用replace链式处理:

a = []
for i in head:
    cleaned = i.text.replace('\n','').replace('\r','').replace('\t','').strip()
    a.append(cleaned)
print(a)

方案3:合并多余空白(可选)

如果你需要同时把文本中连续的多个空白合并为单个空格,可以用split+join的写法,连strip步骤都可以省略:

a = []
for i in head:
    # 按任意空白分割后用单个空格拼接,自动清理首尾空白、合并连续空白
    cleaned = ' '.join(i.text.split())
    a.append(cleaned)
print(a)

这种方式处理示例表头会得到Company Type ?(?前会保留单个空格),适合需要统一空白格式的场景。

内容的提问来源于stack exchange,提问作者Vigneswaran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 14:54:06