Regex替换非字母字符为空格不生效,输出无间隔内容如何解决
代码调整方案
问题根因
- 原正则的匹配逻辑本身可以实现匹配所有非英文字母的需求,核心问题是嵌套列表的处理方式错误:Python 原生列表没有
map方法,直接对嵌套列表调用map会将整个子列表作为参数传入替换函数,无法正确处理每个字符串元素。 - 原正则可简化为
r"[^a-zA-Z]",等价于匹配所有非英文字母字符,逻辑更简洁。
调整后的完整代码
import re edit = [["this, is sample 1."],["this!is @ sample two.*"]] def clearpunch(sentence): clean = re.sub(r"[^a-zA-Z]", r" ", sentence) return clean # 逐层遍历嵌套列表处理每个字符串 edit = [[clearpunch(s) for s in sublist] for sublist in edit] print(edit)
如果你使用的是pandas数据结构,可改用如下写法:
import pandas as pd import re edit = pd.Series(["this, is sample 1.", "this!is @ sample two.*"]) edit = edit.str.replace(r"[^a-zA-Z]", " ", regex=True) print(edit.tolist())
输出结果
运行后得到预期输出:[["this is sample "], ["this is sample two "]]
内容的提问来源于stack exchange,提问作者Yatin Kumar
相关产品推荐
相关产品推荐

