Python程序去除文本标点时引号无法移除的问题求解
问题原因
你代码本身没有逻辑问题,也和IDE/编译器无关,核心原因是你输入文本中使用的是「智能弯引号」(左双引号“、右双引号”),属于Unicode扩展字符,不在Python标准库string.punctuation的覆盖范围内:
string.punctuation仅包含ASCII半角标点,取值为:!"#$%&'()*+,-./:;<=>?@[\]^_{|}~`,没有收录全角、弯引号类特殊标点。
修复方案
只需要调整del_punctuation函数的标点集合,把需要移除的特殊引号加进去即可,同时建议读取文件时指定编码避免乱码:
import string def del_punctuation(item): ''' 该函数用于删除单词中的标点符号 ''' # 在原有ASCII标点基础上,加入常见的弯引号字符 punctuation = string.punctuation + '“”‘’' for c in item: if c in punctuation: item = item.replace(c, '') return item def break_into_words(filename): ''' 该函数读取文件,将内容拆分为全小写的单词列表 ''' # 读取文件指定utf-8编码,避免系统默认编码导致乱码 with open(filename, 'r', encoding='utf-8') as book: words_list = [] for line in book: for item in line.split(): item = del_punctuation(item) item=item.lower() words_list.append(item) return words_list print(break_into_words('input.txt'))
修改后运行即可完全移除所有残留的引号。如果后续遇到其他未被识别的特殊标点,直接加到punctuation变量的拼接字符串里即可。
内容的提问来源于stack exchange,提问作者Safraz
相关产品推荐
相关产品推荐

