如何用Python移除文本文件中的指定词语?
修正你的代码,实现移除特定词语的功能
先说说你当前代码里的几个关键问题:
- 你写的
for cw in "commonWords"是在遍历字符串"commonWords"的每一个字符,而不是读取common.txt里的待移除词语 - 读取
common.txt后,你直接拿到的是整个文件的文本内容,没有把它拆分成单个的单词列表 - 代码里的
text变量没有定义,得先加载你要处理的目标文本文件
接下来是修正后的代码,我会一步步说明逻辑:
步骤1:正确读取待移除词语列表
首先读取common.txt,并把内容拆分成单个单词。假设common.txt里每个词语占一行,我们用splitlines()来分割:
# 读取待移除词语文件,用with语句自动管理文件关闭 with open("common.txt", "r") as common_file: common_words = common_file.read().splitlines()
步骤2:加载要处理的目标文本
假设你要处理的文本文件叫target.txt,先把它的内容读进来:
# 读取目标文本文件 with open("target.txt", "r") as target_file: text = target_file.read()
步骤3:遍历词语并执行替换
现在遍历common_words列表里的每个词语,把它们替换成空格(如果想直接删除就换成空字符串""):
# 移除每个待移除词语,先strip()避免词语带多余空白 for word in common_words: clean_word = word.strip() if clean_word: # 跳过空行或空白词语 text = text.replace(clean_word, " ")
完整可运行代码
把上面的步骤整合起来,完整代码如下:
# 读取待移除词语列表 with open("common.txt", "r") as common_file: common_words = common_file.read().splitlines() # 读取目标文本 with open("target.txt", "r") as target_file: text = target_file.read() # 执行替换操作 for word in common_words: clean_word = word.strip() if clean_word: text = text.replace(clean_word, " ") # 可选:把处理后的文本保存到新文件 with open("cleaned_text.txt", "w") as output_file: output_file.write(text)
额外优化提示
- 如果你的
common.txt里的词语是用空格分隔而不是换行,把splitlines()改成split()即可 - 如果要避免误替换(比如把"cat"从"category"里错误移除),可以用正则表达式匹配完整单词:
import re for word in common_words: clean_word = word.strip() if clean_word: # \b 表示单词边界,re.escape()处理词语里的特殊字符 text = re.sub(rf"\b{re.escape(clean_word)}\b", " ", text)
内容的提问来源于stack exchange,提问作者fskoft
相关产品推荐
相关产品推荐

