Python3中使用Sastrawi处理停用词后,如何将结果保存至新文件?
解决停用词处理结果保存到文件的问题
嘿,你的代码已经搞定了停用词移除的核心逻辑,现在只需要添加上写入文件的步骤就行!我给你调整一下代码,同时解释关键部分:
修改后的完整代码(推荐用with语句更安全)
from Sastrawi.StopWordRemover.StopWordRemoverFactory import StopWordRemoverFactory factory = StopWordRemoverFactory() stopword = factory.create_stop_word_remover() # 用with语句同时管理输入和输出文件,自动处理关闭 with open('dataset/filter.txt', 'r', encoding='utf-8') as input_file, \ open('dataset/processed_result.txt', 'w', encoding='utf-8') as output_file: for data in input_file: # 处理停用词,顺便去除每行末尾的换行符(可选,避免重复换行) processed_line = stopword.remove(data.strip()) # 写入处理后的内容,记得加换行符 output_file.write(processed_line + '\n')
关键步骤解释
- 指定编码
encoding='utf-8':处理印尼语文本时一定要加这个,不然很容易出现乱码问题,保证读写的文本编码一致。 with语句的优势:它会自动帮你关闭文件,哪怕中间代码出错也不会导致文件资源泄漏,比手动调用close()更可靠。strip()和'\n':data.strip()是去掉每行原有的换行符,之后写入时再加'\n',确保输出文件的每行格式整齐,不会出现空行或者换行混乱。
如果你想保持原来的代码结构(不用with),也可以这样写:
from Sastrawi.StopWordRemover.StopWordRemoverFactory import StopWordRemoverFactory factory = StopWordRemoverFactory() stopword = factory.create_stop_word_remover() input_file = open('dataset/filter.txt', 'r', encoding='utf-8') output_file = open('dataset/processed_result.txt', 'w', encoding='utf-8') for data in input_file: processed_line = stopword.remove(data.strip()) output_file.write(processed_line + '\n') # 一定要记得手动关闭文件! input_file.close() output_file.close()
这样处理后,你就能在dataset/processed_result.txt里看到处理好的无停用词文本啦~
内容的提问来源于stack exchange,提问作者hensam
相关产品推荐
相关产品推荐

