如何用NLTK将.txt文件分句并导出为.csv文件?
使用NLTK将TXT文本分句并导出为CSV文件
1. 读取TXT文件作为输入
直接通过Python内置的文件操作读取即可,无需手动复制内容。注意指定编码(如utf-8)避免乱码,同时第一次运行需要下载NLTK分句依赖的模型:
import nltk from nltk.tokenize import sent_tokenize # 第一次运行需执行,下载分句所需的punkt模型 nltk.download('punkt') # 读取整个TXT文件内容 with open('你的报纸语料.txt', 'r', encoding='utf-8') as f: text_content = f.read() # 对全文内容分句 sentences = sent_tokenize(text_content)
如果你的TXT文件中每篇文章单独占一行,也可以逐行处理(跳过空行):
sentences = [] with open('你的报纸语料.txt', 'r', encoding='utf-8') as f: for line in f: cleaned_line = line.strip() if cleaned_line: sentences.extend(sent_tokenize(cleaned_line))
2. 将分句结果导出为CSV文件
用Python标准库csv即可完成,无需额外安装依赖。可以先写入表头,方便后续标注时识别:
import csv # 写入CSV文件,newline=''避免生成多余空行 with open('分句结果.csv', 'w', encoding='utf-8', newline='') as csvfile: writer = csv.writer(csvfile) # 写入表头 writer.writerow(["句子"]) # 逐行写入每个分句 for sent in sentences: writer.writerow([sent])
如果习惯用pandas(需提前执行pip install pandas),代码会更简洁:
import pandas as pd df = pd.DataFrame({"句子": sentences}) df.to_csv('分句结果.csv', encoding='utf-8', index=False)
完整整合代码
把读取和写入逻辑合并后的完整代码:
import nltk from nltk.tokenize import sent_tokenize import csv # 下载分句模型(第一次运行执行) nltk.download('punkt') # 读取TXT文件 with open('报纸语料.txt', 'r', encoding='utf-8') as f: text = f.read() # 分句处理 sentences = sent_tokenize(text) # 写入CSV文件 with open('分句结果.csv', 'w', encoding='utf-8', newline='') as csvfile: writer = csv.writer(csvfile) writer.writerow(["句子"]) for sentence in sentences: writer.writerow([sentence])
内容的提问来源于stack exchange,提问作者Tess
相关产品推荐
相关产品推荐

