You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用NLTK将.txt文件分句并导出为.csv文件?

使用NLTK将TXT文本分句并导出为CSV文件

1. 读取TXT文件作为输入

直接通过Python内置的文件操作读取即可,无需手动复制内容。注意指定编码(如utf-8)避免乱码,同时第一次运行需要下载NLTK分句依赖的模型:

import nltk
from nltk.tokenize import sent_tokenize

# 第一次运行需执行,下载分句所需的punkt模型
nltk.download('punkt')

# 读取整个TXT文件内容
with open('你的报纸语料.txt', 'r', encoding='utf-8') as f:
    text_content = f.read()

# 对全文内容分句
sentences = sent_tokenize(text_content)

如果你的TXT文件中每篇文章单独占一行,也可以逐行处理(跳过空行):

sentences = []
with open('你的报纸语料.txt', 'r', encoding='utf-8') as f:
    for line in f:
        cleaned_line = line.strip()
        if cleaned_line:
            sentences.extend(sent_tokenize(cleaned_line))

2. 将分句结果导出为CSV文件

用Python标准库csv即可完成,无需额外安装依赖。可以先写入表头,方便后续标注时识别:

import csv

# 写入CSV文件,newline=''避免生成多余空行
with open('分句结果.csv', 'w', encoding='utf-8', newline='') as csvfile:
    writer = csv.writer(csvfile)
    # 写入表头
    writer.writerow(["句子"])
    # 逐行写入每个分句
    for sent in sentences:
        writer.writerow([sent])

如果习惯用pandas(需提前执行pip install pandas),代码会更简洁:

import pandas as pd

df = pd.DataFrame({"句子": sentences})
df.to_csv('分句结果.csv', encoding='utf-8', index=False)

完整整合代码

把读取和写入逻辑合并后的完整代码:

import nltk
from nltk.tokenize import sent_tokenize
import csv

# 下载分句模型(第一次运行执行)
nltk.download('punkt')

# 读取TXT文件
with open('报纸语料.txt', 'r', encoding='utf-8') as f:
    text = f.read()

# 分句处理
sentences = sent_tokenize(text)

# 写入CSV文件
with open('分句结果.csv', 'w', encoding='utf-8', newline='') as csvfile:
    writer = csv.writer(csvfile)
    writer.writerow(["句子"])
    for sentence in sentences:
        writer.writerow([sentence])

内容的提问来源于stack exchange,提问作者Tess

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 03:20:14