You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取文本文件并按句号和换行符拆分生成句子列表

解决方案

原代码存在的问题

  • 执行f.read()后,文件读取指针已移动到文件末尾,后续for line in f无法读取任何内容
  • 未实现按句号(.)和换行符(\n)拆分内容并生成目标列表的核心逻辑

修正后的代码

import re

source_file = "textfile.txt"  # 文件内容为:"The cat jumped over the dog\n. The dog ran. etc."

def load_text_file(source_file):
    res = []
    with open(source_file, encoding="utf8", errors='ignore') as f:
        # 读取全部文件内容
        txt = f.read()
        # 将换行符后跟句号的部分替换为单个句号,统一分隔符形式
        txt_processed = txt.replace('\n.', '.')
        # 按". "拆分,得到不带句号的句子片段
        sentence_parts = re.split(r'\. ', txt_processed)
        # 构建目标列表:每个片段加句号,中间插入空格
        for idx, part in enumerate(sentence_parts):
            res.append(f"{part}.")
            # 最后一个片段后不添加空格
            if idx != len(sentence_parts) - 1:
                res.append(" ")
    return res 

list_res = load_text_file(source_file)
print(list_res)

代码说明

  1. 读取文件:直接读取全部内容,避免指针移动导致的后续读取失效
  2. 统一分隔符:将\n.替换为.,把换行和句号的组合转换成标准的句末句号形式
  3. 拆分句子:使用正则表达式按. (句号加空格)拆分,得到纯句子内容片段
  4. 构建目标列表:给每个片段添加句号,在相邻句子之间插入空格,匹配期望的输出格式

内容的提问来源于stack exchange,提问作者milesabc123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 01:45:31