You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用regex.split分割文本并将有效分割项导出为txt到指定文件夹

Python文本分割过滤与批量导出解决方案

步骤2:分割结果过滤

用列表推导式即可实现需求,仅排除空字符串、单独的半角连字符,其余内容全部保留,包括带连字符的复合词。
过滤代码示例:

# split_result为regex.split返回的分割结果
filtered_tokens = [token for token in split_result if token and token != '-']

逻辑说明:

  • if token 自动过滤所有空字符串(空字符串在布尔判断中为False)
  • token != '-' 排除单独出现的半角连字符,长度大于1的带连字符内容(如word-hyphen)会正常保留

步骤3:批量导出为txt文件

导出前先确保目标保存文件夹存在,再遍历过滤后的分词结果逐个写入文件即可,注意使用utf-8编码避免特殊字符乱码。

完整可运行代码

# coding: utf-8
import os
import regex

# 1. 文本分割(原有逻辑)
s = "This sentence is in first place and contain a word-hyphen — Hello I am the second sentence and I'm in second place."
split_result = regex.split("[\s\.;!?…»,«\,]", s)

# 2. 结果过滤
filtered_tokens = [token for token in split_result if token and token != '-']

# 3. 批量导出到指定文件夹
SAVE_DIR = "./token_output"  # 可修改为你需要的保存路径
# 自动创建文件夹,已存在则不报错
os.makedirs(SAVE_DIR, exist_ok=True)

for token in filtered_tokens:
    # 文件名用分词内容,后缀为txt
    file_path = os.path.join(SAVE_DIR, f"{token}.txt")
    # 写入文件,编码用utf-8兼容特殊字符
    with open(file_path, "w", encoding="utf-8") as f:
        f.write(token)

注意事项

  • 如果存在重复的分词结果,同名文件会被覆盖,如需保留所有重复项,可修改文件命名逻辑加上序号后缀:
for idx, token in enumerate(filtered_tokens):
    file_path = os.path.join(SAVE_DIR, f"{token}_{idx}.txt")
  • 部分操作系统不支持文件名包含/ \ : * ? " < > |等特殊字符,若分词结果中存在这类字符,可提前做替换处理。

内容的提问来源于stack exchange,提问作者François

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 21:57:03