You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python拆分TXT文本行并添加分隔符处理NLP语料问题求助

问题原因
  • 变量名错误:定义的读取文件对象是file,但代码中调用了不存在的file1对象读取内容
  • 正则调用错误:re是独立的Python模块,不是列表对象的属性,不能用text.re.split()的写法调用
  • 逻辑错误:没有对拆分后的ID和文本内容做分离处理,也没有过滤按@@拆分后产生的空字符串
解决代码

生成|分隔的自定义格式

import re

# 匹配规则:捕获@@后4-7位数字为ID,后续内容到下一个@@或文本末尾为文本内容
pattern = re.compile(r'@@(\d{4,7})\s(.*?)(?=\s*@@|$)')

with open('/file_directory/file.txt', 'r', encoding='utf-8') as in_file, \
     open('/file_directory/file_cleaned.txt', 'w', encoding='utf-8') as out_file:
    # 写入表头
    out_file.write('ID   | text\n')
    content = in_file.read()
    # 批量匹配所有ID和对应文本
    matches = pattern.findall(content)
    for id_val, text_val in matches:
        out_file.write(f"{id_val} | {text_val}\n")

生成标准CSV格式(推荐,避免文本含分隔符导致格式错乱)

import re
import csv

pattern = re.compile(r'@@(\d{4,7})\s(.*?)(?=\s*@@|$)')

with open('/file_directory/file.txt', 'r', encoding='utf-8') as in_file, \
     open('/file_directory/file_cleaned.csv', 'w', encoding='utf-8', newline='') as out_file:
    writer = csv.writer(out_file)
    # 写入表头
    writer.writerow(['ID', 'text'])
    content = in_file.read()
    matches = pattern.findall(content)
    # 批量写入所有行
    writer.writerows(matches)

内容的提问来源于stack exchange,提问作者Aaron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 17:09:03