You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python按相同Tree名称拆分文本文件并分存独立文件

按Tree名称分组并写入不同文件的Python实现

原始数据

Apple TreeTwo
Banana TreeOne
Juice TreeOne
Pineapple TreeThree
Berries TreeThree

需求目标

将Tree名称相同的行分组,分别写入不同文件:

  • file1.txt
Banana TreeOne
Juice TreeOne
  • file2.txt
Apple TreeTwo
  • file3.txt
Pineapple
Berries

遇到的问题

直接对文件对象调用groupby方法会触发no attribute groupby错误,因为文件对象本身没有该方法。尝试的代码如下:

f = open('data.txt' , 'r')
f_splits = [v for k, v in f.groupby()]
for f_split in f_splits:
    print(f_split, sep = '\n')

解决方案

方法一:用字典手动分组后写入文件

核心逻辑是用字典存储每个Tree对应的行列表,遍历完成后再分别写入文件,同时处理file3.txt的特殊格式需求。

# 初始化字典,键为Tree名称,值为对应行的列表
tree_groups = {}

# 读取原始文件内容
with open('data.txt', 'r', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        if not line:
            continue  # 跳过空行
        # 按空格分割,只分一次,避免内容含空格的情况
        parts = line.split(maxsplit=1)
        if len(parts) != 2:
            continue  # 跳过格式异常的行
        content, tree_name = parts
        # 将行添加到对应Tree的分组中
        if tree_name not in tree_groups:
            tree_groups[tree_name] = []
        tree_groups[tree_name].append(line)

# 遍历分组,写入不同文件
file_num = 1
for tree_name, lines in tree_groups.items():
    filename = f'file{file_num}.txt'
    with open(filename, 'w', encoding='utf-8') as out_f:
        if tree_name == 'TreeThree':
            # 只保留内容部分,去掉TreeThree后缀
            output_lines = [line.split(maxsplit=1)[0] for line in lines]
        else:
            output_lines = lines
        out_f.write('\n'.join(output_lines) + '\n')
    file_num += 1

方法二:使用itertools.groupby(需先排序)

如果想用groupby,需要先按Tree名称排序,因为groupby仅对连续的相同键分组。

from itertools import groupby

# 读取并预处理数据,按Tree名称排序
with open('data.txt', 'r', encoding='utf-8') as f:
    lines = []
    for line in f:
        line = line.strip()
        if line:
            parts = line.split(maxsplit=1)
            if len(parts) == 2:
                lines.append(parts)
    # 按Tree名称排序,确保相同Tree的行连续
    lines.sort(key=lambda x: x[1])

# 分组写入文件
file_num = 1
for tree_name, group in groupby(lines, key=lambda x: x[1]):
    filename = f'file{file_num}.txt'
    with open(filename, 'w', encoding='utf-8') as out_f:
        if tree_name == 'TreeThree':
            out_lines = [item[0] for item in group]
        else:
            out_lines = [f'{item[0]} {item[1]}' for item in group]
        out_f.write('\n'.join(out_lines) + '\n')
    file_num += 1

说明

  • 两种方法都处理了file3.txt的特殊格式要求,同时加入了空行、异常行的过滤逻辑,增强代码鲁棒性。
  • 使用with open语法可自动管理文件关闭,避免资源泄漏。

内容的提问来源于stack exchange,提问作者nora job

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 09:20:42