You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于END关键字的大文本分块需求及代码问题问询

文本文件分割需求及解决方案

需求说明

将大型文本文件分割为较小文件,规则如下:

  • 优先让每个小文件包含5行
  • 若第5行不含END关键字,需继续向下读取,直到找到包含END的行,以此为分界生成小文件
  • 完成分割后,从下一行开始重复上述流程

输入示例

CONTINGENCY 'P11:-12.47:DEI:PURDUE CHP GEN'
SET BUS 249831 GENERATION TO 20 MW
END
CONTINGENCY 'P11:-12.47:DEI:PURDUE TG1-2 GENS'
SET BUS 249831 GENERATION TO 15.5 MW
END
CONTINGENCY 'P11:-13.2:DEI:TATE-LYLE BTM GENS'
OPEN BUS 249936
END
CONTINGENCY 'P11:0.342:DEI:08CR_SOL_GEN:1'
REMOVE MACHINE 1 FROM BUS 251904
END

预期输出

# file 1
CONTINGENCY 'P11:-12.47:DEI:PURDUE CHP GEN'
SET BUS 249831 GENERATION TO 20 MW
END
CONTINGENCY 'P11:-12.47:DEI:PURDUE TG1-2 GENS'
SET BUS 249831 GENERATION TO 15.5 MW
END

# file 2
CONTINGENCY 'P11:-13.2:DEI:TATE-LYLE BTM GENS'
OPEN BUS 249936
END
CONTINGENCY 'P11:0.342:DEI:08CR_SOL_GEN:1'
REMOVE MACHINE 1 FROM BUS 251904
END

当前问题代码

现有代码无法满足END关键字的分割条件,代码如下:

import glob
import pandas as pd
import math
import os

if __name__ == "__main__":
    file_dir = os.path.dirname(__file__)
    if file_dir != "":
       os.getcwd()

read_file = glob.glob("*.con")

with open("combined.con", "wb") as outfile:
    for f in read_file:
        with open (f, "rb") as infile:
            outfile.write(infile.read())
        
df0 = pd.read_csv (file_dir + '/combined.con')
count = len(df0)
row_range = 5
block = count // row_range
for line in df0:
    for i in range(block):   
        if not line.startswith("END"):
            start = i * row_range
            stop = (i+1) * row_range
            while True:
                row_range = row_range + 1
                df2 = df0.iloc[start:stop]
                df2.to_csv(f"Contingency_{i}.con", index=False)
                break

复现用数据

df0.to_dict()的前12行数据:

{"CONTINGENCY 'P11:069:MPW::MPW 7G-G7:NON-BES'": 
 { 0: 'TRIP BRANCH FROM BUS 633408 TO BUS 633007 CKT 1', 
   1: 'REMOVE UNIT 7 FROM BUS 633007', 
   2: 'DISCONNECT BUS 633007', 
   3: 'END', 
   4: "CONTINGENCY 'P11:069:MPW::MPW 8AG-G8A:NON-BES'", 
   5: 'TRIP BRANCH FROM BUS 633408 TO BUS 633018 CKT 1', 
   6: 'REMOVE UNIT A FROM BUS 633018', 
   7: 'DISCONNECT BUS 633018', 
   8: 'END', 
   9: "CONTINGENCY 'P11:069:MPW::MPW 8G-G8:NON-BES'", 
  10: 'TRIP BRANCH FROM BUS 633408 TO BUS 633008 CKT 1', 
  11: 'REMOVE UNIT 8 FROM BUS 633008'
 }
}

解决方案

用逐行读取的方式更易实现分割规则,替代原Pandas的处理逻辑:

import glob
import os

def split_contingency_files():
    # 合并所有.con文件到combined.con
    with open("combined.con", "w", encoding="utf-8") as outfile:
        for f in glob.glob("*.con"):
            with open(f, "r", encoding="utf-8") as infile:
                outfile.write(infile.read())
    
    # 读取合并后的文件并分割
    current_block = []
    file_counter = 1
    
    with open("combined.con", "r", encoding="utf-8") as f:
        for line in f:
            line = line.rstrip("\n")
            current_block.append(line)
            
            # 当块达到5行时,检查是否需要继续读取到END
            if len(current_block) == 5:
                if current_block[-1].strip() != "END":
                    for extra_line in f:
                        extra_line = extra_line.rstrip("\n")
                        current_block.append(extra_line)
                        if extra_line.strip() == "END":
                            break
            
            # 块以END结尾时,写入文件
            if current_block and current_block[-1].strip() == "END":
                with open(f"Contingency_{file_counter}.con", "w", encoding="utf-8") as outfile:
                    outfile.write("\n".join(current_block) + "\n")
                current_block = []
                file_counter += 1
    
    # 处理剩余未完成的内容
    if current_block:
        with open(f"Contingency_{file_counter}.con", "w", encoding="utf-8") as outfile:
            outfile.write("\n".join(current_block) + "\n")

if __name__ == "__main__":
    split_contingency_files()

代码说明

  1. 文件合并:将所有.con文件合并为combined.con,采用文本模式避免二进制写入的编码问题
  2. 逐行分割逻辑:
    • 每读取一行就加入当前块
    • 当块的行数达到5时,若最后一行不是END,则继续读取直到找到包含END的行
    • 一旦块的最后一行是END,立即写入新文件,重置块并递增文件计数器
  3. 剩余内容处理:若文件末尾存在未以END结尾的内容,也写入最后一个文件

该代码严格遵循分割规则,确保每个输出文件尽量接近5行且以END结尾。

内容的提问来源于stack exchange,提问作者Sarvesh Gadre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 20:35:53