You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python处理TXT文件:将头记录日期复制到对应数据记录

处理带日期关联的大型结构化文本数据导入问题

现有如下格式的TXT文件(快照示例):

DC000D20221110012022100019
DC011D           AV0019000300080180003340501031800481200000
DC011D           AV0019000300083180003361901031900071900000
DC011D           AV0019000300089180003378701032100515800000
DC000D20221209012022100019
DC011D           CG0019000300080220100264401000000000000000
DC011D           CG0019000300080220400885101000039990700000
DC011D           CG0019000300080220400885101000040013000000

文件包含约100万条记录,跨度3年。其中DC000开头的是头记录,日期位于字符串的[6:14]位置(比如示例中的20221110、20221209)。需要为每条DC011(及DC012、DC013等)记录添加上对应头记录的日期,期望输出格式如下:

DC000D20221110012022100019
DC011D20221110   AV0019000300080180003340501031800481200000
DC011D20221110   AV0019000300083180003361901031900071900000
DC011D20221110   AV0019000300089180003378701032100515800000
DC000D20221209012022100019
DC011D20221209   CG0019000300080220100264401000000000000000
DC011D20221209   CG0019000300080220400885101000039990700000
DC011D20221209   CG0019000300080220400885101000040013000000

原代码仅提取DC011记录,丢失DC000头记录,无法后续用fillna()补全日期,且需要同时处理多个模块(DC011/012/013等)。


解决方案1:预处理文本文件,生成带日期的模块记录

逐行遍历源文件,跟踪当前头记录的日期,遇到目标模块记录时直接拼接日期,支持一次性处理多个模块:

source_path = r"\path\CRGDEC\CRGDEC.txt"
# 定义需要处理的目标模块列表
target_modules = ['DC011', 'DC012', 'DC013']
# 为每个模块创建输出文件句柄
output_handles = {mod: open(rf"\path\CRGDEC\CRGDEC_{mod}.txt", "w") for mod in target_modules}

current_date = ""

with open(source_path, 'r') as source_file:
    for line in source_file:
        line = line.strip()
        if not line:
            continue
        # 处理头记录,更新当前日期
        if line.startswith('DC000'):
            current_date = line[6:14]
            # 可选:如果需要在输出文件中保留头记录,取消下面的注释
            # for handle in output_handles.values():
            #     handle.write(line + '\n')
        else:
            # 匹配目标模块并生成带日期的记录
            for mod in target_modules:
                if line.startswith(mod):
                    # 提取模块标识后的原始内容
                    content_part = line[len(mod)+1:].strip()
                    new_line = f"{mod}D{current_date}   {content_part}\n"
                    output_handles[mod].write(new_line)
                    break

# 关闭所有输出文件
for handle in output_handles.values():
    handle.close()

解决方案2:直接导入Pandas处理(适合大数据量)

无需生成中间文件,直接读取文件到DataFrame,利用向前填充补全日期,后续可直接用于探索性分析:

import pandas as pd

source_path = r"\path\CRGDEC\CRGDEC.txt"

# 读取文件,每行作为一条原始记录
df = pd.read_csv(source_path, header=None, names=['raw_line'])

# 提取头记录的日期,非头记录标记为缺失值
df['date'] = df['raw_line'].apply(lambda x: x[6:14] if x.startswith('DC000') else pd.NA)

# 向前填充日期,让每条记录继承最近的头记录日期
df['date'] = df['date'].ffill()

# 处理DC011模块
dc011_df = df[df['raw_line'].startswith('DC011')].copy()
dc011_df['formatted_line'] = dc011_df.apply(
    lambda row: f"{row['raw_line'][:5]}D{row['date']}   {row['raw_line'][6:].strip()}",
    axis=1
)
# 保存到文件
dc011_df['formatted_line'].to_csv(r"\path\CRGDEC\CRGDEC_DC011.txt", index=False, header=False)

# 同理处理DC012模块
dc012_df = df[df['raw_line'].startswith('DC012')].copy()
dc012_df['formatted_line'] = dc012_df.apply(
    lambda row: f"{row['raw_line'][:5]}D{row['date']}   {row['raw_line'][6:].strip()}",
    axis=1
)
dc012_df['formatted_line'].to_csv(r"\path\CRGDEC\CRGDEC_DC012.txt", index=False, header=False)

方案优势说明

  • 方案1内存占用低,适合需要严格保留原始文本格式、或不想引入Pandas的场景。
  • 方案2无需中间文件,处理逻辑直观,后续可直接基于DataFrame做探索性分析,利用Pandas的大数据优化能力。

内容的提问来源于stack exchange,提问作者dskmrjt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 09:45:02