如何用Python处理TXT文件:将头记录日期复制到对应数据记录
处理带日期关联的大型结构化文本数据导入问题
现有如下格式的TXT文件(快照示例):
DC000D20221110012022100019 DC011D AV0019000300080180003340501031800481200000 DC011D AV0019000300083180003361901031900071900000 DC011D AV0019000300089180003378701032100515800000 DC000D20221209012022100019 DC011D CG0019000300080220100264401000000000000000 DC011D CG0019000300080220400885101000039990700000 DC011D CG0019000300080220400885101000040013000000文件包含约100万条记录,跨度3年。其中
DC000开头的是头记录,日期位于字符串的[6:14]位置(比如示例中的20221110、20221209)。需要为每条DC011(及DC012、DC013等)记录添加上对应头记录的日期,期望输出格式如下:DC000D20221110012022100019 DC011D20221110 AV0019000300080180003340501031800481200000 DC011D20221110 AV0019000300083180003361901031900071900000 DC011D20221110 AV0019000300089180003378701032100515800000 DC000D20221209012022100019 DC011D20221209 CG0019000300080220100264401000000000000000 DC011D20221209 CG0019000300080220400885101000039990700000 DC011D20221209 CG0019000300080220400885101000040013000000原代码仅提取
DC011记录,丢失DC000头记录,无法后续用fillna()补全日期,且需要同时处理多个模块(DC011/012/013等)。
解决方案1:预处理文本文件,生成带日期的模块记录
逐行遍历源文件,跟踪当前头记录的日期,遇到目标模块记录时直接拼接日期,支持一次性处理多个模块:
source_path = r"\path\CRGDEC\CRGDEC.txt" # 定义需要处理的目标模块列表 target_modules = ['DC011', 'DC012', 'DC013'] # 为每个模块创建输出文件句柄 output_handles = {mod: open(rf"\path\CRGDEC\CRGDEC_{mod}.txt", "w") for mod in target_modules} current_date = "" with open(source_path, 'r') as source_file: for line in source_file: line = line.strip() if not line: continue # 处理头记录,更新当前日期 if line.startswith('DC000'): current_date = line[6:14] # 可选:如果需要在输出文件中保留头记录,取消下面的注释 # for handle in output_handles.values(): # handle.write(line + '\n') else: # 匹配目标模块并生成带日期的记录 for mod in target_modules: if line.startswith(mod): # 提取模块标识后的原始内容 content_part = line[len(mod)+1:].strip() new_line = f"{mod}D{current_date} {content_part}\n" output_handles[mod].write(new_line) break # 关闭所有输出文件 for handle in output_handles.values(): handle.close()
解决方案2:直接导入Pandas处理(适合大数据量)
无需生成中间文件,直接读取文件到DataFrame,利用向前填充补全日期,后续可直接用于探索性分析:
import pandas as pd source_path = r"\path\CRGDEC\CRGDEC.txt" # 读取文件,每行作为一条原始记录 df = pd.read_csv(source_path, header=None, names=['raw_line']) # 提取头记录的日期,非头记录标记为缺失值 df['date'] = df['raw_line'].apply(lambda x: x[6:14] if x.startswith('DC000') else pd.NA) # 向前填充日期,让每条记录继承最近的头记录日期 df['date'] = df['date'].ffill() # 处理DC011模块 dc011_df = df[df['raw_line'].startswith('DC011')].copy() dc011_df['formatted_line'] = dc011_df.apply( lambda row: f"{row['raw_line'][:5]}D{row['date']} {row['raw_line'][6:].strip()}", axis=1 ) # 保存到文件 dc011_df['formatted_line'].to_csv(r"\path\CRGDEC\CRGDEC_DC011.txt", index=False, header=False) # 同理处理DC012模块 dc012_df = df[df['raw_line'].startswith('DC012')].copy() dc012_df['formatted_line'] = dc012_df.apply( lambda row: f"{row['raw_line'][:5]}D{row['date']} {row['raw_line'][6:].strip()}", axis=1 ) dc012_df['formatted_line'].to_csv(r"\path\CRGDEC\CRGDEC_DC012.txt", index=False, header=False)
方案优势说明
- 方案1内存占用低,适合需要严格保留原始文本格式、或不想引入Pandas的场景。
- 方案2无需中间文件,处理逻辑直观,后续可直接基于DataFrame做探索性分析,利用Pandas的大数据优化能力。
内容的提问来源于stack exchange,提问作者dskmrjt
相关产品推荐
相关产品推荐

