You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多文本文件存储的GB级价格数据合并及新增文件名提取日期列的技术实现咨询

Efficient Solution to Merge Large Price Files with Date Column in Python

Got it, let's break down a practical, memory-efficient solution for your problem—since dealing with millions of lines per file means we can't just load everything into RAM at once. Here's how to pull it off smoothly:

Core Approach

Instead of reading entire files into memory, we'll process each file line-by-line, extract the date from its filename, append that date to every line, and write everything to a single merged file. This keeps memory usage low and works reliably even with dozens of 60MB files.

Step-by-Step Code Implementation

First, make sure you have a recent Python version (3.6+) installed. Use this script, and tweak the paths to match your setup:

import os
import re

# Update these paths to your actual file locations
INPUT_DIRECTORY = "/path/to/your/price_files_folder"
OUTPUT_FILE = "merged_prices_with_date.txt"

with open(OUTPUT_FILE, 'w', encoding='utf-8') as output_handle:
    is_first_file = True

    # Loop through all files in the target directory
    for filename in os.listdir(INPUT_DIRECTORY):
        # Match files following the YYMMDD_Prints.txt pattern (adjust regex if your suffix differs)
        date_match = re.match(r'(\d{6})_Prints\.txt', filename)
        if not date_match:
            print(f"Skipping non-matching file: {filename}")
            continue

        # Extract and format the date (convert YYMMDD to YYYY-MM-DD for readability)
        raw_date = date_match.group(1)
        formatted_date = f"20{raw_date[:2]}-{raw_date[2:4]}-{raw_date[4:6]}"
        
        file_path = os.path.join(INPUT_DIRECTORY, filename)
        
        try:
            with open(file_path, 'r', encoding='utf-8') as input_handle:
                # Handle header row (only write it once to avoid duplicates)
                header = input_handle.readline().strip()
                if is_first_file:
                    output_handle.write(f"{header},Date\n")
                    is_first_file = False

                # Process every data line in the file
                for line in input_handle:
                    cleaned_line = line.strip()
                    if cleaned_line:  # Skip empty lines to keep the merged file clean
                        output_handle.write(f"{cleaned_line},{formatted_date}\n")
            
            print(f"Successfully processed: {filename}")
        
        except Exception as e:
            print(f"Failed to process {filename}: {str(e)}")
            continue

print(f"Merge complete! Output saved to {OUTPUT_FILE}")

Key Details to Note

  • Memory Efficiency: By reading/writing one line at a time, we never load more than a single line into memory—critical for handling files with millions of rows.
  • Date Flexibility: The script converts YYMMDD to YYYY-MM-DD (e.g., 220628 → 2022-06-28) which is easier to work with in later analysis. If you prefer the raw 6-digit format, just replace formatted_date with raw_date.
  • Error Resilience: The try-except block ensures that if one file is corrupted or unreadable, the script will skip it and keep processing others.
  • Encoding: We specify utf-8 for read/write operations—adjust this to match your files' actual encoding (e.g., gbk for some Chinese text files) if needed.

Optional Enhancements

  • Progress Tracking: Install the tqdm package (pip install tqdm) and wrap the file loop with tqdm to see a live progress bar:
    from tqdm import tqdm
    for filename in tqdm(os.listdir(INPUT_DIRECTORY)):
    
  • Custom Delimiters: If your input files use tabs or another delimiter instead of commas, adjust the code to split/join with that character.
  • Compression: Write directly to a compressed file (e.g., gzip) using Python's gzip module to save disk space for the large merged output.

内容的提问来源于stack exchange,提问作者jjustcchillin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 01:07:34