多文本文件存储的GB级价格数据合并及新增文件名提取日期列的技术实现咨询
Efficient Solution to Merge Large Price Files with Date Column in Python
Got it, let's break down a practical, memory-efficient solution for your problem—since dealing with millions of lines per file means we can't just load everything into RAM at once. Here's how to pull it off smoothly:
Core Approach
Instead of reading entire files into memory, we'll process each file line-by-line, extract the date from its filename, append that date to every line, and write everything to a single merged file. This keeps memory usage low and works reliably even with dozens of 60MB files.
Step-by-Step Code Implementation
First, make sure you have a recent Python version (3.6+) installed. Use this script, and tweak the paths to match your setup:
import os import re # Update these paths to your actual file locations INPUT_DIRECTORY = "/path/to/your/price_files_folder" OUTPUT_FILE = "merged_prices_with_date.txt" with open(OUTPUT_FILE, 'w', encoding='utf-8') as output_handle: is_first_file = True # Loop through all files in the target directory for filename in os.listdir(INPUT_DIRECTORY): # Match files following the YYMMDD_Prints.txt pattern (adjust regex if your suffix differs) date_match = re.match(r'(\d{6})_Prints\.txt', filename) if not date_match: print(f"Skipping non-matching file: {filename}") continue # Extract and format the date (convert YYMMDD to YYYY-MM-DD for readability) raw_date = date_match.group(1) formatted_date = f"20{raw_date[:2]}-{raw_date[2:4]}-{raw_date[4:6]}" file_path = os.path.join(INPUT_DIRECTORY, filename) try: with open(file_path, 'r', encoding='utf-8') as input_handle: # Handle header row (only write it once to avoid duplicates) header = input_handle.readline().strip() if is_first_file: output_handle.write(f"{header},Date\n") is_first_file = False # Process every data line in the file for line in input_handle: cleaned_line = line.strip() if cleaned_line: # Skip empty lines to keep the merged file clean output_handle.write(f"{cleaned_line},{formatted_date}\n") print(f"Successfully processed: {filename}") except Exception as e: print(f"Failed to process {filename}: {str(e)}") continue print(f"Merge complete! Output saved to {OUTPUT_FILE}")
Key Details to Note
- Memory Efficiency: By reading/writing one line at a time, we never load more than a single line into memory—critical for handling files with millions of rows.
- Date Flexibility: The script converts
YYMMDDtoYYYY-MM-DD(e.g.,220628→2022-06-28) which is easier to work with in later analysis. If you prefer the raw 6-digit format, just replaceformatted_datewithraw_date. - Error Resilience: The
try-exceptblock ensures that if one file is corrupted or unreadable, the script will skip it and keep processing others. - Encoding: We specify
utf-8for read/write operations—adjust this to match your files' actual encoding (e.g.,gbkfor some Chinese text files) if needed.
Optional Enhancements
- Progress Tracking: Install the
tqdmpackage (pip install tqdm) and wrap the file loop withtqdmto see a live progress bar:from tqdm import tqdm for filename in tqdm(os.listdir(INPUT_DIRECTORY)): - Custom Delimiters: If your input files use tabs or another delimiter instead of commas, adjust the code to split/join with that character.
- Compression: Write directly to a compressed file (e.g.,
gzip) using Python'sgzipmodule to save disk space for the large merged output.
内容的提问来源于stack exchange,提问作者jjustcchillin
相关产品推荐
相关产品推荐

