BASH脚本处理LAMMPS大文件时数据重组异常求助
Fixing Your Large-Scale Data Reordering Script
First, let's break down why your script works for small test cases but fails with real-world data:
Key Issues in the Original Script
- Hardcoded Block Size: You used
3as the block size (for your small test) but forgot to update it to121for your real dataset. This causes your script to read 3-line chunks instead of 121-line chunks, leading to way more data being written than expected (hence the 10x larger files). - Inefficient Shell Loops: Using
for a in $(cat rep_$repl)combined with repeatedhead/tailcalls is extremely slow for 10001 lines. Eachhead/tailinvocation reads the entire input file from scratch, which is catastrophic for large files (120k+ lines). - Cumulative Offset Logic: While the cumulative
jvariable works for small files, it's prone to errors with large offsets and unnecessary file I/O.
Solution 1: Fixed Bash Script (For Smaller Datasets)
If you prefer sticking with bash, update the block size and optimize the loop to avoid redundant file reads:
#!/bin/bash # Configuration NP=13 # Number of replicas (matches 14 columns in log.lammps: 1 index + 13 replicas) BLOCK_SIZE=121 # Block size for real dataset (use 3 for small tests) LOG_FILE="log.lammps" # Clean up old output files rm -f torder* ccccd2_H_*_temp_col.bak2 # Initialize target files with the header (blank line from first input file) for target in $(seq 0 $((NP-1))); do head -1 "ccccd2_H_0_col.bak2" > "ccccd2_H_${target}_temp_col.bak2" done # Process each replica for repl in $(seq 0 $((NP-1))); do echo "Processing replica $repl..." input_file="ccccd2_H_${repl}_col.bak2" # Read log column and process each line (step) awk -v repl="$repl" -v block="$BLOCK_SIZE" -v infile="$input_file" ' BEGIN { # Precompute total lines in input file (skip header) total_lines = 0 while ((getline < infile) > 0) total_lines++ close(infile) header_lines = 1 # Skip the first blank line } { step = NR target = $(repl + 2) target_file = "ccccd2_H_" target "_temp_col.bak2" # Calculate start/end lines for the current block (skip header) start = header_lines + (step - 1) * block + 1 end = header_lines + step * block # Extract the block and append to target file cmd = "sed -n '" start "," end "p' " infile while ((getline line < cmd) > 0) { print line >> target_file } close(cmd) } ' "$LOG_FILE" done
Solution 2: Efficient Awk Script (For Large Datasets)
For 10001 lines and large input files, awk is far more efficient because it reads each input file only once and stores blocks in memory (or processes them on the fly):
#!/bin/bash NP=13 BLOCK_SIZE=121 LOG_FILE="log.lammps" awk -v np="$NP" -v block_size="$BLOCK_SIZE" ' # Step 1: Load all replica blocks into memory (skip header) BEGIN { for (repl=0; repl<np; repl++) { filename = "ccccd2_H_" repl "_col.bak2" line_count = 0 block_idx = 0 while ((getline line < filename) > 0) { line_count++ if (line_count == 1) { # Save the header blank line header = line continue } # Assign line to the correct block current_block = int((line_count - 2)/block_size) + 1 blocks[repl][current_block] = blocks[repl][current_block] line "\n" } close(filename) } # Initialize target files with header for (target=0; target<np; target++) { outfile = "ccccd2_H_" target "_temp_col.bak2" print header > outfile } } # Step 2: Process log file to map blocks to target files { step = NR for (repl=0; repl<np; repl++) { target = $(repl + 2) outfile = "ccccd2_H_" target "_temp_col.bak2" # Append the corresponding block to the target file printf "%s", blocks[repl][step] >> outfile } } ' "$LOG_FILE"
Key Notes
- Block Size: Double-check that
BLOCK_SIZEmatches your real dataset (121 lines per block). For your small test, set it to3. - Header Handling: The scripts assume your input files start with a single blank line (like your test case). If your real files don't have this header, remove the header-related code.
- Memory Usage: The awk script stores all blocks in memory. For 13 replicas × 10001 blocks × 121 lines, this is manageable (each line is short text), but if you encounter memory issues, you can modify the awk script to process blocks on the fly instead of storing them.
内容的提问来源于stack exchange,提问作者spheex12
相关产品推荐
相关产品推荐

