You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BASH脚本处理LAMMPS大文件时数据重组异常求助

Fixing Your Large-Scale Data Reordering Script

First, let's break down why your script works for small test cases but fails with real-world data:

Key Issues in the Original Script

  1. Hardcoded Block Size: You used 3 as the block size (for your small test) but forgot to update it to 121 for your real dataset. This causes your script to read 3-line chunks instead of 121-line chunks, leading to way more data being written than expected (hence the 10x larger files).
  2. Inefficient Shell Loops: Using for a in $(cat rep_$repl) combined with repeated head/tail calls is extremely slow for 10001 lines. Each head/tail invocation reads the entire input file from scratch, which is catastrophic for large files (120k+ lines).
  3. Cumulative Offset Logic: While the cumulative j variable works for small files, it's prone to errors with large offsets and unnecessary file I/O.

Solution 1: Fixed Bash Script (For Smaller Datasets)

If you prefer sticking with bash, update the block size and optimize the loop to avoid redundant file reads:

#!/bin/bash
# Configuration
NP=13          # Number of replicas (matches 14 columns in log.lammps: 1 index + 13 replicas)
BLOCK_SIZE=121 # Block size for real dataset (use 3 for small tests)
LOG_FILE="log.lammps"

# Clean up old output files
rm -f torder* ccccd2_H_*_temp_col.bak2

# Initialize target files with the header (blank line from first input file)
for target in $(seq 0 $((NP-1))); do
    head -1 "ccccd2_H_0_col.bak2" > "ccccd2_H_${target}_temp_col.bak2"
done

# Process each replica
for repl in $(seq 0 $((NP-1))); do
    echo "Processing replica $repl..."
    input_file="ccccd2_H_${repl}_col.bak2"
    
    # Read log column and process each line (step)
    awk -v repl="$repl" -v block="$BLOCK_SIZE" -v infile="$input_file" '
        BEGIN {
            # Precompute total lines in input file (skip header)
            total_lines = 0
            while ((getline < infile) > 0) total_lines++
            close(infile)
            header_lines = 1 # Skip the first blank line
        }
        {
            step = NR
            target = $(repl + 2)
            target_file = "ccccd2_H_" target "_temp_col.bak2"
            
            # Calculate start/end lines for the current block (skip header)
            start = header_lines + (step - 1) * block + 1
            end = header_lines + step * block
            
            # Extract the block and append to target file
            cmd = "sed -n '" start "," end "p' " infile
            while ((getline line < cmd) > 0) {
                print line >> target_file
            }
            close(cmd)
        }
    ' "$LOG_FILE"
done

Solution 2: Efficient Awk Script (For Large Datasets)

For 10001 lines and large input files, awk is far more efficient because it reads each input file only once and stores blocks in memory (or processes them on the fly):

#!/bin/bash
NP=13
BLOCK_SIZE=121
LOG_FILE="log.lammps"

awk -v np="$NP" -v block_size="$BLOCK_SIZE" '
    # Step 1: Load all replica blocks into memory (skip header)
    BEGIN {
        for (repl=0; repl<np; repl++) {
            filename = "ccccd2_H_" repl "_col.bak2"
            line_count = 0
            block_idx = 0
            while ((getline line < filename) > 0) {
                line_count++
                if (line_count == 1) {
                    # Save the header blank line
                    header = line
                    continue
                }
                # Assign line to the correct block
                current_block = int((line_count - 2)/block_size) + 1
                blocks[repl][current_block] = blocks[repl][current_block] line "\n"
            }
            close(filename)
        }

        # Initialize target files with header
        for (target=0; target<np; target++) {
            outfile = "ccccd2_H_" target "_temp_col.bak2"
            print header > outfile
        }
    }

    # Step 2: Process log file to map blocks to target files
    {
        step = NR
        for (repl=0; repl<np; repl++) {
            target = $(repl + 2)
            outfile = "ccccd2_H_" target "_temp_col.bak2"
            # Append the corresponding block to the target file
            printf "%s", blocks[repl][step] >> outfile
        }
    }
' "$LOG_FILE"

Key Notes

  • Block Size: Double-check that BLOCK_SIZE matches your real dataset (121 lines per block). For your small test, set it to 3.
  • Header Handling: The scripts assume your input files start with a single blank line (like your test case). If your real files don't have this header, remove the header-related code.
  • Memory Usage: The awk script stores all blocks in memory. For 13 replicas × 10001 blocks × 121 lines, this is manageable (each line is short text), but if you encounter memory issues, you can modify the awk script to process blocks on the fly instead of storing them.

内容的提问来源于stack exchange,提问作者spheex12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:16:29