You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

跨文件匹配loc ID求均值,保留唯一ID对应数值的技术需求

Got it, let's tackle this problem efficiently—especially since you're dealing with hundreds of columns, we need solutions that scale well. Here are two robust approaches that work for both small and large datasets:


1. Awk Script (Fast, Memory-Efficient for Large Files)

Awk is ideal for this kind of text processing because it streams data instead of loading everything into memory, making it perfect for hundreds of columns and large file sizes.

BEGIN {
    OFS = " "  # Set output field separator to space
}

# Automatically detect how many numeric columns follow each key (run only on first line)
NR == 1 {
    m = 0
    for (i = 2; i <= NF; i++) {
        # Adjust this regex to match your key format (example uses "key." prefix)
        if ($i ~ /^key\./) {
            break
        }
        m++
    }
}

# Process each line in both files
{
    # Iterate through each key-value group (key + m numeric columns)
    for (i = 1; i <= NF; i += m+1) {
        key = $i
        count[key]++  # Track how many times the key appears across files
        
        # Accumulate sum for each numeric column
        for (j = 1; j <= m; j++) {
            sum[key,j] += $(i+j)
        }
    }
}

# Calculate averages and print results
END {
    # If you need to preserve the order of keys from the first file, add logic to track order here
    for (key in count) {
        printf "%s", key
        for (j = 1; j <= m; j++) {
            avg = sum[key,j] / count[key]
            # Print as integer if average is whole number, else one decimal place
            if (avg == int(avg)) {
                printf "%s%d", OFS, avg
            } else {
                printf "%s%.1f", OFS, avg
            }
        }
        printf "%s", OFS
    }
    printf "\n"
}

How to use:

Save this as average_keys.awk, then run:

awk -f average_keys.awk file1.txt file2.txt > output.txt

Notes:

  • Adjust the regex /^key\./ to match your actual key format (e.g., if keys start with id_, use /^id_/).
  • If you know the exact number of numeric columns per key upfront, you can hardcode m in the BEGIN block to skip auto-detection (faster).

2. Python Script (Flexible, Easy to Customize)

If you need more control (like adding validation, handling edge cases, or integrating with other tools), a Python script is a great choice. It's readable and easy to modify.

from collections import defaultdict

def calculate_key_averages(file1_path, file2_path):
    # Store key data: count of occurrences, sum of each numeric column
    key_data = defaultdict(lambda: {"count": 0, "sums": []})
    
    def parse_file(file_path):
        with open(file_path, "r") as f:
            for line in f:
                fields = line.strip().split()
                if not fields:
                    continue
                
                num_cols_per_key = None
                idx = 0
                while idx < len(fields):
                    key = fields[idx]
                    
                    # Auto-detect number of numeric columns per key (only once per file)
                    if num_cols_per_key is None:
                        next_key_idx = idx + 1
                        while next_key_idx < len(fields) and not fields[next_key_idx].startswith("key."):
                            next_key_idx += 1
                        num_cols_per_key = next_key_idx - idx - 1
                        # Initialize sum lists for all keys we'll encounter
                        for k in key_data:
                            if not key_data[k]["sums"]:
                                key_data[k]["sums"] = [0.0] * num_cols_per_key
                    
                    # Extract numeric values for current key
                    values = list(map(float, fields[idx+1:idx+1+num_cols_per_key]))
                    
                    # Initialize sum list if this is the first time we see the key
                    if not key_data[key]["sums"]:
                        key_data[key]["sums"] = [0.0] * num_cols_per_key
                    
                    # Update sums and count
                    for i in range(num_cols_per_key):
                        key_data[key]["sums"][i] += values[i]
                    key_data[key]["count"] += 1
                    
                    idx += 1 + num_cols_per_key
    
    # Process both input files
    parse_file(file1_path)
    parse_file(file2_path)
    
    # Build output string
    output_parts = []
    # To preserve key order from the first file, track order during parsing (add a list to store keys)
    for key in sorted(key_data.keys()):
        parts = [key]
        count = key_data[key]["count"]
        for total in key_data[key]["sums"]:
            avg = total / count
            # Format as integer if whole number, else one decimal place
            parts.append(str(int(avg)) if avg.is_integer() else f"{avg:.1f}")
        output_parts.append(" ".join(parts))
    
    return " ".join(output_parts)

# Example usage
if __name__ == "__main__":
    result = calculate_key_averages("file1.txt", "file2.txt")
    print(result)
    # Write to output file
    with open("output.txt", "w") as f:
        f.write(result + "\n")

Notes:

  • Like the awk script, adjust startswith("key.") to match your key format.
  • To preserve the original order of keys (from the first file), add a list variable to track keys when they're first encountered in parse_file, then iterate over that list instead of sorting.

内容的提问来源于stack exchange,提问作者Alex Trevylan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:32:41