You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

坐标数据列内补零与列拆分的批量处理技术求助

Solution to Fix Coordinate Data Gaps and Split Record Numbers

Let's break down how to solve your problem step by step. The core issues are fixing missing leading zeros that split columns, and separating the record number from the first coordinate column. I'll provide a Python script that integrates these fixes with your existing top-row deletion and column filtering logic.

Approach Overview

  1. Batch Process Files: Loop through all target files in your directory.
  2. Skip Top Rows: Maintain your existing logic to remove unwanted header rows.
  3. Fix Column Splits: Detect when the first coordinate column is split into a record number and coordinate (due to missing leading zeros), then combine them.
  4. Pad Leading Zeros: Ensure record numbers have consistent length (e.g., 4 digits for thousands-range records).
  5. Split Record Number: Extract the record number from the merged first coordinate column.
  6. Output Clean Data: Write the final 3-column structure (record number + coordinate 1 + coordinate 2) to new files.

Complete Python Code

import os
import re

# --------------------------
# Configure these parameters
# --------------------------
INPUT_DIR = "./your_input_files"  # Replace with your directory path
OUTPUT_DIR = "./cleaned_output"
SKIP_TOP_ROWS = 5  # Number of top rows to delete (match your existing code)
RECORD_NUM_LENGTH = 4  # Fixed length for record numbers (adjust based on your max record number, e.g., 3 for up to 999)
DELIMITER = "\t"  # Output delimiter (use "," for CSV)

# Create output directory if it doesn't exist
os.makedirs(OUTPUT_DIR, exist_ok=True)

# Process each file in the input directory
for filename in os.listdir(INPUT_DIR):
    # Skip non-data files (adjust extensions as needed)
    if not filename.endswith((".txt", ".csv")):
        continue
    
    input_path = os.path.join(INPUT_DIR, filename)
    output_path = os.path.join(OUTPUT_DIR, f"cleaned_{filename}")
    
    with open(input_path, "r") as infile, open(output_path, "w") as outfile:
        # Skip the specified number of top rows
        for _ in range(SKIP_TOP_ROWS):
            next(infile)
        
        # Write header for output (optional, adjust as needed)
        outfile.write(f"record_number{DELIMITER}coordinate1{DELIMITER}coordinate2\n")
        
        for line_num, line in enumerate(infile, start=SKIP_TOP_ROWS+1):
            line = line.strip()
            if not line:
                continue  # Skip empty lines
            
            tokens = line.split()  # Split line into individual elements
            
            try:
                # Determine if the first coordinate is merged with record number or split
                if re.match(r"^\d+\.\d+$", tokens[0]):
                    # Case 1: Record number is merged with coordinate1 (e.g., "01123.45" or "1123.45")
                    merged_token = tokens[0]
                    # Split into record number (first N digits) and coordinate1 (remaining)
                    record_num_str = merged_token[:RECORD_NUM_LENGTH]
                    coord1_str = merged_token[RECORD_NUM_LENGTH:]
                    coord2_str = tokens[1]
                elif re.match(r"^\d+$", tokens[0]) and re.match(r"^\d+\.\d+$", tokens[1]):
                    # Case 2: Record number is a separate column (e.g., "1 123.45 678.90")
                    record_num_str = tokens[0]
                    coord1_str = tokens[1]
                    coord2_str = tokens[2]
                else:
                    raise ValueError("Unexpected line format")
                
                # Pad record number with leading zeros to fixed length
                record_num = record_num_str.zfill(RECORD_NUM_LENGTH)
                
                # Validate coordinates are numbers
                coord1 = float(coord1_str)
                coord2 = float(coord2_str)
                
                # Write cleaned line to output
                outfile.write(f"{record_num}{DELIMITER}{coord1}{DELIMITER}{coord2}\n")
            
            except (IndexError, ValueError) as e:
                # Log errors for debugging (optional)
                print(f"Warning: Skipping line {line_num} in {filename} - {e}: {line}")
    
    print(f"Successfully processed: {filename} → {output_path}")

print("All files processed!")

Key Adjustments for Your Data

  1. Record Number Length: If your max record number is, say, 999, set RECORD_NUM_LENGTH to 3 instead of 4. This ensures correct splitting of merged record numbers and coordinates.
  2. File Extensions: Update the endswith check to match your file types (e.g., .dat).
  3. Delimiter: Change DELIMITER to "," if you need CSV output instead of tab-separated.
  4. Skip Rows: Adjust SKIP_TOP_ROWS to match the number of header rows your existing code deletes.

How It Works

  • Line-by-Line Processing: This avoids issues with inconsistent column counts caused by missing leading zeros.
  • Regex Matching: Identifies whether the first element is a merged record-coordinate or a standalone record number.
  • Zero Padding: Uses zfill() to add leading zeros, ensuring consistent record number formatting.
  • Error Handling: Skips problematic lines and logs warnings so you can review any edge cases.

内容的提问来源于stack exchange,提问作者gad raifman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:17:08