You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python实现两个不同长度文件的数据搜索与匹配?

Got it, let's tackle this problem step by step. The core goal here is to match the shared ID between your two files so we can append the latitude and longitude from location.dat to the corresponding entries in input.dat.

Approach Breakdown

First, we'll load the reference data into a dictionary where the location ID is the key—this lets us look up coordinates in constant time (O(1)), which is way faster than scanning the entire reference file every time we need a match. Then, we'll iterate through the dataset file, pull the ID from each line, fetch the matching coordinates, and merge the data into a new output file.

Complete Python Script

Here's a fully functional script with comments to walk you through each part:

def load_reference_data(ref_file_path):
    location_dict = {}
    with open(ref_file_path, 'r') as f:
        # Split the entire file into individual space-separated values
        raw_data = f.read().split()
        # Process data in chunks of 3: ID, latitude, longitude
        for idx in range(0, len(raw_data), 3):
            loc_id = raw_data[idx]
            latitude = float(raw_data[idx+1])
            longitude = float(raw_data[idx+2])
            location_dict[loc_id] = (latitude, longitude)
    return location_dict

def match_and_merge(input_file_path, location_map, output_file='merged_output.dat'):
    with open(input_file_path, 'r') as infile, open(output_file, 'w') as outfile:
        for line in infile:
            cleaned_line = line.strip()
            if not cleaned_line:
                continue  # Skip empty lines to avoid errors
            
            line_parts = cleaned_line.split()
            input_id = line_parts[0]
            
            # Check if the ID exists in our reference data
            if input_id in location_map:
                lat, lon = location_map[input_id]
                # Combine original line data with coordinates
                merged_entry = ' '.join(line_parts + [str(lat), str(lon)])
                outfile.write(f"{merged_entry}\n")
            else:
                # Handle missing IDs (adjust this to your needs)
                print(f"Warning: ID {input_id} not found in reference data—skipping this entry.")

# Run the script
if __name__ == "__main__":
    # Load reference coordinates first
    location_data = load_reference_data('location.dat')
    # Merge with input dataset and save results
    match_and_merge('input.dat', location_data)
    print("Merge complete! Check merged_output.dat for your combined data.")

Key Details & Customization Tips

  • Delimiter Handling: The script assumes both files use spaces as delimiters (matches your sample data). If you have tabs or other separators, replace split() with split('\t') or the appropriate delimiter.
  • Large Files: If your files are extremely large (gigabytes), modify the reference file loader to read line-by-line instead of loading the whole file at once—this will save memory.
  • Missing IDs: The current script prints a warning for IDs not found in the reference file. You can adjust this to write the original line without coordinates, or log missing IDs to a separate file instead.

内容的提问来源于stack exchange,提问作者Azam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:56:47