如何使用Python实现两个不同长度文件的数据搜索与匹配?
Got it, let's tackle this problem step by step. The core goal here is to match the shared ID between your two files so we can append the latitude and longitude from location.dat to the corresponding entries in input.dat.
Approach Breakdown
First, we'll load the reference data into a dictionary where the location ID is the key—this lets us look up coordinates in constant time (O(1)), which is way faster than scanning the entire reference file every time we need a match. Then, we'll iterate through the dataset file, pull the ID from each line, fetch the matching coordinates, and merge the data into a new output file.
Complete Python Script
Here's a fully functional script with comments to walk you through each part:
def load_reference_data(ref_file_path): location_dict = {} with open(ref_file_path, 'r') as f: # Split the entire file into individual space-separated values raw_data = f.read().split() # Process data in chunks of 3: ID, latitude, longitude for idx in range(0, len(raw_data), 3): loc_id = raw_data[idx] latitude = float(raw_data[idx+1]) longitude = float(raw_data[idx+2]) location_dict[loc_id] = (latitude, longitude) return location_dict def match_and_merge(input_file_path, location_map, output_file='merged_output.dat'): with open(input_file_path, 'r') as infile, open(output_file, 'w') as outfile: for line in infile: cleaned_line = line.strip() if not cleaned_line: continue # Skip empty lines to avoid errors line_parts = cleaned_line.split() input_id = line_parts[0] # Check if the ID exists in our reference data if input_id in location_map: lat, lon = location_map[input_id] # Combine original line data with coordinates merged_entry = ' '.join(line_parts + [str(lat), str(lon)]) outfile.write(f"{merged_entry}\n") else: # Handle missing IDs (adjust this to your needs) print(f"Warning: ID {input_id} not found in reference data—skipping this entry.") # Run the script if __name__ == "__main__": # Load reference coordinates first location_data = load_reference_data('location.dat') # Merge with input dataset and save results match_and_merge('input.dat', location_data) print("Merge complete! Check merged_output.dat for your combined data.")
Key Details & Customization Tips
- Delimiter Handling: The script assumes both files use spaces as delimiters (matches your sample data). If you have tabs or other separators, replace
split()withsplit('\t')or the appropriate delimiter. - Large Files: If your files are extremely large (gigabytes), modify the reference file loader to read line-by-line instead of loading the whole file at once—this will save memory.
- Missing IDs: The current script prints a warning for IDs not found in the reference file. You can adjust this to write the original line without coordinates, or log missing IDs to a separate file instead.
内容的提问来源于stack exchange,提问作者Azam

