如何实现多结构文本文件统一格式合并?已完成二维列表存储
Got it, let's break this down. You've already nailed the first big steps—loading all your files and splitting lines into a 2D list with strings/nulls. Now we just need to map that messy, mixed-structure data into a standardized format so you can merge everything seamlessly.
Step 1: Define Your Target Unified Structure
First, you need to decide exactly what fields every entry in your merged file should have. Looking at your sample data:
- Some rows look like:
809187.49 226885.80 26934(3 elements: coord X, coord Y, identifier) - Others look like:
2 5509514.58 6558911.86 0.00 80T(5 elements: identifier, coord X, coord Y, extra value, data type)
Let's pick a clear target schema—for example:[coord_x, coord_y, identifier, extra_value, data_type]
(Adjust this to match your actual needs, of course.)
Step 2: Write Mapping Logic to Standardize Rows
We'll loop through your raw 2D list, check each row's structure, and map its elements to the target schema. Missing fields get filled with None (or an empty string, if you prefer). Here's a Python example since it's great for list manipulation:
# Your raw 2D list (from your file processing step) raw_data = [ ["809187.49", "226885.80", "26934"], ["809183.14", "226877.21", "26937a"], ["2", "5509514.58", "6558911.86", "0.00", "80T"], ["3", "5509515.55", "6558913.48", "0.00", "80T"], ["4", "5509516.35", "6558914.56", "0.00", "80T"] ] # Define your target fields (adjust based on your needs) target_fields = ["coord_x", "coord_y", "identifier", "extra_value", "data_type"] # Initialize the unified data array unified_list = [] for row in raw_data: # Start with a row filled with None for all target fields unified_row = [None] * len(target_fields) # Map based on row length (adjust conditions for your actual structures) if len(row) == 3: # Handle 3-element rows: [coord_x, coord_y, identifier] unified_row[0] = row[0] unified_row[1] = row[1] unified_row[2] = row[2] elif len(row) == 5: # Handle 5-element rows: [identifier, coord_x, coord_y, extra_value, data_type] unified_row[2] = row[0] unified_row[0] = row[1] unified_row[1] = row[2] unified_row[3] = row[3] unified_row[4] = row[4] # Add more elif blocks here if you have other row structures unified_list.append(unified_row)
Step 3: Merge & Export the Unified Data
Once you have the standardized 2D list, writing it to a merged file is straightforward. Let's save it as a tab-separated text file (matching your original splitting logic):
# Write the unified data to a new file with open("merged_unified.txt", "w") as output_file: # Optional: Write a header row with field names output_file.write("\t".join(target_fields) + "\n") # Write each standardized row for row in unified_list: # Convert None values to empty strings for clean output cleaned_row = [str(val) if val is not None else "" for val in row] output_file.write("\t".join(cleaned_row) + "\n")
Key Tips for Edge Cases
- Handle null values: If your raw list has
null(not just missing elements), add checks likeif row[0] is not Nonebefore mapping. - Structure detection by content: If rows have the same length but different field orders, use content checks instead of length. For example:
if row[-1].endswith("T"): # This is a 5-element row with a data_type ending in T # Map accordingly - Test with small data: Run your logic on a subset of your files first to make sure every row type gets mapped correctly before processing everything.
内容的提问来源于stack exchange,提问作者DantePh

