如何结合列表多行合并与坐标系编码实现按列生成数据模式?
Hey, it looks like you've already converted your multi-file data into a clean list of coordinate tuples—great start! Generating column-based patterns from this structure is totally manageable, and I’ll walk you through practical implementation steps with Python examples (since it’s perfect for this kind of data wrangling):
First, let’s break down your example: each tuple (col_label, row_position) maps a column identifier to a row where data exists. Your goal is to group these rows by their column labels, then process those groups based on your desired pattern size.
The first core task is to cluster all row positions under their corresponding columns. We can use Python’s collections.defaultdict for this—it’s clean and efficient:
from collections import defaultdict # Your sample coordinate data coords = [('0', 1), ('0', 2), ('0', 3), ('1', 4), ('2', 5), ('3', 6), ('3', 7), ('3', 8), ('2', 9), ('1', 10), ('0', 11), ('-1', 12), ('-2', 13), ('-3', 14), ('-3', 15), ('-3', 16), ('-2', 17), ('-1', 18), ('0', 19), ('0', 20), ('0', 21)] # Group rows by their column label column_groups = defaultdict(list) for col, row in coords: column_groups[col].append(int(row)) # Convert row to integer for easier sorting # Sort columns numerically (so -3 comes before -2, etc.) sorted_columns = sorted(column_groups.keys(), key=int)
Next, you’ll need to adjust these groups to fit your target pattern dimensions. Here are two common scenarios you might encounter:
Scenario 1: Generate a Full Matrix Pattern
If you need to create a binary matrix where each column is a vector (1 = data exists at that row, 0 = no data), first define the full row range of your pattern:
# Get the total row span of your data all_rows = [row for col_rows in column_groups.values() for row in col_rows] min_row = min(all_rows) max_row = max(all_rows) total_rows = max_row - min_row + 1 # Build a pattern vector for each column column_patterns = {} for col in sorted_columns: row_list = column_groups[col] # Initialize a vector of 0s for all rows pattern_vector = [0] * total_rows # Mark positions where data exists for row in row_list: relative_index = row - min_row # Convert to 0-based index pattern_vector[relative_index] = 1 column_patterns[col] = pattern_vector # Example: Print the pattern for column '0' print("Pattern for column '0':", column_patterns['0']) # Output: [1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1]
Scenario 2: Extract Compact Column Blocks
If your pattern is defined by continuous row blocks (like column '3' in your example spans rows 6-8), you can extract key metadata for each column:
column_block_info = {} for col in sorted_columns: sorted_rows = sorted(column_groups[col]) # Check if rows form a continuous block (optional) is_continuous = all(sorted_rows[i+1] == sorted_rows[i] + 1 for i in range(len(sorted_rows)-1)) # Store block details column_block_info[col] = { 'start_row': sorted_rows[0], 'end_row': sorted_rows[-1], 'block_length': len(sorted_rows), 'is_continuous': is_continuous } # Example: Print details for column '3' print("Block info for column '3':", column_block_info['3']) # Output: {'start_row': 6, 'end_row': 8, 'block_length': 3, 'is_continuous': True}
If you’re reading data from multiple files, it’s more efficient to group columns as you read each file (instead of first collecting all coordinates):
def parse_file_columns(file_path): """Read a file and return row positions grouped by column.""" file_col_data = defaultdict(list) with open(file_path, 'r') as f: # Adjust this parsing logic to match your actual file format for line in f: col_label, row_str = line.strip().split(',') file_col_data[col_label].append(int(row_str)) return file_col_data # Merge data from multiple files all_col_data = defaultdict(list) input_files = ['data1.txt', 'data2.txt', 'data3.txt'] # Replace with your files for file in input_files: file_data = parse_file_columns(file) for col, rows in file_data.items(): all_col_data[col].extend(rows) # Now use all_col_data just like column_groups from Step 2
If you have a specific pattern format in mind (like ASCII art, a JSON output, or a custom binary format), feel free to share more details, and we can tweak the implementation to match!
内容的提问来源于stack exchange,提问作者Nauman Shahid

