求基于首列合并行的Python 2.7脚本(不使用pandas)
Hey there, I get how frustrating it is when you've scoured through tons of Q&As and still can't crack a problem. Let's fix this together—since you're stuck with Python 2.7 and can't use pandas, we'll lean on Python's built-in csv module, which is more than capable of handling your large, headerless CSV with duplicate first-column entries.
核心思路
The key here is to use a dictionary to group rows by their first column value. The dictionary's keys will be the unique values from the first column, and the values will store all the corresponding data from the rest of the rows. This way, we can process the file line by line (super memory-friendly for large CSVs) and then combine all the grouped data into merged rows.
方案1:简单追加所有后续列
If your goal is to just append all the extra columns from duplicate first-column rows into a single row (e.g., rows ["A", "1", "2"] and ["A", "3", "4"] become ["A", "1", "2", "3", "4"]), use this script:
import csv def merge_duplicate_rows(input_path, output_path): # Dictionary to hold merged data: key = first column value, value = list of other columns merged_dict = {} # Read the input CSV line by line with open(input_path, 'rb') as infile: csv_reader = csv.reader(infile) for row in csv_reader: if not row: # Skip empty rows to avoid errors continue first_col_val = row[0] remaining_cols = row[1:] # Update the dictionary: append columns if key exists, else create new entry if first_col_val in merged_dict: merged_dict[first_col_val].extend(remaining_cols) else: merged_dict[first_col_val] = remaining_cols # Write the merged data to output CSV with open(output_path, 'wb') as outfile: csv_writer = csv.writer(outfile) for key, combined_cols in merged_dict.items(): merged_row = [key] + combined_cols csv_writer.writerow(merged_row) # Example usage if __name__ == '__main__': merge_duplicate_rows('your_input.csv', 'merged_output.csv')
方案2:按对应列合并(带分隔符)
If you need to merge data by column position (e.g., rows ["A", "x", "y"] and ["A", "m", "n"] become ["A", "x;m", "y;n"]), use this version instead. It groups values by their column index and joins them with a delimiter of your choice:
import csv def merge_duplicate_by_column(input_path, output_path, join_delimiter=';'): merged_dict = {} with open(input_path, 'rb') as infile: csv_reader = csv.reader(infile) for row in csv_reader: if not row: continue first_col_val = row[0] remaining_cols = row[1:] # Initialize entry if key doesn't exist: list of empty lists for each column if first_col_val not in merged_dict: merged_dict[first_col_val] = [[] for _ in range(len(remaining_cols))] # Append each value to its corresponding column list for col_index, value in enumerate(remaining_cols): merged_dict[first_col_val][col_index].append(value) # Write merged data with open(output_path, 'wb') as outfile: csv_writer = csv.writer(outfile) for key, column_groups in merged_dict.items(): # Join each column's values with the delimiter merged_columns = [join_delimiter.join(group) for group in column_groups] merged_row = [key] + merged_columns csv_writer.writerow(merged_row) # Example usage if __name__ == '__main__': merge_duplicate_by_column('your_input.csv', 'col_merged_output.csv')
关键注意事项
- Python 2.7 File Modes: We use
'rb'and'wb'for reading/writing because thecsvmodule in Python 2.7 requires binary mode to handle line endings correctly. - Large Files: Both scripts process the CSV line by line, so they won't hog memory even with huge files—perfect for your use case.
- Custom Delimiters: If your CSV uses something other than commas (like tabs), add
delimiter='\t'to thecsv.readerandcsv.writercalls (e.g.,csv.reader(infile, delimiter='\t')).
内容的提问来源于stack exchange,提问作者Pbree

