Python实现CSV连接统计(无Pandas/Numpy)及多场景适配问询
Got it, let's walk through this problem step by step. I'll cover the solution using only basic Python data structures, plus answers to your specific questions about column lists/loops and handling useless columns:
Solution Overview
Your goal is to take a CSV with columns out_gate, useless_column, in_gate, num_connect, then calculate the total num_connect for each unique out_gate→in_gate pair (e.g., c→b sums to 8, c→a sums to 9). You're restricted to using lists, tuples, dictionaries, or collections.defaultdict (no Pandas/Numpy), and need it to work for 10-40 unique gates.
Key Answers to Your Questions
- Can this be done with column lists + loops? Absolutely yes — this is actually the most straightforward way to handle CSV data without libraries. By working with column indexes (from a header list) and looping through each row, you can precisely extract only the data you need.
- How to handle the useless column? Simply ignore it entirely. When parsing each row, only extract values from the columns you care about (
out_gate,in_gate,num_connect) and skip theuseless_columndata. If the column order might vary, you can first map column names to their indexes using the header row, so you never accidentally grab the wrong data.
Complete Implementation Code
Here's a fully functional script that meets all your requirements:
from collections import defaultdict def calculate_gate_connection_totals(input_csv_path, output_csv_path): # Initialize a defaultdict to track total connections per (out_gate, in_gate) pair connection_totals = defaultdict(int) with open(input_csv_path, 'r', encoding='utf-8') as input_file: # Read and parse the header to find column indexes header = input_file.readline().strip().split(',') out_gate_idx = header.index('out_gate') in_gate_idx = header.index('in_gate') num_connect_idx = header.index('num_connect') # Loop through each data row for line in input_file: line = line.strip() if not line: continue # Skip empty rows row_data = line.split(',') # Extract only the values we need (ignore useless_column entirely) out_gate = row_data[out_gate_idx].strip() in_gate = row_data[in_gate_idx].strip() num_connect = int(row_data[num_connect_idx].strip()) # Add to the total for this gate pair connection_totals[(out_gate, in_gate)] += num_connect # Write the results to the output CSV with open(output_csv_path, 'w', encoding='utf-8') as output_file: # Write header output_file.write('out_gate,in_gate,total_connect\n') # Write each gate pair and its total for (out_gate, in_gate), total in connection_totals.items(): output_file.write(f"{out_gate},{in_gate},{total}\n") # Example usage (replace with your actual file paths) calculate_gate_connection_totals('input.csv', 'output.csv')
Breakdown of the Code
- Handling Useless Columns: By using the header row to find indexes for only the columns we need, we completely bypass the
useless_column— we never read or process its data, so it doesn't affect our results. - Column Lists + Loops: The header is converted to a list, and we use
index()to find the position of our target columns. Then we loop through every row, using those indexes to extract the relevant values. This approach is flexible even if column orders change later. - Efficient Counting:
defaultdict(int)automatically initializes any new(out_gate, in_gate)pair to 0, so we don't have to check if the pair exists before adding to its total. This is perfect for dynamic gate counts (10-40 or more). - Scalability: Since we're dynamically tracking gate pairs as they appear in the input, the script works seamlessly regardless of how many unique gates you have (10, 40, or even more).
Example Input & Output
Input CSV (input.csv)
out_gate,useless_column,in_gate,num_connect c,random_text,b,3 c,another_junk,a,5 c,more_stuff,b,5 c,extra,a,4 a,foo,c,2
Output CSV (output.csv)
out_gate,in_gate,total_connect c,b,8 c,a,9 a,c,2
内容的提问来源于stack exchange,提问作者kts
相关产品推荐
相关产品推荐

