You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现CSV连接统计(无Pandas/Numpy)及多场景适配问询

Got it, let's walk through this problem step by step. I'll cover the solution using only basic Python data structures, plus answers to your specific questions about column lists/loops and handling useless columns:

Solution Overview

Your goal is to take a CSV with columns out_gate, useless_column, in_gate, num_connect, then calculate the total num_connect for each unique out_gate→in_gate pair (e.g., c→b sums to 8, c→a sums to 9). You're restricted to using lists, tuples, dictionaries, or collections.defaultdict (no Pandas/Numpy), and need it to work for 10-40 unique gates.

Key Answers to Your Questions
  • Can this be done with column lists + loops? Absolutely yes — this is actually the most straightforward way to handle CSV data without libraries. By working with column indexes (from a header list) and looping through each row, you can precisely extract only the data you need.
  • How to handle the useless column? Simply ignore it entirely. When parsing each row, only extract values from the columns you care about (out_gate, in_gate, num_connect) and skip the useless_column data. If the column order might vary, you can first map column names to their indexes using the header row, so you never accidentally grab the wrong data.
Complete Implementation Code

Here's a fully functional script that meets all your requirements:

from collections import defaultdict

def calculate_gate_connection_totals(input_csv_path, output_csv_path):
    # Initialize a defaultdict to track total connections per (out_gate, in_gate) pair
    connection_totals = defaultdict(int)
    
    with open(input_csv_path, 'r', encoding='utf-8') as input_file:
        # Read and parse the header to find column indexes
        header = input_file.readline().strip().split(',')
        out_gate_idx = header.index('out_gate')
        in_gate_idx = header.index('in_gate')
        num_connect_idx = header.index('num_connect')
        
        # Loop through each data row
        for line in input_file:
            line = line.strip()
            if not line:
                continue  # Skip empty rows
            
            row_data = line.split(',')
            # Extract only the values we need (ignore useless_column entirely)
            out_gate = row_data[out_gate_idx].strip()
            in_gate = row_data[in_gate_idx].strip()
            num_connect = int(row_data[num_connect_idx].strip())
            
            # Add to the total for this gate pair
            connection_totals[(out_gate, in_gate)] += num_connect
    
    # Write the results to the output CSV
    with open(output_csv_path, 'w', encoding='utf-8') as output_file:
        # Write header
        output_file.write('out_gate,in_gate,total_connect\n')
        # Write each gate pair and its total
        for (out_gate, in_gate), total in connection_totals.items():
            output_file.write(f"{out_gate},{in_gate},{total}\n")

# Example usage (replace with your actual file paths)
calculate_gate_connection_totals('input.csv', 'output.csv')
Breakdown of the Code
  1. Handling Useless Columns: By using the header row to find indexes for only the columns we need, we completely bypass the useless_column — we never read or process its data, so it doesn't affect our results.
  2. Column Lists + Loops: The header is converted to a list, and we use index() to find the position of our target columns. Then we loop through every row, using those indexes to extract the relevant values. This approach is flexible even if column orders change later.
  3. Efficient Counting: defaultdict(int) automatically initializes any new (out_gate, in_gate) pair to 0, so we don't have to check if the pair exists before adding to its total. This is perfect for dynamic gate counts (10-40 or more).
  4. Scalability: Since we're dynamically tracking gate pairs as they appear in the input, the script works seamlessly regardless of how many unique gates you have (10, 40, or even more).
Example Input & Output

Input CSV (input.csv)

out_gate,useless_column,in_gate,num_connect
c,random_text,b,3
c,another_junk,a,5
c,more_stuff,b,5
c,extra,a,4
a,foo,c,2

Output CSV (output.csv)

out_gate,in_gate,total_connect
c,b,8
c,a,9
a,c,2

内容的提问来源于stack exchange,提问作者kts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:11:57