如何用awk按列匹配结果汇总对应列的数值?
Hey there! I get that it can be frustrating when something seems simple but you can't crack it right away. Let's fix that—grouping and summing your values by the first column is actually pretty straightforward once you use the right tools. Here are a few solutions depending on what platform you're working with:
Since your data is already sorted by the first column, using the built-in "Subtotal" feature is the easiest way:
- Select your entire data range (including both columns)
- In Excel: Go to the Data tab → click Subtotal
- In Google Sheets: Go to the Data tab → click Create a filter, then go to Data → Subtotal
- In the settings window:
- Choose your first column as the "At each change in" field
- Select "Sum" as the "Use function" option
- Check the box for your numeric column (the second one)
- Click OK, and you'll get grouped sums. You can then copy the final summary rows to get your desired output:
aaa,42、bbb,333、ccc,101
If you prefer coding, Python has two simple approaches—with or without the pandas library:
Using pandas (most concise)
Pandas is perfect for data aggregation tasks:
import pandas as pd # Define your data (replace this with your actual data source if needed) data = [["aaa", 3], ["aaa", 21], ["aaa", 18], ["bbb", 333], ["ccc", 34], ["ccc", 67]] df = pd.DataFrame(data, columns=["key", "value"]) # Group by the first column and sum the values summary = df.groupby("key")["value"].sum().reset_index() # Print or export the result print(summary) # To get the exact comma-separated format: for _, row in summary.iterrows(): print(f"{row['key']},{row['value']}")
Using basic Python (no external libraries)
If you don't want to install pandas, a simple dictionary will do:
data = [["aaa", 3], ["aaa", 21], ["aaa", 18], ["bbb", 333], ["ccc", 34], ["ccc", 67]] sum_tracker = {} # Iterate through each row and accumulate sums for key, number in data: if key in sum_tracker: sum_tracker[key] += number else: sum_tracker[key] = number # Convert to your desired output format for key, total in sum_tracker.items(): print(f"{key},{total}")
If you're working with a text file (e.g., data.txt where each line is like aaa,3), use awk in your terminal—it's fast and requires no setup:
awk -F ',' '{sum[$1] += $2} END {for (k in sum) print k "," sum[k]}' data.txt
This command will automatically group lines by the first column, sum the numbers, and output exactly what you need.
内容的提问来源于stack exchange,提问作者RobertG

