使用Bash统计CSV文件中各城市学生数及指定城市统计方案
Solution for Student City Counting Bash Script
Let's fix this up for you. First, quick note on your original sort command failure: the semicolon ; is a special shell character that tells bash to run a new command, so you need to quote it or escape it. The correct command would be:
sort -k3 -t';' students.csv
Now, here's a bash script that handles both of your requested features—grouped city counts and targeted single-city counts via pipe or input redirect:
#!/bin/bash # Check if we're getting input from a pipe or file input if [ -t 0 ]; then echo "Error: Please provide CSV data via pipe or input redirect." echo "Usage:" echo " # Count all students per city" echo " cat students.csv | $0" echo " # Count students in a specific city" echo " cat students.csv | $0 NYC" exit 1 fi # Use awk to process the CSV (semicolon-separated) if [ $# -eq 1 ]; then # Single city count: filter rows where column 3 matches the input city awk -F';' -v city="$1" '$3 == city {count++} END {print count + 0}' else # Grouped count: aggregate by city and print totals awk -F';' '{count[$3]++} END {for (city in count) print city ": " count[city]}' fi
How it works:
- Input Check: The
[ -t 0 ]check makes sure we have input coming in (either via pipe likecat students.csv | ./script.shor redirect like./script.sh < students.csv). If not, it shows clear usage instructions. - Single City Mode: When you pass a city name as an argument (e.g.,
./script.sh NYC), awk uses that value to filter rows where the 3rd column matches, then counts and prints the total. The+ 0ensures we get0instead of a blank output if the city doesn't exist in the data. - Grouped Mode: Without arguments, awk creates an associative array where keys are city names and values are the number of students in each city. At the end, it loops through the array to print each city and its corresponding count.
Testing with your sample data:
Sample students.csv content:
michael_s;jpa;NYC john_d;chn;TXS jim_h;usa;POP
Grouped count:
cat students.csv | ./script.shOutput:
NYC: 1 TXS: 1 POP: 1Single city count:
cat students.csv | ./script.sh NYCOutput:
1
Optional: Handle headers
If your CSV has a header row (e.g., Name;Country;City), just add NR>1 to the awk commands to skip the first line:
- For single city:
awk -F';' -v city="$1" 'NR>1 && $3 == city {count++} END {print count + 0}' - For grouped count:
awk -F';' 'NR>1 {count[$3]++} END {for (city in count) print city ": " count[city]}'
内容的提问来源于stack exchange,提问作者not_a_student
相关产品推荐
相关产品推荐

