如何用Pandas计算CSV中位数?新手遇KeyError求助
Hey there! Let's break down why you're running into that KeyError and get your median calculation working right away.
Why the KeyError Happens
Your line median = census[0].median() is trying to access a column named "0" in your DataFrame—but your CSV almost certainly doesn't have a column with that exact name. In Pandas, the square bracket [] operator looks up columns by their label/name, not their position index. Since there's no column called "0", you get the KeyError.
Step-by-Step Fixes & Correct Methods
First, let's confirm your CSV's structure so you know what you're working with. Add these lines right after loading the CSV to see your column names and sample data:
print("Your CSV columns:", census.columns) print("\nSample data from your CSV:\n", census.head())
Once you know your column names or positions, use one of these approaches:
Calculate median for a specific column (by name, recommended)
If you want the median of a column likepopulationorage, reference it directly by its name:# Replace "your_column_name" with the actual name from census.columns median = census["your_column_name"].median()Calculate median for a column by position
If you want the first column (index 0) regardless of its name, useiloc(integer location) to target columns by their position:# : means all rows, 0 means the first column median = census.iloc[:, 0].median()Calculate median for all numeric columns at once
If you want medians for every numeric column in your CSV, use the DataFrame's built-inmedian()method withnumeric_only=Trueto skip non-numeric columns:all_medians = census.median(numeric_only=True) print("Medians for all numeric columns:\n", all_medians)
Full Working Example
Here's how your code might look after fixing it (using the position-based approach as an example):
import pandas census = pandas.read_csv("census(1).csv") # Verify your data structure first print("CSV Columns:", census.columns) print("\nFirst 5 rows:\n", census.head()) # Calculate median for the first column median = census.iloc[:, 0].median() print("\nMedian of the first column:", median)
If you're still stuck, sharing a few rows of your CSV (including the header row) would help debug further—sometimes numeric columns get incorrectly parsed as strings, which can also cause issues with median calculations.
内容的提问来源于stack exchange,提问作者Alex

