如何在Pandas透视表中显示分类标签而非数值编码
Got it, this is a super common pain point when working with encoded categorical datasets—luckily, Pandas has straightforward ways to swap those numeric codes for human-readable labels before (or even while) building your pivot table. Let’s walk through this step by step:
First, translate your codebook into a Python dictionary (or nested dictionary if you have multiple encoded columns). For example:
# Nested dict for multiple columns encoding_mapping = { "gender": {1: "Male", 2: "Female"}, "marital_status": {0: "Single", 1: "Married", 2: "Divorced"}, "education_level": {1: "High School", 2: "Bachelor's", 3: "Master's", 4: "PhD"} } # Or a single dict for one column if that's all you need gender_mapping = {1: "Male", 2: "Female"}
You have two flexible options here depending on whether you want to modify your original DataFrame or not:
Option A: Modify the Original DataFrame (Permanent Change)
Use df.replace() to swap codes for labels across all relevant columns in one go:
# Apply mapping to all columns in the encoding dict df.replace(encoding_mapping, inplace=True)
Or target a single column with map():
# Update just the gender column df["gender"] = df["gender"].map(gender_mapping)
Important note: map() will turn any unmapped numeric values into NaN. If you want to keep those original values instead, use fillna():
df["gender"] = df["gender"].map(gender_mapping).fillna(df["gender"])
Option B: Keep Original DataFrame Intact (Temporary Mapping)
If you don’t want to alter your raw data, you can apply the mapping directly within the pivot_table() call:
pivot = df.pivot_table( values="monthly_salary", # Your numeric value column index=df["gender"].map(gender_mapping), # Map codes to labels for rows columns=df["education_level"].map(encoding_mapping["education_level"]), # Map for columns aggfunc="mean" # Your desired aggregation (sum, count, etc.) )
Once your codes are replaced with labels, creating the pivot table works exactly like normal—except now the rows/columns will show human-readable labels instead of numbers. For example:
# If you modified the original DataFrame final_pivot = df.pivot_table( values="monthly_salary", index="gender", columns="education_level", aggfunc=["mean", "count"], margins=True )
This will give you a pivot table with "Male"/"Female" instead of 1/2, and education labels instead of numeric codes—exactly what you’re looking for!
内容的提问来源于stack exchange,提问作者B.Poe

