使用pd.crosstab构建混淆矩阵触发AssertionError,疑names参数问题
pd.crosstab Hey there! Let's break down why you're hitting that AssertionError and how to fix it. I'll start by clarifying what the names parameter does, then walk through common issues and actionable fixes.
First: What's the names parameter in pd.crosstab?
The names argument lets you set custom labels for the dimensions of your cross-tab (and by extension, your confusion matrix). If you're passing two arrays (true labels Q_test and predicted labels prediction), names needs to be a list of exactly two strings—one for the row dimension (true values) and one for the column dimension (predictions).
If you don't need custom names, you can skip this parameter entirely—pandas will use default labels based on your input arrays. The AssertionError often pops up here when the length of names doesn't match the number of input arrays you're passing to pd.crosstab.
Common Causes & Fixes for the AssertionError
1. Double-check that Q_test and prediction really have the same length
Even if you think they match, it's worth verifying explicitly—your select_neighbor_all function might be dropping or adding a sample by accident. Run this quick check:
print(f"Q_test length: {len(Q_test)}, Prediction length: {len(prediction)}") # If they're pandas Series, also check index alignment if hasattr(Q_test, 'index') and hasattr(prediction, 'index'): print(f"Indices match: {Q_test.index.equals(prediction.index)}")
If lengths don't match, you'll need to debug select_neighbor_all to ensure it returns one prediction per test sample.
2. Align the names parameter length with your input arrays
If you're using names, make sure its length matches how many arrays you're passing to pd.crosstab. For a standard confusion matrix (two arrays), it should look like this:
import pandas as pd # Correct usage with names confusion_matrix = pd.crosstab( Q_test, prediction, names=['True Class', 'Predicted Class'] # Exactly 2 names for 2 input arrays )
Or, if you prefer more explicit control, use rownames and colnames instead (this avoids confusion with names):
confusion_matrix = pd.crosstab( Q_test, prediction, rownames=['True Class'], # Label for true values (rows) colnames=['Predicted Class'] # Label for predictions (columns) )
If you pass a names list with 1 or 3+ elements here, pandas will throw an AssertionError because it expects the count to match the number of input arrays.
3. Check the exact error message
Take a look at the full AssertionError traceback—pandas usually tells you exactly what's wrong. For example:
- If it says
arrays must all be same length, yourQ_testandpredictionhave mismatched lengths. - If it says
names must match number of input arrays, yournameslist length is off.
Example Working Code
Here's a complete snippet using your binary 'Low'/'High' labels:
import pandas as pd # Sample data matching your setup Q_test = ['Low', 'High', 'Low', 'Low', 'High'] prediction = ['Low', 'High', 'High', 'Low', 'High'] # Verify lengths first assert len(Q_test) == len(prediction), "Mismatched lengths between true and predicted labels!" # Build confusion matrix confusion_matrix = pd.crosstab( Q_test, prediction, rownames=['True Class'], colnames=['Predicted Class'] ) print(confusion_matrix)
This will output a clean confusion matrix showing counts of True Positives, True Negatives, etc.
内容的提问来源于stack exchange,提问作者k3513662

