如何在Python中从DataFrame两列生成规整的二进制矩阵
Ah, I see the issue with your current code—you're using the raw ID and P values (including duplicates) as your index and columns, which is why you're getting unmerged, unsorted rows and columns. The fix is to use unique, sorted values for both axes, then mark the intersections where an ID has a corresponding P value.
Here's the simplest way to get the clean, ordered binary matrix you want using pandas' built-in crosstab function, which handles unique/sorted axes automatically:
import pandas as pd # Your sample data df = pd.DataFrame({ 'ID': [2, 1, 3, 1, 1, 2, 3], 'P': [1, 2, 2, 3, 4, 5, 5] }) # Generate the binary matrix binary_matrix = pd.crosstab(df['ID'], df['P']).astype(int) print(binary_matrix)
This outputs exactly the structured matrix you're looking for:
P 1 2 3 4 5 ID 1 0 1 1 1 0 2 1 0 0 0 1 3 0 1 0 0 1
How this works:
pd.crosstabautomatically uses the unique, sorted values ofIDas rows andPas columns.- It counts occurrences of each ID-P pair, and
astype(int)converts counts (which are 1 for existing pairs, 0 otherwise) to binary values.
If you need to include all P values up to a maximum (like your example shows columns 1-7):
If you want columns to span every number from the minimum to maximum P value (even if some aren't present in your data), you can reindex the columns:
# Let's say you want columns up to 7 max_p = 7 binary_matrix = binary_matrix.reindex(columns=range(1, max_p + 1), fill_value=0) print(binary_matrix)
This adds missing columns (6 and 7) filled with 0s:
P 1 2 3 4 5 6 7 ID 1 0 1 1 1 0 0 0 2 1 0 0 0 1 0 0 3 0 1 0 0 1 0 0
Manual approach (if you prefer not to use crosstab):
If you want to build the matrix manually (similar to your original code but fixed), you can do this:
# Get unique sorted IDs and P values unique_ids = sorted(df['ID'].unique()) unique_p = sorted(df['P'].unique()) # Create a list of lists for the binary data binary_data = [] for id_val in unique_ids: # Check if each P exists for the current ID row = [1 if p in df[df['ID'] == id_val]['P'].values else 0 for p in unique_p] binary_data.append(row) # Create the DataFrame binary_matrix = pd.DataFrame(binary_data, index=unique_ids, columns=unique_p)
This gives you the same result as the crosstab method, though crosstab is more efficient for larger datasets.
内容的提问来源于stack exchange,提问作者C_psy

