将分类行转为列标题(无聚合):寻求Pandas更优实现方案
The issue with using pivot() directly here is that it relies on the existing index to align values. Since each original index row only has one value for 'Art' (either 'blue' or 'red'), the other column ends up with NaNs because there's no matching value for that index.
A cleaner, more efficient approach than your apply() workaround is to first create a sequential identifier for each entry within its 'Art' group, then pivot using that identifier as the index. Here's how:
Step 1: Add a group-specific ID column
First, we generate an ID that increments for each entry in the same 'Art' category. This ensures we have a common index to align 'blue' and 'red' entries properly:
import pandas as pd data = pd.DataFrame({ 'Art':['blue', 'red', 'blue', 'red', 'blue', 'red', 'blue', 'red'], 'Description':['Some text 1', 'Some text 2', 'Some text 3', 'Some text 4', 'Some text 5', 'Some text 6', 'Some text 7', 'Some text 8'] }) # Add sequential ID per 'Art' group data['group_id'] = data.groupby('Art').cumcount()
Step 2: Pivot using the group ID
Now we can pivot using the group_id as the index, which will align the 'blue' and 'red' entries without NaNs:
pivoted_data = data.pivot( index='group_id', columns='Art', values='Description' ).reset_index(drop=True) print(pivoted_data)
Output:
Art blue red 0 Some text 1 Some text 2 1 Some text 3 Some text 4 2 Some text 5 Some text 6 3 Some text 7 Some text 8
Why this works better than your workaround:
- Vectorized operation: This method uses pandas' built-in groupby and pivot functions, which are optimized for speed (unlike
apply()which loops through each column, making it slower for large datasets). - Readability: The logic is explicit—we're creating a clear alignment key before pivoting, which makes the code easier to understand and maintain.
Alternative (using unstack):
You can also achieve the same result with groupby and unstack, though it's slightly more concise:
pivoted_data = data.groupby( ['Art', data.groupby('Art').cumcount()] )['Description'].unstack(0).reset_index(drop=True)
Either way, both methods avoid NaNs and are more efficient than your original apply() approach.
内容的提问来源于stack exchange,提问作者AMD

