You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将分类行转为列标题(无聚合):寻求Pandas更优实现方案

Solution: Convert Categorical Rows to Columns Without Aggregation

The issue with using pivot() directly here is that it relies on the existing index to align values. Since each original index row only has one value for 'Art' (either 'blue' or 'red'), the other column ends up with NaNs because there's no matching value for that index.

A cleaner, more efficient approach than your apply() workaround is to first create a sequential identifier for each entry within its 'Art' group, then pivot using that identifier as the index. Here's how:

Step 1: Add a group-specific ID column

First, we generate an ID that increments for each entry in the same 'Art' category. This ensures we have a common index to align 'blue' and 'red' entries properly:

import pandas as pd

data = pd.DataFrame({
    'Art':['blue', 'red', 'blue', 'red', 'blue', 'red', 'blue', 'red'],
    'Description':['Some text 1', 'Some text 2', 'Some text 3', 'Some text 4', 'Some text 5', 'Some text 6', 'Some text 7', 'Some text 8']
})

# Add sequential ID per 'Art' group
data['group_id'] = data.groupby('Art').cumcount()

Step 2: Pivot using the group ID

Now we can pivot using the group_id as the index, which will align the 'blue' and 'red' entries without NaNs:

pivoted_data = data.pivot(
    index='group_id',
    columns='Art',
    values='Description'
).reset_index(drop=True)

print(pivoted_data)

Output:

Art          blue         red
0      Some text 1  Some text 2
1      Some text 3  Some text 4
2      Some text 5  Some text 6
3      Some text 7  Some text 8

Why this works better than your workaround:

  • Vectorized operation: This method uses pandas' built-in groupby and pivot functions, which are optimized for speed (unlike apply() which loops through each column, making it slower for large datasets).
  • Readability: The logic is explicit—we're creating a clear alignment key before pivoting, which makes the code easier to understand and maintain.

Alternative (using unstack):

You can also achieve the same result with groupby and unstack, though it's slightly more concise:

pivoted_data = data.groupby(
    ['Art', data.groupby('Art').cumcount()]
)['Description'].unstack(0).reset_index(drop=True)

Either way, both methods avoid NaNs and are more efficient than your original apply() approach.

内容的提问来源于stack exchange,提问作者AMD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:47:51