如何高效为带括号的Pandas DataFrame水果名称添加新行?
问题
现有如下Pandas DataFrame:
Fruit Color Size Lemon(Fruit) Green 1 Apple(Fruit) Green 1.1 Banana(Fruit) Yellow 2.5 Banana Black 1
需求:仅为名称带括号的水果生成新行,新行仅保留纯水果名称;同时保留原有无括号的Banana行,期望输出如下:
Fruit Color Size Lemon(Fruit) Green 1 Lemon Green 1 Apple(Fruit) Green 1.1 Apple Green 1.1 Banana(Fruit) Yellow 2.5 Banana Black 1
此前通过循环遍历行实现,但效率极低,求更优方法。
高效解决方案
利用Pandas的矢量化操作替代循环,大幅提升处理效率,具体实现如下:
1. 准备原始数据
import pandas as pd df = pd.DataFrame({ 'Fruit': ['Lemon(Fruit)', 'Apple(Fruit)', 'Banana(Fruit)', 'Banana'], 'Color': ['Green', 'Green', 'Yellow', 'Black'], 'Size': [1, 1.1, 2.5, 1] })
2. 提取带括号的行并生成纯名称行
通过矢量化字符串操作筛选目标行、提取纯名称,构造新的DataFrame:
# 筛选Fruit列包含括号的行 mask = df['Fruit'].str.contains('\(', regex=True) paren_rows = df[mask].copy() # 提取括号前的纯水果名称 paren_rows['Fruit'] = paren_rows['Fruit'].str.split('(', expand=True)[0]
3. 合并数据并调整顺序
将原始数据和新生成的纯名称行合并,再通过分组排序让原行与对应新行相邻:
# 合并两个数据集 combined_df = pd.concat([df, paren_rows], ignore_index=True) # 新增辅助列用于分组排序 combined_df['pure_name'] = combined_df['Fruit'].str.split('(', expand=True)[0] # 先按纯名称排序,再让带括号的行排在纯名称行前面 combined_df = combined_df.sort_values(by=['pure_name', 'Fruit'], ascending=[True, False]) # 删除辅助列并重置索引 final_df = combined_df.drop('pure_name', axis=1).reset_index(drop=True)
运行后得到的final_df就是期望输出:
Fruit Color Size 0 Lemon(Fruit) Green 1.0 1 Lemon Green 1.0 2 Apple(Fruit) Green 1.1 3 Apple Green 1.1 4 Banana(Fruit) Yellow 2.5 5 Banana Black 1.0
效率说明
全程使用Pandas内置的矢量化方法(str.contains、str.split等),这类方法底层基于C实现,比Python逐行循环快数个数量级;合并、排序操作也属于Pandas的批量高效处理逻辑,完全规避了逐行遍历的性能损耗。
内容的提问来源于stack exchange,提问作者Cheburashka
相关产品推荐
相关产品推荐

