基于DataFrame某列值新增列的实现及代码优化需求
问题描述
现有如下结构的DataFrame:
wellname size bin 01 3 bin 02 6 bin 03 2 bin 04 5 john 01 2 john 02 8 john 02 5 pet 05 7 pet 06 10
需要基于wellname列的前缀匹配,新增fieldname列,匹配规则为:
- 前缀为
bin的行,fieldname值为tiger - 前缀为
john的行,fieldname值为leopard - 前缀为
pet的行,fieldname值为bat
用户自行编写的add_col函数未得到预期结果,预期输出如下:
fieldname wellname size tiger bingo 3 leopard jean 5 fish john 2 bat peter 8
请求提供正确的函数实现及等效的列表推导式。
解决方案
1. 正确的函数实现
利用pandas的str.contains()结合正则^锚定前缀,通过条件赋值完成列新增:
import pandas as pd def add_fieldname(df): df_copy = df.copy() # 正则匹配前缀并赋值 df_copy.loc[df_copy['wellname'].str.contains(r'^bin'), 'fieldname'] = 'tiger' df_copy.loc[df_copy['wellname'].str.contains(r'^john'), 'fieldname'] = 'leopard' df_copy.loc[df_copy['wellname'].str.contains(r'^pet'), 'fieldname'] = 'bat' # 为未匹配的行设置默认值(对应预期输出中的fish) df_copy['fieldname'] = df_copy['fieldname'].fillna('fish') return df_copy # 测试代码 data = { 'wellname': ['bin 01', 'john 01', 'pet 05', 'test 01'], 'size': [3, 2, 7, 4] } df = pd.DataFrame(data) print(add_fieldname(df))
2. 等效的列表推导式
遍历wellname列逐个判断前缀,生成fieldname列表后赋值给DataFrame:
import pandas as pd data = { 'wellname': ['bin 01', 'john 01', 'pet 05', 'test 01'], 'size': [3, 2, 7, 4] } df = pd.DataFrame(data) # 基于startswith的列表推导式(简洁高效) df['fieldname'] = [ 'tiger' if well.startswith('bin') else 'leopard' if well.startswith('john') else 'bat' if well.startswith('pet') else 'fish' for well in df['wellname'] ] # 若必须使用str.contains正则匹配,可改为: # df['fieldname'] = [ # 'tiger' if pd.Series([well]).str.contains(r'^bin').iloc[0] # else 'leopard' if pd.Series([well]).str.contains(r'^john').iloc[0] # else 'bat' if pd.Series([well]).str.contains(r'^pet').iloc[0] # else 'fish' # for well in df['wellname'] # ] print(df)
内容的提问来源于stack exchange,提问作者Chuks
相关产品推荐
相关产品推荐

