使用Lambda为DataFrame批量填充多列时遇ValueError问题求助
问题
我有如下DataFrame:
TABLE_NAME jace building equipment 0 R338_1_FAHU2_SUPAIR_TEMP NaN NaN NaN 1 R1001_1_R1005_1_FAHU_1_CO2_SEN2 NaN NaN NaN
我编写了一个子函数fill,可解析TABLE_NAME列的内容并返回三个值:
import re def fill(tablename='R338_1_FAHU2_SUPAIR_TEMP'): jace,building,equipment=re.findall('(^[RP].*?_[1-9])_*(.*?)_(F.*)',tablename)[0] if not len(building): building=re.findall('(.*)_',jace)[0] return jace,building,equipment
该函数对第一行返回结果为:
('R338_1', 'R338', 'FAHU2_SUPAIR_TEMP')
我希望将这些值填充到DataFrame的jace、building、equipment列中,尝试了以下代码:
df[['jace','building','equipment']]=df['TABLE_NAME'].apply(lambda x: (fill(x)))
但触发错误:
ValueError: Must have equal len keys and value when setting with an iterable
我也曾在apply中添加axis=1,但似乎与Lambda存在冲突。请问如何实现需求?(我可以使用fill(x)[0]、fill(x)[1]这类硬编码方式,但希望避免)
解决方案
可以通过以下几种方式实现,避免硬编码索引:
方法1:使用apply+pd.Series
在lambda中将fill的返回值转为pd.Series,apply会生成包含三列的DataFrame,直接赋值即可:
import pandas as pd df[['jace', 'building', 'equipment']] = df['TABLE_NAME'].apply(lambda x: pd.Series(fill(x)))
方法2:使用apply的expand参数(pandas 0.23+支持)
直接在apply中设置expand=True,会将返回的元组自动展开为多列:
df[['jace', 'building', 'equipment']] = df['TABLE_NAME'].apply(fill, expand=True)
方法3:修改fill函数返回pd.Series
调整fill函数的返回值为Series,apply后直接得到可匹配目标列的结构:
def fill(tablename='R338_1_FAHU2_SUPAIR_TEMP'): jace,building,equipment=re.findall('(^[RP].*?_[1-9])_*(.*?)_(F.*)',tablename)[0] if not len(building): building=re.findall('(.*)_',jace)[0] return pd.Series([jace, building, equipment], index=['jace', 'building', 'equipment']) df[['jace', 'building', 'equipment']] = df['TABLE_NAME'].apply(fill)
以上方法都能直接将fill返回的三个值对应填充到目标列中,无需硬编码索引。
内容的提问来源于stack exchange,提问作者Moataz Imad
相关产品推荐
相关产品推荐

