Pandas中用Series筛选DataFrame返回空值的问题及解决方法
问题:用Series/Index筛选DataFrame返回空值的原因及解决方法
原始数据
Date Category Sales Paid 8/12/2020 Table 1 table Yes 8/12/2020 Chair 3chairs Yes 13/1/2020 Cushion 8 cushions Yes 24/5/2020 Table 3Tables Yes 31/10/2020 Chair 12 Chairs No 11/7/2020 Mats 12Mats Yes 11/7/2020 Mats 4Mats Yes
问题重现
执行以下代码获取销量总数最高的日期:
ddate=df['Sales'].str.extract('^(\d+)', expand=False).astype(int).groupby(df['Date']).agg(sums='count').idxmax()
返回结果显示为['8/12/2020'],但用该对象筛选原DataFrame时返回空值:
df[df['Date'].eq(ddate)] # 返回空DataFrame
而直接用字符串筛选却能得到正确结果:
df[df['Date'].eq('8/12/2020')] # 正常返回数据
原因分析
你看到的['8/12/2020']并不是普通列表或字符串,而是Pandas的Index对象(索引类型)。当你用eq(ddate)比较时,是把Date列的每个字符串和整个Index对象做相等判断,这种匹配逻辑不成立,会返回全False的布尔序列,最终得到空DataFrame。
解决方法
方法1:提取Index中的单个字符串值
直接通过索引位置或iloc拿到字符串,再进行筛选:
# 提取第一个元素(这里只有一个最大值日期) target_date = ddate[0] # 或者用iloc更稳妥 target_date = ddate.iloc[0] df[df['Date'].eq(target_date)]
方法2:使用isin()方法匹配
isin()支持接受Index、列表等可迭代对象作为匹配条件,直接传入ddate即可:
df[df['Date'].isin(ddate)]
方法3:优化原代码直接获取字符串
在获取日期的代码末尾直接提取值,一步到位:
ddate = df['Sales'].str.extract('^(\d+)', expand=False).astype(int).groupby(df['Date']).agg(sums='count').idxmax()[0] # 此时ddate是字符串,直接筛选即可 df[df['Date'].eq(ddate)]
内容的提问来源于stack exchange,提问作者Dan
相关产品推荐
相关产品推荐

