如何在两个Pandas表上应用自定义函数并输出结果?
问题描述
我有如下两个Pandas DataFrame:
import pandas as pd import numpy as np df1 = pd.DataFrame(data={'1': ['john', '10', 'john'], '2': ['mike', '30', 'ana'], '3': ['ana', '20', 'mike'], '4': ['eve', 'eve', 'eve'], '5': ['10', np.NaN, '10'], '6': [np.NaN, np.NaN, '20']}, index=pd.Series(['ind1', 'ind2', 'ind3'], name='index')) # df1输出: # 1 2 3 4 5 6 # index # ind1 john mike ana eve 10 NaN # ind2 10 30 20 eve NaN NaN # ind3 john ana mike eve 10 20 df2 = pd.DataFrame(data={'first_n': [4, 4, 3]}, index=pd.Series(['ind1', 'ind2', 'ind3'], name='index')) # df2输出: # first_n # index # ind1 4 # ind2 4 # ind3 3
我还有一个自定义函数get_rev_first_n,功能是反转列表并获取前n个非NA元素:
def get_rev_first_n(row, top_n): rev_row = [x for x in row[::-1] if x == x] # x == x 用于判断非NaN return rev_row[:top_n] # 示例调用: # get_rev_first_n(['john', 'mike', 'ana', 'eve', '10', np.NaN], 4) # 输出:['10', 'eve', 'ana', 'mike']
请问如何将该函数同时作用于df1和df2,输出列表或列?
解决方法
可以通过pandas.DataFrame.apply逐行处理df1,同时利用索引匹配df2中对应的first_n值作为函数参数,具体实现如下:
方法1:输出为列表形式的新列
# 按行处理df1,每一行调用自定义函数并传入df2中对应的top_n result = df1.apply( lambda row: get_rev_first_n(row, df2.loc[row.name, 'first_n']), axis=1 ) print(result)
输出结果:
index ind1 [10, eve, ana, mike] ind2 [eve, 20, 30, 10] ind3 [20, 10, eve] dtype: object
方法2:将列表展开为多列
如果需要把列表中的元素拆分成单独的列,可以用pd.DataFrame将结果转换为新的DataFrame:
expanded_result = pd.DataFrame(result.tolist(), index=result.index) print(expanded_result)
输出结果:
0 1 2 3 index ind1 10 eve ana mike ind2 eve 20 30 10 ind3 20 10 eve NaN
核心说明
axis=1指定按行处理df1的每一行数据row.name获取当前行的索引,以此从df2中匹配对应的first_n参数x == x是判断元素非NaN的简洁写法(因为NaN不等于自身)
内容的提问来源于stack exchange,提问作者bltSandwich21
相关产品推荐
相关产品推荐

