如何使用Pandas将DataFrame指定行转为新列匹配对应下属条目
实现方案
Pandas 实现(推荐)
核心思路是识别出仅第一列有值的分类行,通过前向填充把分类值向下传递到所有对应数据行,最后过滤掉原始分类行即可:
import pandas as pd import numpy as np # 构造原始DataFrame table = {'0': {6: 'Becks', 7: '307NRR', 8: '321NRR', 9: '342NRR', 10: 'Campbell', 11: '329NRR', 12: '347NRR', 13: 'Crows', 14: 'C3001R'}, '1': {6: np.nan, 7: 'R', 8: 'R', 9: 'R', 10: np.nan, 11: 'R', 12: 'R', 13: np.nan, 14: 'R'}, '2': {6: np.nan, 7: 'CM,SG', 8: 'CM,SG', 9: 'CM,SG', 10: np.nan, 11: 'None', 12: 'None', 13: np.nan, 14: 'None'}, '3': {6: np.nan, 7: 3.0, 8: 3.2, 9: 3.4, 10: np.nan, 11: 3.2, 12: 3.4, 13: np.nan, 14: 3.0}} df = pd.DataFrame(table) # 新建分类列,仅分类行保留第一列值,其余为NaN df['category'] = np.where(df[['1','2','3']].isna().all(axis=1), df['0'], np.nan) # 前向填充分类列,把分类值传递到对应数据行 df['category'] = df['category'].ffill() # 过滤分类行,调整列顺序 result = df[~df[['1','2','3']].isna().all(axis=1)][['category','0','1','2','3']].reset_index(drop=True)
Python 基础模块实现
无需引入第三方依赖,按索引顺序遍历原始数据,记录当前分类后拼接数据即可:
import numpy as np table = {'0': {6: 'Becks', 7: '307NRR', 8: '321NRR', 9: '342NRR', 10: 'Campbell', 11: '329NRR', 12: '347NRR', 13: 'Crows', 14: 'C3001R'}, '1': {6: np.nan, 7: 'R', 8: 'R', 9: 'R', 10: np.nan, 11: 'R', 12: 'R', 13: np.nan, 14: 'R'}, '2': {6: np.nan, 7: 'CM,SG', 8: 'CM,SG', 9: 'CM,SG', 10: np.nan, 11: 'None', 12: 'None', 13: np.nan, 14: 'None'}, '3': {6: np.nan, 7: 3.0, 8: 3.2, 9: 3.4, 10: np.nan, 11: 3.2, 12: 3.4, 13: np.nan, 14: 3.0}} sorted_index = sorted(table['0'].keys()) current_category = None result = [] for idx in sorted_index: col0, col1, col2, col3 = table['0'][idx], table['1'][idx], table['2'][idx], table['3'][idx] # 判断是否为分类行 if np.isna(col1) and np.isna(col2) and np.isna(col3): current_category = col0 else: result.append([current_category, col0, col1, col2, col3])
内容的提问来源于stack exchange,提问作者Kabocha Porter
相关产品推荐
相关产品推荐

