使用Pandas处理表头与子表头:日期列顺序校验代码问题
多级表头DataFrame列访问与时间顺序校验问题
我有一个带**表头(Header)和子表头(Subheader)**的DataFrame,能轻松获取这些表头/子表头的名称,但访问特定表头或子表头对应的列时遇到困难。
我的核心需求是:以表头+子表头作为键,校验对应列中的所有数据是否按正确的时间顺序排列。目前已经提取了按应有的时间顺序排列的表头/子表头元组列表,这些元组来自表格中的日期类表头/子表头。
表头/子表头的格式如下:
Header1 Header2 Header3 participant_number Subheader1-1 Subheader1-2 Subheader2-1 (no subheader)
可正常运行的代码段
这段代码能正确将指定列转换为日期类型:
date_columns= [(list of tuples,'')] for column_name, column_labels in date_columns.items(): header_label = column_labels[0] subheader_label = column_labels[1] column_index = None for idx, col in enumerate(df.columns): if col[0] == header_label or col[1] == subheader_label: column_index = idx break if column_index is not None: df[column_name] = pd.to_datetime(df.iloc[:, column_index], dayfirst=True, errors='coerce')
无法正常运行的代码段
这段用于校验时间顺序的代码无法正常工作,希望排查问题:
participant_errors = {} for current_column, next_column in zip(date_columns, date_columns[1:]): mask = ~(df[current_column].apply(pd.isna) | df[next_column].apply(pd.isna) | (df[current_column] <= df[next_column])) errors = df.loc[mask, [('key', '')] + (current_column[0], current_column[1]) + (next_column[0], next_column[1])] for row in errors.itertuples(): participant_number = row[('key', '')] current_value = getattr(row, current_column[0])[row.Index] next_value = getattr(row, next_column[0])[row.Index] if participant_number not in participant_errors: participant_errors[participant_number] = {current_column[0]: f"check data: {current_column[0]} ({current_value},) should occur before {next_column[0]} ({next_value})"} else: participant_errors[participant_number][current_column[0]] = f"check data: {current_column[0]} ({current_value},) should occur before {next_column[0]} ({next_value})" print(participant_errors)
内容的提问来源于stack exchange,提问作者HoopStart
相关产品推荐
相关产品推荐

