矩阵后处理功能优化:移动空列表头并移除空列需求
解决CSV矩阵处理中的表头左移与空列移除问题
需求说明
现有用于CSV转换的Python矩阵post_processing函数,已实现字符串替换、空单元格合并、空行移除功能。当前输出存在两类问题:部分表头(如'Red eyes')对应下方单元格全为空,且存在全空列。需在原有功能基础上添加两个逻辑:
- 将孤立的非空表头左移,填补前方空白表头位置
- 移除所有全空列
输入数组示例
csv_arr_output = [ ["", "", "Average", "", "Red" ], ["", "", "", "", "Red eyes"], ["", "height", "weight", "", "" ], ["Males", "1.9", "0.003", "40%", "" ], ["Females", "1.7", "0.002", "43%", "" ], ]
现有代码及当前输出
现有函数代码
def post_processing(csv_array, common_percent_threshold=0.5): # Replace longer next cell if they have a certain percent of common string for row in range(len(csv_array) - 1): for col in range(len(csv_array[row])): current_cell = csv_array[row][col] next_cell = csv_array[row + 1][col] if calculate_common_percent(current_cell, next_cell) >= common_percent_threshold: if len(current_cell) < len(next_cell): csv_array[row][col] = next_cell csv_array[row + 1][col] = "" # Merge empty cells in the same column for col in range(len(csv_array[0])): column_cells = [row[col] for row in csv_array] non_empty_cells = [cell for cell in column_cells if cell != ""] empty_cells = [""] * (len(column_cells) - len(non_empty_cells)) for row in range(len(csv_array)): if csv_array[row][col] == "": csv_array[row][col] = empty_cells.pop(0) # Remove empty rows csv_array = [row for row in csv_array if any(cell != "" for cell in row)] return csv_array def calculate_common_percent(cell1, cell2): set1 = set(cell1.split()) set2 = set(cell2.split()) # Check if both sets are empty if not set1 and not set2: return 0.0 common_words = set1.intersection(set2) total_words = set1.union(set2) # Avoid division by zero if not total_words: return 0.0 return len(common_words) / len(total_words)
当前输出
[ ["", "", "Average", "", "Red eyes"], ["", "height", "weight", "", "" ], ["Males", "1.9", "0.003", "40%", "" ], ["Females", "1.7", "0.002", "43%", "" ], ]
修改后的代码
在原有函数末尾新增表头左移处理和移除全空列两个步骤,具体实现如下:
def post_processing(csv_array, common_percent_threshold=0.5): # 1. 字符串替换逻辑(原有) for row in range(len(csv_array) - 1): for col in range(len(csv_array[row])): current_cell = csv_array[row][col] next_cell = csv_array[row + 1][col] if calculate_common_percent(current_cell, next_cell) >= common_percent_threshold: if len(current_cell) < len(next_cell): csv_array[row][col] = next_cell csv_array[row + 1][col] = "" # 2. 空单元格合并(原有) for col in range(len(csv_array[0])): column_cells = [row[col] for row in csv_array] non_empty_cells = [cell for cell in column_cells if cell != ""] empty_cells = [""] * (len(column_cells) - len(non_empty_cells)) for row in range(len(csv_array)): if csv_array[row][col] == "": csv_array[row][col] = empty_cells.pop(0) # 3. 移除空行(原有) csv_array = [row for row in csv_array if any(cell != "" for cell in row)] # 4. 新增:表头左移处理(针对前两行表头,将非空表头填补左侧空白) header_rows = 2 for row_idx in range(header_rows): non_empty_headers = [cell for cell in csv_array[row_idx] if cell != ""] new_row = [""] * len(csv_array[row_idx]) fill_pos = 0 for header in non_empty_headers: # 找到第一个空白位置填充 while fill_pos < len(new_row) and new_row[fill_pos] != "": fill_pos += 1 if fill_pos < len(new_row): new_row[fill_pos] = header fill_pos += 1 csv_array[row_idx] = new_row # 5. 新增:移除所有全空列 non_empty_col_indices = [] for col_idx in range(len(csv_array[0])): if any(row[col_idx] != "" for row in csv_array): non_empty_col_indices.append(col_idx) # 仅保留非空列 csv_array = [[row[col] for col in non_empty_col_indices] for row in csv_array] return csv_array def calculate_common_percent(cell1, cell2): set1 = set(cell1.split()) set2 = set(cell2.split()) # Check if both sets are empty if not set1 and not set2: return 0.0 common_words = set1.intersection(set2) total_words = set1.union(set2) # Avoid division by zero if not total_words: return 0.0 return len(common_words) / len(total_words)
最终输出
运行修改后的函数,输入示例将得到如下结果:
[ ["", "", "Average", "Red eyes"], ["", "height", "weight", ""], ["Males", "1.9", "0.003", "40%"], ["Females", "1.7", "0.002", "43%"] ]
逻辑说明
- 表头行中的'Red eyes'被左移至原第四列的空白位置,避免孤立在全空列
- 原第五列因全空被移除,所有无数据的列均被清理
- 原有功能不受影响,仅新增逻辑优化表头与列结构
内容的提问来源于stack exchange,提问作者RONY SALEM
相关产品推荐
相关产品推荐

