You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

矩阵后处理功能优化:移动空列表头并移除空列需求

解决CSV矩阵处理中的表头左移与空列移除问题

需求说明

现有用于CSV转换的Python矩阵post_processing函数,已实现字符串替换、空单元格合并、空行移除功能。当前输出存在两类问题:部分表头(如'Red eyes')对应下方单元格全为空,且存在全空列。需在原有功能基础上添加两个逻辑:

  • 将孤立的非空表头左移,填补前方空白表头位置
  • 移除所有全空列

输入数组示例

csv_arr_output = [
    ["",        "",       "Average", "",    "Red"     ],
    ["",        "",       "",        "",    "Red eyes"],
    ["",        "height", "weight",  "",    ""        ],
    ["Males",   "1.9",    "0.003",   "40%", ""        ],
    ["Females", "1.7",    "0.002",   "43%", ""        ],
]

现有代码及当前输出

现有函数代码

def post_processing(csv_array, common_percent_threshold=0.5):
    # Replace longer next cell if they have a certain percent of common string
    for row in range(len(csv_array) - 1):
        for col in range(len(csv_array[row])):
            current_cell = csv_array[row][col]
            next_cell = csv_array[row + 1][col]

            if calculate_common_percent(current_cell, next_cell) >= common_percent_threshold:
                if len(current_cell) < len(next_cell):
                    csv_array[row][col] = next_cell
                    csv_array[row + 1][col] = ""

    # Merge empty cells in the same column
    for col in range(len(csv_array[0])):
        column_cells = [row[col] for row in csv_array]
        non_empty_cells = [cell for cell in column_cells if cell != ""]
        empty_cells = [""] * (len(column_cells) - len(non_empty_cells))

        for row in range(len(csv_array)):
            if csv_array[row][col] == "":
                csv_array[row][col] = empty_cells.pop(0)

    # Remove empty rows
    csv_array = [row for row in csv_array if any(cell != "" for cell in row)]

    return csv_array

def calculate_common_percent(cell1, cell2):
    set1 = set(cell1.split())
    set2 = set(cell2.split())
    
    # Check if both sets are empty
    if not set1 and not set2:
        return 0.0
    
    common_words = set1.intersection(set2)
    total_words = set1.union(set2)
    
    # Avoid division by zero
    if not total_words:
        return 0.0
    
    return len(common_words) / len(total_words)

当前输出

[
    ["",        "",       "Average", "",    "Red eyes"],
    ["",        "height", "weight",  "",    ""        ],
    ["Males",   "1.9",    "0.003",   "40%", ""        ],
    ["Females", "1.7",    "0.002",   "43%", ""        ],
]

修改后的代码

在原有函数末尾新增表头左移处理和移除全空列两个步骤,具体实现如下:

def post_processing(csv_array, common_percent_threshold=0.5):
    # 1. 字符串替换逻辑(原有)
    for row in range(len(csv_array) - 1):
        for col in range(len(csv_array[row])):
            current_cell = csv_array[row][col]
            next_cell = csv_array[row + 1][col]

            if calculate_common_percent(current_cell, next_cell) >= common_percent_threshold:
                if len(current_cell) < len(next_cell):
                    csv_array[row][col] = next_cell
                    csv_array[row + 1][col] = ""

    # 2. 空单元格合并(原有)
    for col in range(len(csv_array[0])):
        column_cells = [row[col] for row in csv_array]
        non_empty_cells = [cell for cell in column_cells if cell != ""]
        empty_cells = [""] * (len(column_cells) - len(non_empty_cells))

        for row in range(len(csv_array)):
            if csv_array[row][col] == "":
                csv_array[row][col] = empty_cells.pop(0)

    # 3. 移除空行(原有)
    csv_array = [row for row in csv_array if any(cell != "" for cell in row)]

    # 4. 新增:表头左移处理(针对前两行表头,将非空表头填补左侧空白)
    header_rows = 2
    for row_idx in range(header_rows):
        non_empty_headers = [cell for cell in csv_array[row_idx] if cell != ""]
        new_row = [""] * len(csv_array[row_idx])
        fill_pos = 0
        for header in non_empty_headers:
            # 找到第一个空白位置填充
            while fill_pos < len(new_row) and new_row[fill_pos] != "":
                fill_pos += 1
            if fill_pos < len(new_row):
                new_row[fill_pos] = header
                fill_pos += 1
        csv_array[row_idx] = new_row

    # 5. 新增:移除所有全空列
    non_empty_col_indices = []
    for col_idx in range(len(csv_array[0])):
        if any(row[col_idx] != "" for row in csv_array):
            non_empty_col_indices.append(col_idx)
    # 仅保留非空列
    csv_array = [[row[col] for col in non_empty_col_indices] for row in csv_array]

    return csv_array

def calculate_common_percent(cell1, cell2):
    set1 = set(cell1.split())
    set2 = set(cell2.split())
    
    # Check if both sets are empty
    if not set1 and not set2:
        return 0.0
    
    common_words = set1.intersection(set2)
    total_words = set1.union(set2)
    
    # Avoid division by zero
    if not total_words:
        return 0.0
    
    return len(common_words) / len(total_words)

最终输出

运行修改后的函数,输入示例将得到如下结果:

[
    ["", "", "Average", "Red eyes"],
    ["", "height", "weight", ""],
    ["Males", "1.9", "0.003", "40%"],
    ["Females", "1.7", "0.002", "43%"]
]

逻辑说明

  • 表头行中的'Red eyes'被左移至原第四列的空白位置,避免孤立在全空列
  • 原第五列因全空被移除,所有无数据的列均被清理
  • 原有功能不受影响,仅新增逻辑优化表头与列结构

内容的提问来源于stack exchange,提问作者RONY SALEM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 00:16:01