You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析多竖线分隔表格时的换行符处理问题

解决JIRA表格数据解析换行问题

问题背景

  • 需求:将从JIRA提取的表格数据导入Excel
  • 表格规则:双竖线||包裹列名,单竖线|包裹列值

原始JIRA表格数据

||Name||Age||Address||Phones||size||

|Edwards|22|London|06 45 06 06 06

06 75 85 06 06

07 85 22 15 48|180cm|

期望解析结果

| Name | Age | Address | Phones | size |
| --- | --- | --- | --- | --- |
| Edwards | 22 | London | 06 45 06 06 \n 06 75 85 06 06 \n 07 85 22 15 48 | 180cm |

当前问题

解析时换行后的内容全部被合并到最后一列单元格中,现有Python处理函数如下:

def process_description_column(df):

    data_dictionary = {}
    for index, row in df.iterrows():
        modified_text = row['Description'].replace('*', '')
        description_text = ""

        lines = modified_text.strip().split('\n')
        in_description_section = False
        line_counter = 0
        
        for line in lines:
            line = line.strip()
            in_description_section, line_counter, description_text = process_line_description(
                line, in_description_section, line_counter, description_text)
            
            if '||' in line:
                column_names = re.split(r'\|\|', line)
                columns = [col.strip() for col in column_names if col.strip()]
                for col in columns:
                    if col not in data_dictionary:
                        data_dictionary[col] = []
            
            elif '|' in line and columns:
                line_without_bar = line.replace('|', '').strip()
                
                if '\n' in line_without_bar:
                    values = re.split(r'\|', line)
                    values = [value.strip() for value in values if value.strip()]
                    
                    while len(values) < len(columns):
                        values.append('')
                    for col, value in zip(columns, values):
                        data_dictionary[col].append(value)
                else:
                    if data_dictionary[columns[-1]]:
                        data_dictionary[columns[-1]][-1] += ' ' + line_without_bar
                    else:
                        values = re.split(r'\|', line)
                        values = [value.strip() for value in values if value.strip()]
                        
                        if len(columns) == len(values):
                            for col, value in zip(columns, values):
                                data_dictionary[col].append(value)
                        else:
                            while len(values) < len(columns):
                                values.append('')
                            for col, value in zip(columns, values):
                                data_dictionary[col].append(value)
        
        data_dictionary['Description'] = [description_text.strip()]
    return data_dictionary

解决方案

核心问题是没正确识别多行单元格的边界——JIRA表格中,多行值被同一对|包裹,换行只是单元格内的内容,不是新行的开始。修改思路:先合并同一表格行的所有内容,再分割列值,保留单元格内换行。

修改后的代码:

import re

def process_description_column(df):
    data_dictionary = {}
    for index, row in df.iterrows():
        modified_text = row['Description'].replace('*', '')
        description_text = ""
        # 保留每行原始换行,避免提前丢失格式
        lines = [line.rstrip('\n') for line in modified_text.splitlines()]
        
        in_description_section = False
        line_counter = 0
        columns = []
        current_table_row = None
        
        for line in lines:
            stripped_line = line.strip()
            # 保留原有的非表格内容处理逻辑
            in_description_section, line_counter, description_text = process_line_description(
                line, in_description_section, line_counter, description_text)
            
            # 处理列名行
            if '||' in stripped_line:
                column_names = re.split(r'\|\|', stripped_line)
                columns = [col.strip() for col in column_names if col.strip()]
                # 初始化字典键
                for col in columns:
                    if col not in data_dictionary:
                        data_dictionary[col] = []
                # 处理未完成的表格行(防止表格中断)
                if current_table_row is not None:
                    _process_table_row(current_table_row, columns, data_dictionary)
                    current_table_row = None
                continue
            
            # 处理表格内容行
            if columns:
                # 以|开头的是新表格行
                if stripped_line.startswith('|'):
                    # 先处理上一行未完成的内容
                    if current_table_row is not None:
                        _process_table_row(current_table_row, columns, data_dictionary)
                    # 开始新行,保留原始内容
                    current_table_row = line.rstrip()
                elif current_table_row is not None:
                    # 追加当前行到单元格,保留换行
                    current_table_row += '\n' + line.rstrip()
        
        # 处理最后一行未完成的表格行
        if current_table_row is not None and columns:
            _process_table_row(current_table_row, columns, data_dictionary)
        
        # 更新Description字段
        if 'Description' not in data_dictionary:
            data_dictionary['Description'] = []
        data_dictionary['Description'].append(description_text.strip())
    
    return data_dictionary

def _process_table_row(row_text, columns, data_dict):
    """辅助函数:处理合并后的完整表格行"""
    # 分割列值,跳过转义的|(适配特殊场景)
    values = re.split(r'(?<!\\)\|', row_text)
    # 清理空值,保留单元格内的换行
    values = [val.strip() for val in values if val.strip()]
    # 补全缺失的列值
    while len(values) < len(columns):
        values.append('')
    # 写入字典,将单元格内换行转为\n(按示例要求)
    for col, val in zip(columns, values):
        data_dict[col].append(val.replace('\n', '\\n'))

关键优化点

  • 新增_process_table_row辅助函数,拆分处理逻辑,更易维护
  • 以|开头作为新表格行的标志,后续非|行归为当前单元格的多行内容
  • 用正则(?<!\\)\|分割列值,避免误拆分单元格内的转义|
  • 完整保留单元格内的换行,转为示例要求的\n格式

内容的提问来源于stack exchange,提问作者user27963554

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 17:31:11