Python解析多竖线分隔表格时的换行符处理问题
解决JIRA表格数据解析换行问题
问题背景
- 需求:将从JIRA提取的表格数据导入Excel
- 表格规则:双竖线
||包裹列名,单竖线|包裹列值
原始JIRA表格数据
||Name||Age||Address||Phones||size|| |Edwards|22|London|06 45 06 06 06 06 75 85 06 06 07 85 22 15 48|180cm|
期望解析结果
| Name | Age | Address | Phones | size | | --- | --- | --- | --- | --- | | Edwards | 22 | London | 06 45 06 06 \n 06 75 85 06 06 \n 07 85 22 15 48 | 180cm |
当前问题
解析时换行后的内容全部被合并到最后一列单元格中,现有Python处理函数如下:
def process_description_column(df): data_dictionary = {} for index, row in df.iterrows(): modified_text = row['Description'].replace('*', '') description_text = "" lines = modified_text.strip().split('\n') in_description_section = False line_counter = 0 for line in lines: line = line.strip() in_description_section, line_counter, description_text = process_line_description( line, in_description_section, line_counter, description_text) if '||' in line: column_names = re.split(r'\|\|', line) columns = [col.strip() for col in column_names if col.strip()] for col in columns: if col not in data_dictionary: data_dictionary[col] = [] elif '|' in line and columns: line_without_bar = line.replace('|', '').strip() if '\n' in line_without_bar: values = re.split(r'\|', line) values = [value.strip() for value in values if value.strip()] while len(values) < len(columns): values.append('') for col, value in zip(columns, values): data_dictionary[col].append(value) else: if data_dictionary[columns[-1]]: data_dictionary[columns[-1]][-1] += ' ' + line_without_bar else: values = re.split(r'\|', line) values = [value.strip() for value in values if value.strip()] if len(columns) == len(values): for col, value in zip(columns, values): data_dictionary[col].append(value) else: while len(values) < len(columns): values.append('') for col, value in zip(columns, values): data_dictionary[col].append(value) data_dictionary['Description'] = [description_text.strip()] return data_dictionary
解决方案
核心问题是没正确识别多行单元格的边界——JIRA表格中,多行值被同一对|包裹,换行只是单元格内的内容,不是新行的开始。修改思路:先合并同一表格行的所有内容,再分割列值,保留单元格内换行。
修改后的代码:
import re def process_description_column(df): data_dictionary = {} for index, row in df.iterrows(): modified_text = row['Description'].replace('*', '') description_text = "" # 保留每行原始换行,避免提前丢失格式 lines = [line.rstrip('\n') for line in modified_text.splitlines()] in_description_section = False line_counter = 0 columns = [] current_table_row = None for line in lines: stripped_line = line.strip() # 保留原有的非表格内容处理逻辑 in_description_section, line_counter, description_text = process_line_description( line, in_description_section, line_counter, description_text) # 处理列名行 if '||' in stripped_line: column_names = re.split(r'\|\|', stripped_line) columns = [col.strip() for col in column_names if col.strip()] # 初始化字典键 for col in columns: if col not in data_dictionary: data_dictionary[col] = [] # 处理未完成的表格行(防止表格中断) if current_table_row is not None: _process_table_row(current_table_row, columns, data_dictionary) current_table_row = None continue # 处理表格内容行 if columns: # 以|开头的是新表格行 if stripped_line.startswith('|'): # 先处理上一行未完成的内容 if current_table_row is not None: _process_table_row(current_table_row, columns, data_dictionary) # 开始新行,保留原始内容 current_table_row = line.rstrip() elif current_table_row is not None: # 追加当前行到单元格,保留换行 current_table_row += '\n' + line.rstrip() # 处理最后一行未完成的表格行 if current_table_row is not None and columns: _process_table_row(current_table_row, columns, data_dictionary) # 更新Description字段 if 'Description' not in data_dictionary: data_dictionary['Description'] = [] data_dictionary['Description'].append(description_text.strip()) return data_dictionary def _process_table_row(row_text, columns, data_dict): """辅助函数:处理合并后的完整表格行""" # 分割列值,跳过转义的|(适配特殊场景) values = re.split(r'(?<!\\)\|', row_text) # 清理空值,保留单元格内的换行 values = [val.strip() for val in values if val.strip()] # 补全缺失的列值 while len(values) < len(columns): values.append('') # 写入字典,将单元格内换行转为\n(按示例要求) for col, val in zip(columns, values): data_dict[col].append(val.replace('\n', '\\n'))
关键优化点
- 新增
_process_table_row辅助函数,拆分处理逻辑,更易维护 - 以
|开头作为新表格行的标志,后续非|行归为当前单元格的多行内容 - 用正则
(?<!\\)\|分割列值,避免误拆分单元格内的转义| - 完整保留单元格内的换行,转为示例要求的
\n格式
内容的提问来源于stack exchange,提问作者user27963554
相关产品推荐
相关产品推荐

