如何用Python将竖线分隔文本转换为每行最多4个单元格的格式?
需求说明
需要将竖线分隔的多行文本转换为每行最多包含4个单元格的格式,规则如下:
- 若原行单元格数量超过4,超出部分拆分为新行
- 拆分后的新行必须保留原行的前两个单元格
- 每行单元格数不得超过4
输入文本示例
text = """A | B | Lorem | Ipsum | is | simply | dummy C | D | text | of | the | printing | and E | F | typesetting | industry. | Lorem G | H | more | recently | with | desktop | publishing | software | like | Aldus I | J | Ipsum | has | been | the | industry's K | L | standard | dummy | text | ever | since | the | 1500s M | N | took | a O | P | scrambled | it | to | make | a | type | specimen | book"""
期望输出示例
A | B | Lorem | Ipsum A | B | is | simply A | B | dummy C | D | text | of C | D | the | printing C | D | and E | F | typesetting | industry. E | F | Lorem G | H | more | recently G | H | with | desktop G | H | publishing | software G | H | like | Aldus ...
未完成代码
new_text = "" for i in text.split('\n'): row = i.split(' | ') if len(row) == 4: new_text = new_text + i + '\n' elif len(row) > 4: for j in range(len(row)): if j < 3: new_text = new_text + row[0] + ' | ' + row[1] + ...
完整实现代码
text = """A | B | Lorem | Ipsum | is | simply | dummy C | D | text | of | the | printing | and E | F | typesetting | industry. | Lorem G | H | more | recently | with | desktop | publishing | software | like | Aldus I | J | Ipsum | has | been | the | industry's K | L | standard | dummy | text | ever | since | the | 1500s M | N | took | a O | P | scrambled | it | to | make | a | type | specimen | book""" new_lines = [] # 遍历每一行原始文本 for line in text.split('\n'): cells = line.split(' | ') # 取出固定保留的前两个单元格 prefix = [cells[0], cells[1]] # 剩下需要拆分的内容从第三个单元格开始 remaining = cells[2:] # 按每组最多2个单元格拆分(加上前缀正好凑4个) for i in range(0, len(remaining), 2): current_group = prefix + remaining[i:i+2] new_lines.append(' | '.join(current_group)) # 拼接所有新行得到最终文本 new_text = '\n'.join(new_lines) print(new_text)
代码逻辑说明
- 先把每一行按
|拆成单元格列表 - 提取每行的前两个单元格作为所有拆分后新行的固定开头
- 从第三个单元格开始,每2个单元格为一组(加上前缀刚好4个),循环处理所有剩余内容
- 把前缀和每组内容拼接成新行,最后合并所有行输出
内容的提问来源于stack exchange,提问作者Mohit Aswani
相关产品推荐
相关产品推荐

