如何合并Python列表中标记为continued的元素至后续首个非标记元素
实现方法
核心思路
遍历所有元素,先把连续的continued文本暂存起来,直到碰到第一个非continued的元素时,把暂存的文本和这个元素的文本合并,保留该元素的label,最后把合并后的结果存入最终列表。
代码实现
# 先把原始文本按行拆分,得到每一行的内容 raw_lines = """continued In the film, "Girl Interrupted," Winona Ryder plays an 18-year-old continued who enters a mental institution for continued what anecdote is diagnosed as borderline personality disorder anecdote The year is 1967 anecdote the country is in turmoil over Vietnam and civil rights continued While continued lying on her bed one night continued and continued watching TV continued , anecdote she sees a news report about a demonstration continued The narrator says something continued that might apply to today's turmoil continued : continued "We live in a time of doubt continued . continued The institutions continued we once trusted no longer anecdote seem reliable." continued As 2014 ends Modd-NU statistics the stock market is at record highs assumption our traditional institutions and self-confidence are in decline continued A Pew Research Center study confirms one trend testimony that has been obvious over several years assumption The "typical" American family is no longer typical statistics Just 46 percent of American children now live in homes with their married, heterosexual parents statistics Five percent have no parents at home continued They most likely are living with grandparents continued , testimony says the study assumption These startling figures about the decline of the American family contrast with the year 1960 continued when Modd-NU statistics 73 percent of American children lived in traditional families assumption A major contributor to this trend has been the assault on marriage and other institutions by the Baby Boom generation """.splitlines() # 初始化暂存continued文本的列表和结果列表 pending_continued = [] result = [] for line in raw_lines: # 按制表符拆分label和text,只拆一次,避免text里的空格影响 if '\t' in line: label, text = line.split('\t', 1) else: # 处理原始文本中可能用多个空格代替制表符的情况 split_idx = line.find(' ') label = line[:split_idx].strip() text = line[split_idx:].strip() if label == 'continued': # 暂存当前continued文本,去除首尾多余空格 pending_continued.append(text) else: # 合并暂存文本与当前元素的文本,用空格连接保证语句通顺 combined_text = ' '.join(pending_continued) + ' ' + text # 把合并后的(label, 文本)存入结果列表,去除首尾多余空格 result.append( (label, combined_text.strip()) ) # 清空暂存列表,准备处理下一批continued pending_continued = [] # 输出结果示例 for label, text in result: print(f"{label}\t{text}")
代码说明
- 先将原始文本拆分为单行,再拆分每行的
label和text:优先用制表符拆分,避免文本内容里的空格干扰;如果没有制表符,就按第一个空格位置拆分。 - 遇到
continued标记时,将对应文本存入暂存列表;碰到非continued元素时,把暂存的所有文本和当前文本拼接,保证语句连贯后,将(label, 合并文本)存入结果列表。 - 最后可以遍历结果列表,直接输出或做后续处理。
内容的提问来源于stack exchange,提问作者Haochuan Liu
相关产品推荐
相关产品推荐

