You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并Python列表中标记为continued的元素至后续首个非标记元素

实现方法

核心思路

遍历所有元素,先把连续的continued文本暂存起来,直到碰到第一个非continued的元素时,把暂存的文本和这个元素的文本合并,保留该元素的label,最后把合并后的结果存入最终列表。

代码实现

# 先把原始文本按行拆分,得到每一行的内容
raw_lines = """continued   In the film, "Girl Interrupted," Winona Ryder plays an 18-year-old
continued   who enters a mental institution for
continued   what
anecdote    is diagnosed as borderline personality disorder
anecdote    The year is 1967
anecdote    the country is in turmoil over Vietnam and civil rights
continued   While
continued   lying on her bed one night
continued   and
continued   watching TV
continued   ,
anecdote    she sees a news report about a demonstration
continued   The narrator says something
continued   that might apply to today's turmoil
continued   :
continued   "We live in a time of doubt
continued   .
continued   The institutions
continued   we once trusted no longer
anecdote    seem reliable."
continued   As 2014 ends    Modd-NU
statistics  the stock market is at record highs
assumption  our traditional institutions and self-confidence are in decline
continued   A Pew Research Center study confirms one trend
testimony   that has been obvious over several years
assumption  The "typical" American family is no longer typical
statistics  Just 46 percent of American children now live in homes with their married, heterosexual parents
statistics  Five percent have no parents at home
continued   They most likely are living with grandparents
continued   ,
testimony   says the study
assumption  These startling figures about the decline of the American family contrast with the year 1960
continued   when    Modd-NU
statistics  73 percent of American children lived in traditional families
assumption  A major contributor to this trend has been the assault on marriage and other institutions by the Baby Boom generation
""".splitlines()

# 初始化暂存continued文本的列表和结果列表
pending_continued = []
result = []

for line in raw_lines:
    # 按制表符拆分label和text,只拆一次,避免text里的空格影响
    if '\t' in line:
        label, text = line.split('\t', 1)
    else:
        # 处理原始文本中可能用多个空格代替制表符的情况
        split_idx = line.find(' ')
        label = line[:split_idx].strip()
        text = line[split_idx:].strip()
    
    if label == 'continued':
        # 暂存当前continued文本,去除首尾多余空格
        pending_continued.append(text)
    else:
        # 合并暂存文本与当前元素的文本,用空格连接保证语句通顺
        combined_text = ' '.join(pending_continued) + ' ' + text
        # 把合并后的(label, 文本)存入结果列表,去除首尾多余空格
        result.append( (label, combined_text.strip()) )
        # 清空暂存列表,准备处理下一批continued
        pending_continued = []

# 输出结果示例
for label, text in result:
    print(f"{label}\t{text}")

代码说明

  • 先将原始文本拆分为单行,再拆分每行的label和text:优先用制表符拆分,避免文本内容里的空格干扰;如果没有制表符,就按第一个空格位置拆分。
  • 遇到continued标记时,将对应文本存入暂存列表;碰到非continued元素时,把暂存的所有文本和当前文本拼接,保证语句连贯后,将(label, 合并文本)存入结果列表。
  • 最后可以遍历结果列表,直接输出或做后续处理。

内容的提问来源于stack exchange,提问作者Haochuan Liu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 08:28:40