You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何实现仅相邻重叠多词短语的合并功能

相邻重叠多词短语合并实现方案

需求描述

我有一个多词短语列表,示例如下:

['President Barack', 'Barack Obama', 'New York', 'York City', 'United States', 'States of America', 'This is not overlapping']

希望合并其中相邻的重叠多词短语,最终得到如下结果:

['President Barack Obama', 'New York City', 'United States of America', 'This is not overlapping']

补充说明:非相邻的重叠短语不需要合并,例如列表为['President Barack', 'Some word', 'Other word', 'Barack Obama']时,前后两个含Barack的短语不需要合并。

初始代码问题

最初参考同类问题编写了如下代码:

strFrag = ['President Barack', 'Barack Obama', 'New York', 'York City', 'United States', 'States of America', 'This is not overlapping']

for repeat in range(0, len(strFrag)-1):
    bestMatch = [2, '', ''] #overlap score (minimum value 3), otherStr index, assembled str portion
    for otherStr in strFrag[1:]:
        for x in range(0,len(otherStr)):
            if otherStr[x:] == strFrag[0][:len(otherStr[x:])]:
                if len(otherStr)-x > bestMatch[0]:
                    bestMatch = [len(otherStr)-x, strFrag.index(otherStr), otherStr[:x]+strFrag[0]]
            if otherStr[:-x] == strFrag[0][-len(otherStr[x:]):]:
                if x > bestMatch[0]:
                    bestMatch = [x, strFrag.index(otherStr), strFrag[0]+otherStr[-x:]]
    if bestMatch[0] > 2:
        strFrag[0] = bestMatch[2]
        strFrag = strFrag[:bestMatch[1]]+strFrag[bestMatch[1]+1:]

该代码仅能合并列表的第一组重叠短语,运行后得到的结果不符合要求:

['President Barack Obama', 'New York', 'York City', 'United States', 'States of America', 'This is not overlapping']

后续实现代码及评估

后续重新编写了如下代码,输出符合预期:

strFrag = ['President Barack', 'Barack Obama', 'Obama of the USA', 'New York', 'York City', 'Test', 'Hello how', 'how you doin?']

for i in range(len(strFrag)):
  strFrag[i] = strFrag[i].split()

for i in range(len(strFrag)-1,-1,-1):
    if (strFrag[i][0] == strFrag[i-1][-1]):
      strFrag[i-1].remove(strFrag[i-1][-1])
      strFrag[i] = strFrag[i-1] + strFrag[i]
      strFrag.remove(strFrag[i-1])

for i in range(len(strFrag)):
  strFrag[i] = ' '.join(strFrag[i])

运行输出结果:

['President Barack Obama of the USA',
 'New York City',
 'Test',
 'Hello how you doin?']

代码优化建议

  • 倒序遍历处理相邻元素的逻辑合理,避免了正序遍历时索引偏移导致的漏处理问题
  • 可补充边界判断,避免当i=0时访问strFrag[i-1]出现的索引越界问题
  • 当前仅支持单个词重叠的场景,如果有多个词重叠需求(例如['a b c', 'b c d']需要合并为['a b c d']),可扩展匹配逻辑,对比后缀和前缀的最长公共词序列,而不是仅对比首尾单个词

内容的提问来源于stack exchange,提问作者Radix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 13:27:00