Python如何实现仅相邻重叠多词短语的合并功能
相邻重叠多词短语合并实现方案
需求描述
我有一个多词短语列表,示例如下:
['President Barack', 'Barack Obama', 'New York', 'York City', 'United States', 'States of America', 'This is not overlapping']
希望合并其中相邻的重叠多词短语,最终得到如下结果:
['President Barack Obama', 'New York City', 'United States of America', 'This is not overlapping']
补充说明:非相邻的重叠短语不需要合并,例如列表为['President Barack', 'Some word', 'Other word', 'Barack Obama']时,前后两个含Barack的短语不需要合并。
初始代码问题
最初参考同类问题编写了如下代码:
strFrag = ['President Barack', 'Barack Obama', 'New York', 'York City', 'United States', 'States of America', 'This is not overlapping'] for repeat in range(0, len(strFrag)-1): bestMatch = [2, '', ''] #overlap score (minimum value 3), otherStr index, assembled str portion for otherStr in strFrag[1:]: for x in range(0,len(otherStr)): if otherStr[x:] == strFrag[0][:len(otherStr[x:])]: if len(otherStr)-x > bestMatch[0]: bestMatch = [len(otherStr)-x, strFrag.index(otherStr), otherStr[:x]+strFrag[0]] if otherStr[:-x] == strFrag[0][-len(otherStr[x:]):]: if x > bestMatch[0]: bestMatch = [x, strFrag.index(otherStr), strFrag[0]+otherStr[-x:]] if bestMatch[0] > 2: strFrag[0] = bestMatch[2] strFrag = strFrag[:bestMatch[1]]+strFrag[bestMatch[1]+1:]
该代码仅能合并列表的第一组重叠短语,运行后得到的结果不符合要求:
['President Barack Obama', 'New York', 'York City', 'United States', 'States of America', 'This is not overlapping']
后续实现代码及评估
后续重新编写了如下代码,输出符合预期:
strFrag = ['President Barack', 'Barack Obama', 'Obama of the USA', 'New York', 'York City', 'Test', 'Hello how', 'how you doin?'] for i in range(len(strFrag)): strFrag[i] = strFrag[i].split() for i in range(len(strFrag)-1,-1,-1): if (strFrag[i][0] == strFrag[i-1][-1]): strFrag[i-1].remove(strFrag[i-1][-1]) strFrag[i] = strFrag[i-1] + strFrag[i] strFrag.remove(strFrag[i-1]) for i in range(len(strFrag)): strFrag[i] = ' '.join(strFrag[i])
运行输出结果:
['President Barack Obama of the USA', 'New York City', 'Test', 'Hello how you doin?']
代码优化建议
- 倒序遍历处理相邻元素的逻辑合理,避免了正序遍历时索引偏移导致的漏处理问题
- 可补充边界判断,避免当
i=0时访问strFrag[i-1]出现的索引越界问题 - 当前仅支持单个词重叠的场景,如果有多个词重叠需求(例如
['a b c', 'b c d']需要合并为['a b c d']),可扩展匹配逻辑,对比后缀和前缀的最长公共词序列,而不是仅对比首尾单个词
内容的提问来源于stack exchange,提问作者Radix
相关产品推荐
相关产品推荐

