Python如何从含连续近似重复的字符串列表中保留最短重复项
解决方案
适用严格子串包含的场景
你给出的示例中同组连续近似重复字符串满足「长字符串完全包含组内最短字符串」的特征,可以用以下方法实现,完全保留原列表顺序,仅保留每组最短字符串:
def deduplicate_continuous_approx(lst): if not lst: return [] result = [] # 初始化当前近似组的最短字符串 current_shortest = lst[0] for s in lst[1:]: # 判断是否属于同一近似组:互相存在子串包含关系 if current_shortest in s or s in current_shortest: # 属于同一组,更新组内最短字符串 if len(s) < len(current_shortest): current_shortest = s else: # 进入新组,存储上一组的最短结果 result.append(current_shortest) current_shortest = s # 存入最后一组的最短结果 result.append(current_shortest) return result # 测试示例 liste = ['I am googling for the solution for an hour now', 'I am googling for the solution for an hour now --Sent via mail--', 'I am googling for the solution for an hour now --Sent via mail-- What are you doing?', 'Hello I am good thanks >> How are you?', 'Hello I am good thanks', 'Hello I am good thanks >>'] print(deduplicate_continuous_approx(liste))
运行输出:
['I am googling for the solution for an hour now', 'Hello I am good thanks']
适用通用文本相似度的场景
如果你的近似重复不是严格的子串包含,只是文本语义/字符相似度较高,可以基于编辑距离判断是否属于同一组,需要先安装依赖库python-Levenshtein:
from Levenshtein import ratio def deduplicate_continuous_similar(lst, similarity_threshold=0.8): if not lst: return [] result = [] current_group = [lst[0]] for s in lst[1:]: # 相似度高于阈值判定为同一近似组 if ratio(s, current_group[0]) >= similarity_threshold: current_group.append(s) else: # 取当前组最短字符串存入结果 result.append(min(current_group, key=len)) current_group = [s] result.append(min(current_group, key=len)) return result
你可以根据自己的实际数据调整similarity_threshold阈值,取值范围为0~1,数值越高判断同组的标准越严格。
内容的提问来源于stack exchange,提问作者soulwreckedyouth
相关产品推荐
相关产品推荐

