如何移除数组中被包含的子串元素?三种场景详解
如何避免数组中出现重复语义的元素?
以下为三种场景说明:
- 场景一:希望输出保留全部3个元素,包括
Phil。Phil是Phillies的子串,但不应被移除(仅当短字符串是长字符串中独立单词时才判定为重复语义)。 - 场景二:希望输出保留"New York Mets"、"Philadelphia Phillies",但移除
Phillies,因为它是"Philadelphia Phillies"中的一个独立单词。 - 场景三:希望输出仅保留
Cristiano Ronaldo,因为Cristiano是该长串的一部分(作为独立单词存在于长字符串中),应被移除。
原有的remove_subwords函数仅通过子串匹配判断,无法区分"子串"和"独立单词",会导致场景一误删Phil,因此需要针对不同场景调整判断逻辑:
针对场景一的解决方案
需要判断短字符串是否是长字符串中的独立单词(通过空格分割后是否存在该单词),而不是单纯的子串:
def keep_non_word_substrings(players): result = [] # 先把每个元素拆成单词集合,方便后续判断 word_sets = [set(item.split()) for item in players] for idx, player in enumerate(players): # 检查当前元素是否是其他元素中的独立单词 is_word_sub = any(player in word_set and player != other for other, word_set in zip(players, word_sets) if player != other) if not is_word_sub: result.append(player) return result # 测试场景一 football_players = ["New York Mets", "Philadelphia Phillies", "Phil"] print(keep_non_word_substrings(football_players)) # 输出: ['New York Mets', 'Philadelphia Phillies', 'Phil']
针对场景二、三的解决方案
如果需要移除作为独立单词存在于长字符串中的短元素,可以直接使用基于单词集合的判断:
def remove_word_substrings(players): result = [] word_sets = [set(item.split()) for item in players] for idx, player in enumerate(players): # 检查当前元素是否是其他元素的子单词 is_sub_word = any(player in word_set and len(player.split()) < len(other.split()) for other, word_set in zip(players, word_sets) if player != other) if not is_sub_word: result.append(player) return result # 测试场景二 football_players = ["New York Mets", "Philadelphia Phillies", "Phillies"] print(remove_word_substrings(football_players)) # 输出: ['New York Mets', 'Philadelphia Phillies'] # 测试场景三 football_players = ["Cristiano", "Cristiano Ronaldo"] print(remove_word_substrings(football_players)) # 输出: ['Cristiano Ronaldo']
通用可配置方案
如果需要根据不同场景切换判断逻辑,可以写一个带参数的函数:
def filter_semantic_duplicates(players, keep_non_word_subs=False): result = [] word_sets = [set(item.split()) for item in players] for player in players: should_remove = False for other in players: if player == other: continue if keep_non_word_subs: # 场景一规则:仅移除作为独立单词的子串 if player in word_sets[players.index(other)]: should_remove = True break else: # 场景二、三规则:移除作为独立单词或前缀的短串 if player in word_sets[players.index(other)] or (len(player) < len(other) and other.startswith(player + " ")): should_remove = True break if not should_remove: result.append(player) return result # 场景一调用 print(filter_semantic_duplicates(["New York Mets", "Philadelphia Phillies", "Phil"], keep_non_word_subs=True)) # 场景二调用 print(filter_semantic_duplicates(["New York Mets", "Philadelphia Phillies", "Phillies"])) # 场景三调用 print(filter_semantic_duplicates(["Cristiano", "Cristiano Ronaldo"]))
内容的提问来源于stack exchange,提问作者user352290
相关产品推荐
相关产品推荐

