You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除数组中被包含的子串元素?三种场景详解

如何避免数组中出现重复语义的元素?

以下为三种场景说明:

  • 场景一:希望输出保留全部3个元素,包括Phil。Phil是Phillies的子串,但不应被移除(仅当短字符串是长字符串中独立单词时才判定为重复语义)。
  • 场景二:希望输出保留"New York Mets"、"Philadelphia Phillies",但移除Phillies,因为它是"Philadelphia Phillies"中的一个独立单词。
  • 场景三:希望输出仅保留Cristiano Ronaldo,因为Cristiano是该长串的一部分(作为独立单词存在于长字符串中),应被移除。

原有的remove_subwords函数仅通过子串匹配判断,无法区分"子串"和"独立单词",会导致场景一误删Phil,因此需要针对不同场景调整判断逻辑:

针对场景一的解决方案

需要判断短字符串是否是长字符串中的独立单词(通过空格分割后是否存在该单词),而不是单纯的子串:

def keep_non_word_substrings(players):
    result = []
    # 先把每个元素拆成单词集合,方便后续判断
    word_sets = [set(item.split()) for item in players]
    for idx, player in enumerate(players):
        # 检查当前元素是否是其他元素中的独立单词
        is_word_sub = any(player in word_set and player != other 
                         for other, word_set in zip(players, word_sets) if player != other)
        if not is_word_sub:
            result.append(player)
    return result

# 测试场景一
football_players = ["New York Mets", "Philadelphia Phillies", "Phil"]
print(keep_non_word_substrings(football_players))
# 输出: ['New York Mets', 'Philadelphia Phillies', 'Phil']

针对场景二、三的解决方案

如果需要移除作为独立单词存在于长字符串中的短元素,可以直接使用基于单词集合的判断:

def remove_word_substrings(players):
    result = []
    word_sets = [set(item.split()) for item in players]
    for idx, player in enumerate(players):
        # 检查当前元素是否是其他元素的子单词
        is_sub_word = any(player in word_set and len(player.split()) < len(other.split())
                         for other, word_set in zip(players, word_sets) if player != other)
        if not is_sub_word:
            result.append(player)
    return result

# 测试场景二
football_players = ["New York Mets", "Philadelphia Phillies", "Phillies"]
print(remove_word_substrings(football_players))
# 输出: ['New York Mets', 'Philadelphia Phillies']

# 测试场景三
football_players = ["Cristiano", "Cristiano Ronaldo"]
print(remove_word_substrings(football_players))
# 输出: ['Cristiano Ronaldo']

通用可配置方案

如果需要根据不同场景切换判断逻辑,可以写一个带参数的函数:

def filter_semantic_duplicates(players, keep_non_word_subs=False):
    result = []
    word_sets = [set(item.split()) for item in players]
    for player in players:
        should_remove = False
        for other in players:
            if player == other:
                continue
            if keep_non_word_subs:
                # 场景一规则:仅移除作为独立单词的子串
                if player in word_sets[players.index(other)]:
                    should_remove = True
                    break
            else:
                # 场景二、三规则:移除作为独立单词或前缀的短串
                if player in word_sets[players.index(other)] or (len(player) < len(other) and other.startswith(player + " ")):
                    should_remove = True
                    break
        if not should_remove:
            result.append(player)
    return result

# 场景一调用
print(filter_semantic_duplicates(["New York Mets", "Philadelphia Phillies", "Phil"], keep_non_word_subs=True))
# 场景二调用
print(filter_semantic_duplicates(["New York Mets", "Philadelphia Phillies", "Phillies"]))
# 场景三调用
print(filter_semantic_duplicates(["Cristiano", "Cristiano Ronaldo"]))

内容的提问来源于stack exchange,提问作者user352290

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 10:27:34