You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python变位词校验函数应用于整句去重场景?

实现方案

核心思路

按照「拆分句子→特征标记→过滤去重→拼接输出」的流程处理即可,核心是给每个单词生成唯一的变位词特征(和你现有anagram函数的判断逻辑完全对齐),通过特征快速判断是否属于重复的变位词组。

实现代码

可直接复用你已有的变位词判断函数,完整实现如下:

# 你现有的变位词判断函数
def anagram(a, b):
  if len(a)==len(b) and sorted(a)==sorted(b):
    return True
  else:
    return False

def filter_anagram_sentence(sentence: str, remove_all_group: bool = False) -> str:
    # 按空格拆分句子为独立单词列表
    words = sentence.split()
    # 统计每个变位词组的出现次数
    feature_map = {}
    word_features = []
    for word in words:
        # 生成变位词唯一特征:排序后的字符串,和anagram判断逻辑一致
        feature = ''.join(sorted(word))
        word_features.append(feature)
        feature_map[feature] = feature_map.get(feature, 0) + 1
    
    result = []
    seen_features = set()
    for idx, word in enumerate(words):
        current_feat = word_features[idx]
        if remove_all_group:
            # 匹配你的示例需求:只要变位词组出现次数≥2,该组所有单词全部删除
            if feature_map[current_feat] < 2:
                result.append(word)
        else:
            # 常规去重需求:同一变位词组只保留第一个出现的单词
            if current_feat not in seen_features:
                seen_features.add(current_feat)
                result.append(word)
    # 拼接为完整句子返回
    return ' '.join(result)

调用示例

test_sentence = 'hello vola alvo my name is ...'
# 按你给出的示例需求调用(删除所有互为变位词的重复项)
output = filter_anagram_sentence(test_sentence, remove_all_group=True)
print(output) 
# 输出结果:hello my name is ...

可选优化:大小写不敏感适配

如果需要匹配你提到的voLa和alVo判断为变位词的需求,只需修改特征生成逻辑,统一转为小写/大写后再排序即可:

# 将特征生成行修改为如下代码即可
feature = ''.join(sorted(word.lower()))

内容的提问来源于stack exchange,提问作者repulsive2020

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 19:45:03