You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何过滤string vector中存在完整子串匹配的元素

Filter Out Strings That Are Complete Substrings of Other Elements in a Vector

Got it, let's tackle this problem head-on. The goal is to take your input string vector and strip out any element that acts as a full substring of another element in the same list. Here's a clear, practical solution with explanations tailored to your example:

Core Logic Breakdown

For every string in the vector, we need to ask: Is there another string (that's longer than it) which contains it as a complete substring? If yes, we filter it out. If no, we keep it in the final result.

Example Implementation (C++)

Since you referenced a string vector, I'll use C++ for the code example—it aligns directly with the data structure you mentioned:

#include <vector>
#include <string>

std::vector<std::string> filterSubstringElements(const std::vector<std::string>& inputVec) {
    std::vector<std::string> filteredVec;

    for (const auto& currentStr : inputVec) {
        bool shouldFilter = false;
        // Compare the current string to every other element
        for (const auto& otherStr : inputVec) {
            // Skip comparing the string to itself
            if (&currentStr == &otherStr) continue;
            // Check if the other string is longer AND contains currentStr as a substring
            if (otherStr.size() > currentStr.size() && otherStr.find(currentStr) != std::string::npos) {
                shouldFilter = true;
                break; // No need to check further once we find a match
            }
        }
        // Keep the string only if it's not a substring of any longer element
        if (!shouldFilter) {
            filteredVec.push_back(currentStr);
        }
    }

    return filteredVec;
}

How It Works With Your Input

Let's verify against your specific input to confirm it matches your expected output:

  • "I love Mangoes.": Filtered out (it’s a substring of "I love Mangoes and Apples.")
  • "I love Mangoes and Apples.": Filtered out (it’s a substring of "I love Mangoes and Apples and Strawberries.")
  • "Apples are good for health": Kept (no longer element contains it)
  • "I live in America": Kept (no longer element contains it)
  • "I love Mangoes and Apples and Strawberries.": Kept (it’s the longest in its related group, no other element contains it)
  • "Mangoes and Apples.": Filtered out (it’s a substring of "I love Mangoes and Apples.")
  • "Mangoes and Apples and Honey": Kept (no longer element contains it)

The resulting vector is exactly what you described.

Quick Notes

  • If you're using another language (like Python), the logic stays identical—just adjust the syntax (e.g., use Python's in operator for substring checks).
  • This approach runs in O(n² * m) time (n = number of strings, m = average string length). For small vectors like your example, this is totally efficient. For larger datasets, you could optimize by sorting strings by length first to cut down on unnecessary comparisons.

内容的提问来源于stack exchange,提问作者ALEX MATHEW

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:42:10