如何过滤string vector中存在完整子串匹配的元素
Got it, let's tackle this problem head-on. The goal is to take your input string vector and strip out any element that acts as a full substring of another element in the same list. Here's a clear, practical solution with explanations tailored to your example:
Core Logic Breakdown
For every string in the vector, we need to ask: Is there another string (that's longer than it) which contains it as a complete substring? If yes, we filter it out. If no, we keep it in the final result.
Example Implementation (C++)
Since you referenced a string vector, I'll use C++ for the code example—it aligns directly with the data structure you mentioned:
#include <vector> #include <string> std::vector<std::string> filterSubstringElements(const std::vector<std::string>& inputVec) { std::vector<std::string> filteredVec; for (const auto& currentStr : inputVec) { bool shouldFilter = false; // Compare the current string to every other element for (const auto& otherStr : inputVec) { // Skip comparing the string to itself if (¤tStr == &otherStr) continue; // Check if the other string is longer AND contains currentStr as a substring if (otherStr.size() > currentStr.size() && otherStr.find(currentStr) != std::string::npos) { shouldFilter = true; break; // No need to check further once we find a match } } // Keep the string only if it's not a substring of any longer element if (!shouldFilter) { filteredVec.push_back(currentStr); } } return filteredVec; }
How It Works With Your Input
Let's verify against your specific input to confirm it matches your expected output:
- "I love Mangoes.": Filtered out (it’s a substring of "I love Mangoes and Apples.")
- "I love Mangoes and Apples.": Filtered out (it’s a substring of "I love Mangoes and Apples and Strawberries.")
- "Apples are good for health": Kept (no longer element contains it)
- "I live in America": Kept (no longer element contains it)
- "I love Mangoes and Apples and Strawberries.": Kept (it’s the longest in its related group, no other element contains it)
- "Mangoes and Apples.": Filtered out (it’s a substring of "I love Mangoes and Apples.")
- "Mangoes and Apples and Honey": Kept (no longer element contains it)
The resulting vector is exactly what you described.
Quick Notes
- If you're using another language (like Python), the logic stays identical—just adjust the syntax (e.g., use Python's
inoperator for substring checks). - This approach runs in O(n² * m) time (n = number of strings, m = average string length). For small vectors like your example, this is totally efficient. For larger datasets, you could optimize by sorting strings by length first to cut down on unnecessary comparisons.
内容的提问来源于stack exchange,提问作者ALEX MATHEW

