将函数返回的vector赋值时触发Segmentation fault问题求助
Hey there! Let's break down why you're hitting that segmentation fault and how to fix your approach.
First, the Direct Cause of the Segmentation Fault
Look at the loop in find_duplicates_with_sorting:
for( unsigned int i = 0; i < strings.size() - 1; ++i ) {
When you call find_duplicates_with_sorting(dups) in the hashtable function, if dups is empty (which it will be the first time you detect a duplicate string), strings.size() is 0. Since size() returns an unsigned integer, subtracting 1 from 0 results in a massive positive number (unsigned integers wrap around instead of going negative). This makes your loop condition always true, and you'll keep trying to access strings[i] and strings[i+1]—positions that don't exist in an empty vector. This out-of-bounds memory access is exactly what triggers the segmentation fault.
Next, the Logical Misstep
Your hashtable function's logic is misaligned with your goal here. The line dups = find_duplicates_with_sorting(dups); doesn't serve any purpose for collecting new duplicates:
find_duplicates_with_sortingis meant to find duplicates in the input vector, but you're passing itdups—the list of duplicates you've already found. This doesn't help you add the new duplicate string you just detected.- You already confirmed the current string isn't in
dupswithstd::find(dups.begin(), dups.end(), *i) == dups.end()—so you don't need another function call; you just need to add the string directly todups.
Fix Directions
- Patch the Segfault Immediately: In
find_duplicates_with_sorting, add a check to handle empty or single-element vectors before entering the loop. Ifstrings.size() <= 1, there can't be any duplicates, so return an emptydupsvector right away. - Fix the Hashtable Function Logic: Replace the problematic line with code that adds the current duplicate string (
*i) todups. Since you already verified it's not indups, a simpledups.push_back(*i);will work. - Optional Efficiency Upgrade: Your hashtable uses
std::unordered_map<std::string, std::string>but you only need to track whether a string has been seen before. Astd::unordered_set<std::string>would be more efficient (no unnecessary key-value pairs). Alternatively, usestd::unordered_map<std::string, int>to track counts—this lets you skip thestd::findcheck entirely: just add the string todupswhen its count increases from 1 to 2.
内容的提问来源于stack exchange,提问作者Dayem Saeed

