如何更优化地实现字符串数组重复元素置null(不复制数组、不硬编码长度且保持原顺序)
Great question! Your current implementation works, but it has two key areas for improvement: a slow O(n²) time complexity from nested loops, plus a potential fragility in edge cases with multiple duplicates. Let's walk through better, more efficient approaches that fix these issues.
First: Clarify the Core Requirement
Your original code sets both the first and subsequent duplicates to null. If that's exactly what you want (all instances of a repeated element become null), we can optimize that. If instead you want to keep the first occurrence and only null out later duplicates, we have a simpler, faster solution for that too.
1. Keep First Occurrence, Nullify Duplicates (Most Common Use Case)
This approach uses a HashSet to track elements we've already seen, allowing us to process the array in a single pass (O(n) average time complexity). This is drastically faster than nested loops for large arrays.
import java.util.Scanner; import java.util.HashSet; import java.util.Arrays; public class RemoveDuplicates { public static void main(String[] args) { Scanner sc = new Scanner(System.in); int arrSize = 10; String[] strings = new String[arrSize]; // Populate the array with user input for (int i = 0; i < arrSize; i++) { strings[i] = sc.nextLine(); } HashSet<String> seenElements = new HashSet<>(); for (int i = 0; i < strings.length; i++) { String current = strings[i]; if (current == null) continue; // Skip already nulled entries if (seenElements.contains(current)) { strings[i] = null; // Mark duplicate as null } else { seenElements.add(current); // Track first occurrence } } // Print the result Arrays.stream(strings).forEach(System.out::println); } }
Why this is better:
- Runs in O(n) average time instead of O(n²) — critical for arrays with hundreds or thousands of elements.
- Logic is straightforward and easy for other developers to follow.
- Avoids fragile dependencies on mid-loop variable state (like your original
checkvariable).
2. Nullify All Instances of Repeated Elements (Match Your Original Logic)
If you need every occurrence of a duplicate element to become null (not just later ones), use a HashMap to count occurrences first, then mark elements with a count > 1 as null. This still runs in O(n) time.
import java.util.Scanner; import java.util.HashMap; import java.util.Arrays; public class MarkAllDuplicatesAsNull { public static void main(String[] args) { Scanner sc = new Scanner(System.in); int arrSize = 10; String[] strings = new String[arrSize]; // Populate array for (int i = 0; i < arrSize; i++) { strings[i] = sc.nextLine(); } // Count how many times each element appears HashMap<String, Integer> elementCount = new HashMap<>(); for (String s : strings) { if (s != null) { elementCount.put(s, elementCount.getOrDefault(s, 0) + 1); } } // Nullify any element that appears more than once for (int i = 0; i < strings.length; i++) { String current = strings[i]; if (current != null && elementCount.get(current) > 1) { strings[i] = null; } } // Output the result Arrays.stream(strings).forEach(System.out::println); } }
Why this beats your original code:
- Still O(n) time, so it scales way better for large arrays.
- No risk of missing duplicates (your original code relies on the
checkvariable to keep comparing after settingstrings[i]to null, which is fragile). - Logic is explicit and easier to debug if edge cases pop up.
What's Less Ideal About the Original Approach?
- Slow Performance: Nested loops mean for an array of size n, you do ~n²/2 comparisons. For n=1000, that's 499,500 operations — compared to 1000 operations with the HashSet approach.
- Fragile Logic: The code depends on the
checkvariable retaining the original value after you setstrings[i]to null. While it works for simple cases, it's harder to maintain and could lead to bugs if someone modifies the code later.
内容的提问来源于stack exchange,提问作者Batiievskyi

