如何用正则表达式移除文件中非连续出现的重复单词?
Got it, I see the issue here—your current regex only handles consecutive duplicate words, but you want to eliminate every instance of a word after its first occurrence, no matter what's in between. The problem with regex alone here is that it can't "remember" words it's already matched earlier in the text—it only works with sequential patterns.
Let's fix this by combining regex (to properly identify words and punctuation) with a Java collection that tracks which words we've already seen. Here's how to adjust your code:
First, Why Your Original Regex Fails
Your regex \b(\p{IsAlphabetic}+)(\s+\1\b)+ looks for a word, followed by one or more spaces and the exact same word again. That works for back-to-back duplicates like university university, but it can't catch university ... university because there's other text breaking the sequence. Regex doesn't have a built-in way to track previously matched words across non-consecutive positions.
The Solution: Track Seen Words with a Collection
We'll use a LinkedHashSet to keep track of words we've already encountered—this collection preserves the order of first occurrence while automatically rejecting duplicates. We'll also use regex to properly split out words (and handle punctuation attached to them, like radio.).
Modified Full Code
Replace your line processing logic with this updated version:
File dir = new File("C:/Users/Arnoldas/workspace/uplo/"); String source = dir.getCanonicalPath() + File.separator + "Output.txt"; String dest = dir.getCanonicalPath() + File.separator + "Final.txt"; File fin = new File(source); FileInputStream fis = new FileInputStream(fin); BufferedReader in = new BufferedReader(new InputStreamReader(fis, "UTF-8")); OutputStreamWriter fstream = new OutputStreamWriter(new FileOutputStream(dest, true), "UTF-8"); BufferedWriter out = new BufferedWriter(fstream); // Regex to match words (alphabetic part + optional trailing punctuation) String wordRegex = "\\b(\\p{IsAlphabetic}+)(\\p{Punct}*)\\b"; Pattern wordPattern = Pattern.compile(wordRegex, Pattern.CASE_INSENSITIVE); String aLine; while ((aLine = in.readLine()) != null) { Matcher wordMatcher = wordPattern.matcher(aLine); LinkedHashSet<String> seenWords = new LinkedHashSet<>(); StringBuilder processedLine = new StringBuilder(); int lastPosition = 0; while (wordMatcher.find()) { // Add any non-word text between the last match and current match processedLine.append(aLine.substring(lastPosition, wordMatcher.start())); String fullWord = wordMatcher.group(); String lowercaseAlpha = wordMatcher.group(1).toLowerCase(); // Normalize for case-insensitive check // Only keep the word if we haven't seen its alphabetic part before if (!seenWords.contains(lowercaseAlpha)) { seenWords.add(lowercaseAlpha); processedLine.append(fullWord); } lastPosition = wordMatcher.end(); } // Add any remaining text after the last word processedLine.append(aLine.substring(lastPosition)); out.write(processedLine.toString()); out.newLine(); out.flush(); } // Don't forget to close your streams! in.close(); out.close();
How This Works
- Word Matching: The regex
\b(\p{IsAlphabetic}+)(\p{Punct}*)\\bcaptures two parts of each word: the alphabetic core (e.g.,beganfrombegan.) and any trailing punctuation (e.g.,.). - Case Insensitivity: We convert the alphabetic core to lowercase so
Universityanduniversityare treated as duplicates. - Duplicate Tracking:
LinkedHashSetstores each unique lowercase alphabetic core. If a word's core is already in the set, we skip adding it to the processed line. - Preserving Formatting: We keep all non-word text (spaces, punctuation not attached to words) intact, so your original sentence structure stays the same—only duplicate words are removed.
Testing this with your example input:
Input:
The university of Hawaii university began using began radio.
Output:The university of Hawaii began using radio.
Which is exactly what you're looking for!
内容的提问来源于stack exchange,提问作者Arnas Arnelis

