Python实现去除近似重复项并保留最长条目问题求助
Fixing Your Near-Duplicate Removal Code (Keep Longest Entry)
First, let's break down the issues in your existing code that are causing it to fail:
- Incorrect loop syntax:
for x in len(ListofLine):doesn't iterate over line indices. You needfor x in range(len(ListofLine))instead. - Index out-of-bounds error: When
xis the last index of the list,x+1will go beyond the list length, throwing anIndexError. - Wrong comparison operator:
=is for assignment; you need==to check if the first words match. - Incomplete logic: Your code only compares adjacent lines, not all lines that share the same starting words. It also closes the files mid-loop, which stops any further writes.
- Unhandled edge cases: Empty lines or lines with no words will cause
split()[0]to throw anIndexError.
Corrected Implementation
This code follows your original idea: grouping lines by their first N words (we'll use 2 as an example, adjust as needed), then keeping the longest line in each group.
def keep_longest_near_duplicates(input_file, output_file, num_start_words=2): # Dictionary to track groups: key = tuple of first N words, value = longest line so far line_groups = {} # Safely read input file using with statement (auto-closes file) with open(input_file, 'r', encoding='utf-8') as f1: for line in f1: stripped_line = line.strip() if not stripped_line: # Skip empty lines continue # Split line into words, handle lines with fewer than N words words = stripped_line.split() if len(words) < num_start_words: group_key = tuple(words) else: group_key = tuple(words[:num_start_words]) # Update group with the longest line if group_key not in line_groups or len(stripped_line) > len(line_groups[group_key]): line_groups[group_key] = stripped_line + '\n' # Add back newline for formatting # Write longest lines to output file with open(output_file, 'w', encoding='utf-8') as f2: for longest_line in line_groups.values(): f2.write(longest_line) # Run the function with your file names keep_longest_near_duplicates('a.txt', 'n.txt')
Key Improvements
- Uses
withstatements to handle files safely (no manual closing needed) - Groups lines by the first N words (adjust
num_start_wordsto match your definition of "near-duplicate") - Skips empty lines and handles short lines gracefully
- Efficiently tracks the longest line per group using a dictionary
- Avoids index errors entirely by iterating directly over lines instead of indices
Input/Output Example
Input (a.txt):
Hello world this is a long line
Hello world short
Python programming basics
Python programming for advanced users
Test line
Output (n.txt):
Hello world this is a long line
Python programming for advanced users
Test line
内容的提问来源于stack exchange,提问作者Harry Hau
相关产品推荐
相关产品推荐

