You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现去除近似重复项并保留最长条目问题求助

Fixing Your Near-Duplicate Removal Code (Keep Longest Entry)

First, let's break down the issues in your existing code that are causing it to fail:

  • Incorrect loop syntax: for x in len(ListofLine): doesn't iterate over line indices. You need for x in range(len(ListofLine)) instead.
  • Index out-of-bounds error: When x is the last index of the list, x+1 will go beyond the list length, throwing an IndexError.
  • Wrong comparison operator: = is for assignment; you need == to check if the first words match.
  • Incomplete logic: Your code only compares adjacent lines, not all lines that share the same starting words. It also closes the files mid-loop, which stops any further writes.
  • Unhandled edge cases: Empty lines or lines with no words will cause split()[0] to throw an IndexError.

Corrected Implementation

This code follows your original idea: grouping lines by their first N words (we'll use 2 as an example, adjust as needed), then keeping the longest line in each group.

def keep_longest_near_duplicates(input_file, output_file, num_start_words=2):
    # Dictionary to track groups: key = tuple of first N words, value = longest line so far
    line_groups = {}

    # Safely read input file using with statement (auto-closes file)
    with open(input_file, 'r', encoding='utf-8') as f1:
        for line in f1:
            stripped_line = line.strip()
            if not stripped_line:  # Skip empty lines
                continue
            
            # Split line into words, handle lines with fewer than N words
            words = stripped_line.split()
            if len(words) < num_start_words:
                group_key = tuple(words)
            else:
                group_key = tuple(words[:num_start_words])
            
            # Update group with the longest line
            if group_key not in line_groups or len(stripped_line) > len(line_groups[group_key]):
                line_groups[group_key] = stripped_line + '\n'  # Add back newline for formatting

    # Write longest lines to output file
    with open(output_file, 'w', encoding='utf-8') as f2:
        for longest_line in line_groups.values():
            f2.write(longest_line)

# Run the function with your file names
keep_longest_near_duplicates('a.txt', 'n.txt')

Key Improvements

  • Uses with statements to handle files safely (no manual closing needed)
  • Groups lines by the first N words (adjust num_start_words to match your definition of "near-duplicate")
  • Skips empty lines and handles short lines gracefully
  • Efficiently tracks the longest line per group using a dictionary
  • Avoids index errors entirely by iterating directly over lines instead of indices

Input/Output Example

Input (a.txt):

Hello world this is a long line
Hello world short
Python programming basics
Python programming for advanced users
Test line

Output (n.txt):

Hello world this is a long line
Python programming for advanced users
Test line

内容的提问来源于stack exchange,提问作者Harry Hau

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:22:31