You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF解析文本断句合并问题及Python代码优化求助

Fixing Your PDF Text Merging Code (Duplicate Entry Issue)

Let's break down why your current code is causing duplicate entries, then fix it step by step.

What's Wrong with the Original Code?

  1. Index Logic Flaw: You're iterating over every index i and always accessing i-1, which means you'll process the same pair multiple times (e.g., when i=1 you merge 0+1, then when i=2 you try to merge 1+2—even though 1 was already part of the merged 0+1 entry).
  2. Broken Condition Check: The or in dictlist[i-1][-1:] != ('.') or dictlist[i-1][-1:] != ('. ') is a logical mistake—this condition will always be true, since a string can't end with both . and . at the same time. You probably meant to use and here, or adjust the check to target unfinished sentences.
  3. Syntax Error: new.append[dict_list[i]] uses square brackets instead of parentheses, which will throw an error.

Corrected Solution

Instead of iterating through each index and looking back, use a pointer to track your current position in the list. This way you can skip over elements that have already been merged:

def cleanlist(dictlist):
    new = []
    i = 0
    while i < len(dictlist):
        current = dictlist[i]
        # Check if we need to merge with the next element
        if i + 1 < len(dictlist):
            next_str = dictlist[i+1]
            # Case 1: Current ends with a hyphen (split word)
            if current.endswith('-'):
                # Remove the hyphen and merge with next string
                merged = current[:-1] + next_str
                new.append(merged)
                i += 2  # Skip the next element since we merged it
                continue
            # Case 2: Current doesn't end with a period (unfinished sentence)
            elif not current.endswith(('.', '. ')):
                merged = current + next_str
                new.append(merged)
                i += 2
                continue
        # If no merge needed, add current and move to next
        new.append(current)
        i += 1
    return new

# Test it with your example
test_list = ['I am a ', 'sentence.', 'Please join me toge-', 'ther. Thanks for your help.']
print(cleanlist(test_list))
# Output: ['I am a sentence.', 'Please join me together. Thanks for your help.']

How This Works

  • We use a while loop with a pointer i that we manually increment, so we can skip elements once they're merged.
  • First, we check if there's a next element to merge with.
  • For hyphenated splits: we remove the trailing hyphen before merging, then jump two indices ahead (since we've used both current and next).
  • For unfinished sentences (no ending period): we merge directly and jump two indices.
  • If no merge is needed, we add the current element and move to the next index.

This avoids duplicate entries because we never process the same element more than once.

内容的提问来源于stack exchange,提问作者JSC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:43:41