PDF解析文本断句合并问题及Python代码优化求助
Fixing Your PDF Text Merging Code (Duplicate Entry Issue)
Let's break down why your current code is causing duplicate entries, then fix it step by step.
What's Wrong with the Original Code?
- Index Logic Flaw: You're iterating over every index
iand always accessingi-1, which means you'll process the same pair multiple times (e.g., wheni=1you merge 0+1, then wheni=2you try to merge 1+2—even though 1 was already part of the merged 0+1 entry). - Broken Condition Check: The
orindictlist[i-1][-1:] != ('.') or dictlist[i-1][-1:] != ('. ')is a logical mistake—this condition will always be true, since a string can't end with both.and.at the same time. You probably meant to useandhere, or adjust the check to target unfinished sentences. - Syntax Error:
new.append[dict_list[i]]uses square brackets instead of parentheses, which will throw an error.
Corrected Solution
Instead of iterating through each index and looking back, use a pointer to track your current position in the list. This way you can skip over elements that have already been merged:
def cleanlist(dictlist): new = [] i = 0 while i < len(dictlist): current = dictlist[i] # Check if we need to merge with the next element if i + 1 < len(dictlist): next_str = dictlist[i+1] # Case 1: Current ends with a hyphen (split word) if current.endswith('-'): # Remove the hyphen and merge with next string merged = current[:-1] + next_str new.append(merged) i += 2 # Skip the next element since we merged it continue # Case 2: Current doesn't end with a period (unfinished sentence) elif not current.endswith(('.', '. ')): merged = current + next_str new.append(merged) i += 2 continue # If no merge needed, add current and move to next new.append(current) i += 1 return new # Test it with your example test_list = ['I am a ', 'sentence.', 'Please join me toge-', 'ther. Thanks for your help.'] print(cleanlist(test_list)) # Output: ['I am a sentence.', 'Please join me together. Thanks for your help.']
How This Works
- We use a
whileloop with a pointerithat we manually increment, so we can skip elements once they're merged. - First, we check if there's a next element to merge with.
- For hyphenated splits: we remove the trailing hyphen before merging, then jump two indices ahead (since we've used both current and next).
- For unfinished sentences (no ending period): we merge directly and jump two indices.
- If no merge is needed, we add the current element and move to the next index.
This avoids duplicate entries because we never process the same element more than once.
内容的提问来源于stack exchange,提问作者JSC
相关产品推荐
相关产品推荐

