非正则表达式实现标点拆分与单词拼接的代码调试问题
Let's break down what's wrong with your current code and fix it step by step to meet your requirement.
What's Wrong with Your Current Code
- Incorrect splitting logic: Using
sentence.split('_')splits the entire input at every underscore—including the spaces around them. For your input, this creates["Democr ", "acy , is overrat ", "ed ."], and joining these gives messy text with extra spaces instead of properly combining word fragments. - Misplaced, useless condition checks: The variable
cstarts as an empty string, soc not in punctuationis always true. Theelifblock never runs unless you modifyc—and even if it did, looping through every character in the sentence and splitting on eachcwould just spam random lists, which explains your 29-list output. - No punctuation spacing handling: Your code doesn't address the core need to attach punctuation directly to the preceding word (like turning
acy ,intoacy,).
Step-by-Step Fix
We'll skip regex entirely and focus on two key tasks: combining word fragments separated by _ (space-underscore-space), and cleaning up punctuation spacing to match your expected output.
Here's the corrected code:
sentence = input().strip() # Define all punctuation marks we need to adjust punctuation = {"!", "\"", "#", "$", "%", "&", "'", "(", ")", "*", "+", ",", "-", ".", "/", ":", ";", "<", "=", ">", "?", "@", "[", "\\", "]", "^", "`", "{", "|", "}", "~"} # Step 1: Merge word fragments split by " _ " processed = sentence.replace(" _ ", "") # Step 2: Remove spaces before punctuation to attach them to words for p in punctuation: processed = processed.replace(f" {p}", p) print(processed)
How It Works
- Combining word fragments: The
replace(" _ ", "")call targets the exact pattern separating your word parts. For your input,Democr _acybecomesDemocracyandoverrat _edbecomesoverrated. - Cleaning punctuation: We loop through each punctuation mark and replace any instance of
[space][punctuation]with just the punctuation. This fixesacy ,toacy,anded .toed.while leaving normal spaces between words (likeis overrated) intact.
Testing with your input Democr _acy , is overrat _ed . gives the expected output:
Democracy, is overrated.
内容的提问来源于stack exchange,提问作者0phase
相关产品推荐
相关产品推荐

