You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式移除字符串标点后,如何还原标点至原位置?

How to Restore Punctuation After Modifying Text (Python)

Got it, let's work through this problem step by step. The key here is that you need to track exactly where punctuation was located before removing it—that way, you can slot it back into the right spot once you've edited your words.

Here's a straightforward, reliable approach with code examples:

Step 1: Split the Original String into "Tokens"

First, we'll break down your original string into a list of every individual component: words, spaces, and punctuation marks. This preserves the exact order and position of everything in the original text.

import re

s = "I want to remove all the punctuation, and then put it, back where it was."

# Split into tokens: matches words, spaces, and punctuation
tokens = re.findall(r'\w+|\s+|[^\w\s]', s)
# Result: ['I', ' ', 'want', ' ', 'to', ' ', 'remove', ' ', 'all', ' ', 'the', ' ', 'punctuation', ',', ' ', 'and', ' ', 'then', ' ', 'put', ' ', 'it', ',', ' ', 'back', ' ', 'where', ' ', 'it', ' ', 'was', '.']

Step 2: Extract and Modify the Words

Next, we'll pull out just the words from the token list, make your desired changes, and keep track of where those words were positioned in the original token list.

# Record indices of all word tokens
word_indices = [i for i, token in enumerate(tokens) if re.match(r'\w+', token)]

# Extract the words to modify
words = [tokens[i] for i in word_indices]

# Example modification: capitalize every word (replace this with your own edits!)
words = [word.capitalize() for word in words]

Step 3: Restore Punctuation by Reinserting Modified Words

Now, we'll take the original token list, replace the old words with your modified ones, and join everything back into a single string. The punctuation and spaces stay exactly where they were!

# Create a copy of the original tokens to avoid overwriting
reconstructed_tokens = tokens.copy()

# Replace old words with modified ones in their original positions
for word_idx, token_idx in enumerate(word_indices):
    reconstructed_tokens[token_idx] = words[word_idx]

# Join tokens back into a string
final_string = ''.join(reconstructed_tokens)
print(final_string)
# Output: "I Want To Remove All The Punctuation, And Then Put It, Back Where It Was."

Why This Works

  • The regex r'\w+|\s+|[^\w\s]' ensures we don't miss any part of the original text—every word, space, and punctuation mark gets captured as a separate token.
  • By tracking the indices of word tokens, we know exactly where to put your modified words back without disrupting punctuation or spacing.
  • This method is flexible: you can make any changes to the words list (like replacing words, altering case, or filtering) and the punctuation will always snap back to its original spot.

内容的提问来源于stack exchange,提问作者hippocampus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 00:44:04