课程作业:书籍加密程序超长字符串去标点性能问题求助
Ah, the classic "works great for small chunks but dies on full-length books" problem—total relatable pain point. Let’s figure out why your punctuation-stripping function is choking on Pride and Prejudice and fix it fast.
Why Your Current Function Is Slow
Chances are, your current implementation is doing something like iterating over every character in the text, then checking if it’s in a list of punctuation marks (e.g., if char in string.punctuation for each character). That’s an O(n*m) operation—for every character (n, which is millions in a full novel), you’re checking against every punctuation mark (m, ~30-40). That adds up really quick.
Optimized Solutions (Python Examples)
Let’s swap that inefficient loop for operations that run in C under the hood—they’ll handle the full book in milliseconds, not minutes.
Option 1: Pre-Compiled Regex (Clean and Fast)
Regex engines are optimized for bulk string operations, and compiling the pattern once (instead of every time you call the function) eliminates unnecessary overhead.
import re import string # Compile the regex ONCE, outside your function punctuation_pattern = re.compile(f'[{re.escape(string.punctuation)}]') def remove_punctuation(text): return punctuation_pattern.sub('', text)
- The
re.escape()ensures any punctuation with special regex meaning (like.or*) is treated as literal characters. - This does a single pass over the text to replace all punctuation with empty strings.
Option 2: Translation Table (Fastest Possible)
Python’s str.translate() uses a pre-built table to delete characters in one go, and it’s implemented directly in C—this is usually the fastest method for this task.
import string # Create the translation table ONCE, outside your function # This tells translate() to delete all characters in string.punctuation punct_trans_table = str.maketrans('', '', string.punctuation) def remove_punctuation(text): return text.translate(punct_trans_table)
- If you need to keep certain punctuation (like apostrophes in "Mr. Bennet" or "Elizabeth’s"), just adjust the punctuation set first:
# Keep apostrophes, remove everything else custom_punct = string.punctuation.replace("'", "") punct_trans_table = str.maketrans('', '', custom_punct)
Quick Pro Tips
- Don’t re-create your pattern/table inside the function: If you’re generating the regex or translation table every time you call
remove_punctuation, that’s adding unnecessary overhead. Define them once, outside the function. - Process in bulk: Avoid splitting the book into pages/chunks unless you have to—bulk operations are always faster than processing small pieces repeatedly.
- Test with a subset first: If you want to verify, run the function on the first 10k characters of the book to confirm it works, then scale up.
Either of these methods should handle the full text of Pride and Prejudice in well under a second—no more interrupt mode crashes!
内容的提问来源于stack exchange,提问作者D. Christ

