Python正则表达式代码报错:TypeError: expected string or buffer 求助
Hey there, let's work through this regex error you're running into. You've already confirmed lorem_ipsum is a string, so we can rule out the most obvious culprit—but there are still a few common missteps that could trigger that TypeError. Let's break down the fixes for your two tasks, plus how to avoid this issue going forward.
First: Diagnose the Root Cause
Even if you started with a string, double-check these two things before diving into fixes:
- Did you mix up regex method parameters? Python's
remodule requires the pattern first, then the string (e.g.,re.findall(pattern, string), not the other way around). Swapping these will throw the exact error you're seeing. - Did you accidentally reassign
lorem_ipsum? If you ran something likelorem_ipsum = re.split(r'\s+', lorem_ipsum)earlier, that variable is now a list—not a string. Printtype(lorem_ipsum)right before your regex code to confirm.
Fix 1: Count Non-Alphanumeric Characters (Expected: 144)
Here are two reliable ways to get that count, both guaranteed to work if lorem_ipsum is a valid string:
Method 1: Use re.findall to Match All Non-Alnum Characters
import re # Quick sanity check to confirm type assert isinstance(lorem_ipsum, str), "lorem_ipsum got changed to a non-string type!" # Match any character that's NOT a letter or number non_alnum_chars = re.findall(r'[^a-zA-Z0-9]', lorem_ipsum) non_alnum_count = len(non_alnum_chars) print(non_alnum_count) # Should output 144
Method 2: Replace Alnum Characters and Count Remaining Length
This is slightly more efficient for large strings:
import re # Replace all letters/numbers with empty string, count what's left non_alnum_count = len(re.sub(r'[a-zA-Z0-9]', '', lorem_ipsum)) print(non_alnum_count)
Fix 2: Find All "……" Content
I'm assuming you mean either finding all instances of the ellipsis itself, or content wrapped between ellipses. Here are both scenarios:
Scenario A: Find All Ellipsis Instances
For Chinese ellipses (……) or English multi-dot sequences:
# Match all Chinese ellipses all_ellipses = re.findall(r'……', lorem_ipsum) print(all_ellipses) # Match all sequences of 3+ English dots (e.g., "...", "....") all_english_ellipses = re.findall(r'\.{3,}', lorem_ipsum) print(all_english_ellipses)
Scenario B: Find Content Wrapped Between Ellipses
If you need text that's surrounded by …… (e.g., ……sample text……), use non-greedy matching to avoid capturing too much:
# Capture content between ellipses (non-greedy to stop at the next ellipsis) wrapped_content = re.findall(r'……(.*?)……', lorem_ipsum, re.DOTALL) # re.DOTALL makes "." match newlines too, remove if you don't need that print(wrapped_content)
Final Quick Checks
- Always print
type(lorem_ipsum)right before your regex code to confirm it's still a string. - If you're using any other regex methods (like
re.matchorre.search), double-check their parameter order—pattern first, string second.
内容的提问来源于stack exchange,提问作者t.hill

