You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则表达式代码报错:TypeError: expected string or buffer 求助

Fixing the "TypeError: expected string or buffer" in Your Regex Code

Hey there, let's work through this regex error you're running into. You've already confirmed lorem_ipsum is a string, so we can rule out the most obvious culprit—but there are still a few common missteps that could trigger that TypeError. Let's break down the fixes for your two tasks, plus how to avoid this issue going forward.

First: Diagnose the Root Cause

Even if you started with a string, double-check these two things before diving into fixes:

  • Did you mix up regex method parameters? Python's re module requires the pattern first, then the string (e.g., re.findall(pattern, string), not the other way around). Swapping these will throw the exact error you're seeing.
  • Did you accidentally reassign lorem_ipsum? If you ran something like lorem_ipsum = re.split(r'\s+', lorem_ipsum) earlier, that variable is now a list—not a string. Print type(lorem_ipsum) right before your regex code to confirm.

Fix 1: Count Non-Alphanumeric Characters (Expected: 144)

Here are two reliable ways to get that count, both guaranteed to work if lorem_ipsum is a valid string:

Method 1: Use re.findall to Match All Non-Alnum Characters

import re

# Quick sanity check to confirm type
assert isinstance(lorem_ipsum, str), "lorem_ipsum got changed to a non-string type!"

# Match any character that's NOT a letter or number
non_alnum_chars = re.findall(r'[^a-zA-Z0-9]', lorem_ipsum)
non_alnum_count = len(non_alnum_chars)
print(non_alnum_count)  # Should output 144

Method 2: Replace Alnum Characters and Count Remaining Length

This is slightly more efficient for large strings:

import re

# Replace all letters/numbers with empty string, count what's left
non_alnum_count = len(re.sub(r'[a-zA-Z0-9]', '', lorem_ipsum))
print(non_alnum_count)

Fix 2: Find All "……" Content

I'm assuming you mean either finding all instances of the ellipsis itself, or content wrapped between ellipses. Here are both scenarios:

Scenario A: Find All Ellipsis Instances

For Chinese ellipses (……) or English multi-dot sequences:

# Match all Chinese ellipses
all_ellipses = re.findall(r'……', lorem_ipsum)
print(all_ellipses)

# Match all sequences of 3+ English dots (e.g., "...", "....")
all_english_ellipses = re.findall(r'\.{3,}', lorem_ipsum)
print(all_english_ellipses)

Scenario B: Find Content Wrapped Between Ellipses

If you need text that's surrounded by …… (e.g., ……sample text……), use non-greedy matching to avoid capturing too much:

# Capture content between ellipses (non-greedy to stop at the next ellipsis)
wrapped_content = re.findall(r'……(.*?)……', lorem_ipsum, re.DOTALL)
# re.DOTALL makes "." match newlines too, remove if you don't need that
print(wrapped_content)

Final Quick Checks

  • Always print type(lorem_ipsum) right before your regex code to confirm it's still a string.
  • If you're using any other regex methods (like re.match or re.search), double-check their parameter order—pattern first, string second.

内容的提问来源于stack exchange,提问作者t.hill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:06:50