You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

soup.get_text()获取的文本无法被re.search匹配的问题求助

问题原因与解决方法

为什么正则匹配失败?

你输出的salary_range里,数字之间的空格不是普通空格(U+0020),而是非断空格(NBSP,Unicode编码U+00A0)。而你写的正则表达式r"^(\d{0,2} ?\d{3})"里的?只匹配普通空格,自然找不到对应内容,re.search返回None,调用.group(1)就会抛出AttributeError。

解决方法

方法1:先替换非断空格为普通空格

在正则匹配前,把字符串里的非断空格替换成普通空格:

salary_range = soup2.get_text().strip()
# 替换非断空格为普通空格
salary_range = salary_range.replace('\u00A0', ' ')
bottom_salary = re.search(r"^(\d{0,2} ?\d{3})", salary_range).group(1)
print(bottom_salary)  # 输出"10 000"

方法2:修改正则匹配非断空格

让正则同时匹配普通空格和非断空格:

bottom_salary = re.search(r"^(\d{0,2}[\s\u00A0]?\d{3})", salary_range).group(1)

这里[\s\u00A0]表示匹配任意空白字符(包括普通空格)或者非断空格。

方法3:跳过空格直接提取数字(更简洁)

不需要正则,直接拆分字符串后提取所有数字字符:

# 按分隔符"–"拆分,取前半部分
bottom_part = salary_range.split('–')[0]
# 提取所有数字字符拼接
bottom_salary_int = ''.join([c for c in bottom_part if c.isdigit()])
print(bottom_salary_int)  # 输出10000

内容的提问来源于stack exchange,提问作者Mikołaj Przygoda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 09:34:50