You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取属性值包含非换行空格(nbsp)的HTML标签?

解决BeautifulSoup匹配包含非换行空格的aria-label问题

嘿,这个坑我之前踩过!问题根源在于BeautifulSoup对HTML实体的解析逻辑:当它加载并解析HTML页面时,会自动把 这类实体转换成对应的Unicode非换行空格字符(U+00A0),而你代码里写的是字符串"Siguiente ",相当于在匹配实体的字面量文本,自然找不到目标标签啦。

给你几个靠谱的解决方案:

方案1:直接使用Unicode非换行空格

把代码里的 替换成Unicode转义符\u00A0,直接匹配解析后的字符:

a_url = soup.find_all('a', {"aria-label": "Siguiente\u00A0"})

方案2:借助BeautifulSoup的实体常量

BeautifulSoup内置了实体常量,你可以直接调用它来拼接字符串,避免手动写转义符:

from bs4 import BeautifulSoup
from bs4.element import Entities

soup = BeautifulSoup(your_html_content, 'html.parser')
a_url = soup.find_all('a', {"aria-label": "Siguiente" + Entities.nbsp})

方案3:用正则表达式模糊匹配(适合不确定空白类型的场景)

如果不确定目标标签的aria-label里是普通空格还是非换行空格,可以用正则的\s(注意开启Unicode模式)来匹配所有空白字符:

import re
from bs4 import BeautifulSoup

soup = BeautifulSoup(your_html_content, 'html.parser')
a_url = soup.find_all('a', {"aria-label": re.compile(r"Siguiente\s", re.UNICODE)})

这三种方法里,方案1最直接高效,要是页面里的 没有被其他特殊处理,用它基本就能搞定啦。

内容的提问来源于stack exchange,提问作者gian ruiz archenti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 18:27:41