使用BS4替换文本内容并保留标签时遇AttributeError问题求助
问题与解决方案
我有一个包含text-transform: uppercase;样式的HTML文档,需要在保留所有HTML标签的前提下,将该样式作用范围内的所有文本转换为大写。目前尝试了两种方法,但都存在问题:
- 第一种方法能实现文本转大写,但会直接移除所有HTML标签,只保留纯文本
- 第二种方法试图遍历文本节点进行替换,却抛出
AttributeError: 'NoneType' object has no attribute 'next_element'错误
示例代码如下:
from bs4 import BeautifulSoup, NavigableString, Tag import re html = ''' <div style="text-transform: uppercase;"> Foo0 <font>Foo0</font> <div>Foo1 <div>Foo2</div> </div> </div> ''' upper_patt = re.compile('(?i)text-transform:\s*uppercase') # 可行但会移除HTML标签 soup = BeautifulSoup(html, "html.parser") for node in soup.find_all(attrs={'style': upper_patt}): node.replace_with(node.text.upper()) # 报错方案 soup = BeautifulSoup(html, "html.parser") for node in soup.find_all(attrs={'style': upper_patt}): for txt in node.strings: txt.replace_with(txt.upper())
错误原因
第二种方案报错是因为node.strings返回的是动态生成器,在遍历过程中调用replace_with()修改DOM结构,会打乱生成器的迭代顺序,导致后续迭代时获取到None对象,进而触发属性错误。
正确解决方案
先把需要处理的文本节点一次性收集到列表中,再遍历列表进行替换,避免遍历过程中修改DOM结构影响迭代:
from bs4 import BeautifulSoup, NavigableString import re html = ''' <div style="text-transform: uppercase;"> Foo0 <font>Foo0</font> <div>Foo1 <div>Foo2</div> </div> </div> ''' upper_patt = re.compile('(?i)text-transform:\s*uppercase') soup = BeautifulSoup(html, "html.parser") # 遍历所有匹配样式的节点 for container in soup.find_all(attrs={'style': upper_patt}): # 收集所有文本节点到列表(避免遍历生成器时修改DOM导致的问题) text_nodes = [node for node in container.descendants if isinstance(node, NavigableString)] # 遍历列表替换文本为大写 for node in text_nodes: # 保留原空白格式,仅转换字母大小写 node.replace_with(node.upper()) # 输出处理后的HTML print(soup.prettify())
处理结果
运行上述代码后,输出的HTML会保留所有标签结构,且目标文本全部转为大写:
<div style="text-transform: uppercase;"> FOO0 <font>FOO0</font> <div>FOO1 <div>FOO2</div> </div> </div>
如果需要过滤掉纯空白的文本节点(比如多余的换行、空格),可以在收集节点时添加判断:
text_nodes = [node for node in container.descendants if isinstance(node, NavigableString) and node.strip()]
内容的提问来源于stack exchange,提问作者Peter
相关产品推荐
相关产品推荐

