如何修改BeautifulSoup中soup.html.text或soup.html.string的值?
解决BeautifulSoup中无法直接赋值
soup.html.text的问题 在BeautifulSoup里,soup.html.text是只读的动态属性——它的作用是拼接并返回当前节点下所有子文本节点的内容,并没有提供对应的赋值方法,所以直接执行soup.html.text = 'some text'或soup.html.text = soup.html.text都会触发AttributeError。
要实现替换<html>标签内全部文本的需求,有两种常用方案:
方案1:清空所有子节点后添加新文本
如果不需要保留原有的标签结构,直接清空<html>下的所有内容,再添加新文本:
from bs4 import BeautifulSoup html_doc = """ <html> <head><title>Test Page</title></head> <body><p>Sample text</p></body> </html> """ soup = BeautifulSoup(html_doc, 'html.parser') # 清空html标签下的所有子节点 for child in soup.html.children: child.extract() # 添加新文本内容 soup.html.append(soup.new_string('some text')) print(soup.prettify())
执行后,<html>标签内只会保留指定的文本,原有标签结构会被移除。
方案2:递归替换所有文本节点(保留标签结构)
如果需要保留原有的HTML标签结构,只替换所有文本内容,可以通过递归遍历所有节点,替换其中的文本节点:
def replace_all_text(node, new_text): # 遍历当前节点的所有子节点(用list避免遍历过程中节点变化的问题) for child in list(node.children): if child.name is None: # 判断是否为文本节点(无标签名的节点) child.replace_with(soup.new_string(new_text)) else: # 递归处理子标签节点 replace_all_text(child, new_text) # 调用函数替换html下所有文本 replace_all_text(soup.html, 'some text') print(soup.prettify())
执行后,原有标签结构会被保留,所有标签内的文本都会被替换为指定内容。
内容的提问来源于stack exchange,提问作者Abdullah Zaman Babar
相关产品推荐
相关产品推荐

