Python修改HTML替换font标签内容报AttributeError如何解决
报错原因
- 执行
i.find(text=re.compile(...))时,如果当前font标签内没有匹配当前n值的文本,find()方法会返回None,调用replace_with方法自然抛出AttributeError range(129)生成的是0128的整数,和你需求里1128的数字范围不符,多了无效的n=0遍历步骤- 写文件逻辑放在两层循环内部,每修改一条内容就写一次文件,存在大量多余IO操作,执行效率极低
最优解决方案
不需要嵌套range循环,直接用正则捕获文本中的数字部分,一次性替换所有符合格式的内容,不会出现匹配失败返回None的问题:
from bs4 import BeautifulSoup import re # 读取原HTML文件 with open("c:/users/dell/desktop/se2.html", "r", encoding="utf-8") as f: d = BeautifulSoup(f.read(), "html.parser") # 筛选所有符合颜色要求的font标签 target_fonts = d.find_all("font", {"color": "#FFFFFF"}) for font in target_fonts: raw_text = font.get_text(strip=True) if not raw_text: continue # 正则捕获原数字,仅替换名字部分 new_text = re.sub( r"PAT-204635 - LAICHE AHMED ILYES - Number (\d{1,3})", r"PAT-204635 - LAICHE MOHAMED ISLAM - Number \1", raw_text ) font.string.replace_with(new_text) # 全部修改完成后一次性写入文件,建议先输出到新文件验证后再覆盖原文件 with open("c:/users/dell/desktop/se2_modified.html", "wb") as f: f.write(d.prettify("utf-8"))
原有逻辑修复方案
如果你要保留原来的嵌套循环写法,增加非空判断即可:
import re from bs4 import BeautifulSoup with open("c:/users/dell/desktop/se2.html", "r", encoding="utf-8") as f: d = BeautifulSoup(f.read(), "html.parser") old = d.find_all("font", {"color": "#FFFFFF"}) # 调整range范围为1~128 for n in range(1, 129): for i in old: matched = i.find(text=re.compile(f"PAT-204635 - LAICHE AHMED ILYES - Number {n}")) # 匹配成功才执行替换 if matched: matched.replace_with(f"PAT-204635 - LAICHE MOHAMED ISLAM - Number {n}") # 统一写文件 with open ("c:/users/dell/desktop/se2.html","wb") as ff: ff.write(d.prettify("utf-8"))
内容的提问来源于stack exchange,提问作者serhani ahmed saned
相关产品推荐
相关产品推荐

