Python提取HTML全量文本并识别文本是否加粗的实现方法
实现思路
核心逻辑:对每一个提取到的文本节点,从其父节点开始向上遍历整条DOM祖先链,只要任意一个祖先节点命中粗体判定规则,就标记该文本为粗体,否则为非粗体。
判定时需要覆盖两类规则:
- 文本被
<b>标签直接或间接包裹 - 任意包裹文本的标签,style属性中设置
font-weight值为bold或700
注意不要硬匹配style属性的完整字符串,实际HTML中style可能包含多余空格、和其他样式属性混写,需要拆分样式键值对做匹配,避免漏判。
完整实现代码
from bs4 import BeautifulSoup import pandas as pd def check_text_bold(text_node): # 向上遍历所有祖先标签 for parent_tag in text_node.parents: # 规则1:被<b>标签包裹 if parent_tag.name == "b": return True # 规则2:标签设置了粗体font-weight样式 style_attr = parent_tag.get("style", "") if not style_attr: continue # 拆分style条目,兼容空格、多样式混写场景 style_entries = [entry.strip() for entry in style_attr.split(";") if entry.strip()] for entry in style_entries: if ":" not in entry: continue prop, value = [seg.strip().lower() for seg in entry.split(":", 1)] if prop == "font-weight" and value in ("bold", "700"): return True return False if __name__ == "__main__": # 替换为你的HTML文件路径 html_file = "test.html" with open(html_file, "r", encoding="utf-8") as f: soup = BeautifulSoup(f.read(), "html.parser") # 提取所有文本节点 all_text_nodes = soup.find_all(text=True, recursive=True) res = {"Text": [], "Bold": []} for node in all_text_nodes: text = node.strip() # 跳过纯空白的换行、缩进节点 if not text: continue res["Text"].append(text) res["Bold"].append("yes" if check_text_bold(node) else "no") # 生成DataFrame df = pd.DataFrame(res) print(df)
测试效果说明
用题目给出的测试HTML样例运行代码,会输出共15条文本记录,其中bold text (8)、bold text (10)、bold text (12)、bold text (14)4条记录的Bold字段值为yes,其余为no,完全匹配预期结果。
扩展提示
- 如果需要识别
<strong>标签默认粗体的场景,只需要在规则1的判断中增加parent_tag.name == "strong"的条件即可 - 如果需要兼容font-weight为500以上数值即判定粗体的业务规则,只需要调整value判断的逻辑,把数值做类型转换后判断阈值即可
- 如果HTML中存在通过CSS类设置粗体的场景,可以额外引入CSS解析库读取对应类的样式属性,再追加到判定逻辑中
内容的提问来源于stack exchange,提问作者constiii
相关产品推荐
相关产品推荐

