BeautifulSoup遍历h3标签间type_grid div并构建嵌套字典问题
问题1解答
你写的print([child.name for child in heading.parent.children])基本是正确实现,但是存在小瑕疵:BeautifulSoup的children迭代器会包含标签之间的空白、换行等NavigableString类型节点,这类节点的name属性为None,会出现在结果列表中。
如果要严格只保留HTML标签对象列表,更严谨的写法是先判断节点是否为Tag类型:
from bs4 import Tag # 仅获取标签对象列表 tag_list = [child for child in heading.parent.children if isinstance(child, Tag)] # 仅获取标签名列表 tag_name_list = [child.name for child in tag_list]
你补充的代码里加了child.name in allowed_tags的过滤条件,会自动过滤掉None值,实际运行也能得到符合要求的结果。
问题2解答
针对这种平级无嵌套的标签结构,更简洁高效的方案是状态遍历法:遍历父节点下所有子标签时,用变量记录当前正在归属的h3值,碰到对应标签就更新状态,碰到type__grid类div就解析内容存入对应字典层级,一次遍历就能完成组装,不需要提前记录索引、也不需要复杂的while循环,示例代码如下:
from bs4 import BeautifulSoup, Tag # 假设已完成页面解析拿到soup对象 h2_tag = soup.find("h2", class_="type__heading") h2_key = h2_tag.get_text(strip=True) result = {h2_key: {}} current_h3 = None # 遍历h2父节点的所有子标签 for child in h2_tag.parent.children: # 跳过非标签节点 if not isinstance(child, Tag): continue # 匹配h3标签,更新当前归属的h3,初始化对应字典 if child.name == "h3" and "type__purpose-heading" in child.get("class", []): current_h3 = child.get_text(strip=True) result[h2_key][current_h3] = {} # 匹配type__grid类div,解析内容存入当前h3对应的字典 elif child.name == "div" and "type__grid" in child.get("class", []): if not current_h3: continue h4_tag = child.find("h4", class_="type__category-heading") ul_tag = child.find("ul", class_="type__category-items") if not h4_tag or not ul_tag: continue h4_key = h4_tag.get_text(strip=True) li_values = [li.get_text(strip=True) for li in ul_tag.find_all("li")] result[h2_key][current_h3][h4_key] = li_values
内容的提问来源于stack exchange,提问作者KArrow'sBest
相关产品推荐
相关产品推荐

