如何提取两个Header标签间的数据并生成键值对?Header标签下文本提取方法
解决方案
一、提取Header标签下的文本
你的现有代码已经能获取Header的文本,这里补充优化写法和不同场景的提取方式:
- 直接获取标签下所有文本(含子标签内容):
item.text.strip()(加strip()去除前后空白) - 仅提取直接子节点的文本(排除嵌套标签内容):用
' '.join(item.stripped_strings),自动过滤空白字符与换行 - 简化现有代码的写法:
import re from bs4 import BeautifulSoup soup = BeautifulSoup(page.content, 'html.parser') # 用列表推导式简化提取逻辑 htag_list = [{tag.name: tag.text.strip()} for tag in soup.find_all(re.compile('^h[1-6]'))] print(htag_list)
二、提取两个Header之间的内容并生成键值对
核心逻辑是遍历每个Header,捕获它到下一个Header之间的所有内容,以下是两种实用实现方式:
方法1:利用find_next_siblings + 终止条件
import re from bs4 import BeautifulSoup soup = BeautifulSoup(page.content, 'html.parser') headers = soup.find_all(re.compile('^h[1-6]')) header_content_pairs = {} for idx, header in enumerate(headers): # 用当前Header的文本作为键 key = header.text.strip() content = [] # 遍历当前Header之后的所有兄弟节点 for sibling in header.find_next_siblings(): # 遇到下一个Header就停止遍历 if re.match('^h[1-6]$', sibling.name): break # 过滤空文本后加入内容列表 sibling_text = sibling.text.strip() if sibling_text: content.append(sibling_text) # 将内容拼接成字符串存入字典 header_content_pairs[key] = ' '.join(content) # 输出结果 for key, val in header_content_pairs.items(): print(f"{key}: {val}")
方法2:用next_sibling逐个遍历(精细控制)
适合处理复杂节点结构(比如内容嵌套在div中):
import re from bs4 import BeautifulSoup soup = BeautifulSoup(page.content, 'html.parser') headers = soup.find_all(re.compile('^h[1-6]')) header_content_pairs = {} for idx, header in enumerate(headers): key = header.text.strip() content = [] current_sibling = header.next_sibling while current_sibling is not None: # 遇到下一个Header就终止循环 if current_sibling.name and re.match('^h[1-6]$', current_sibling.name): break # 处理标签节点的文本 if hasattr(current_sibling, 'text'): text = current_sibling.text.strip() if text: content.append(text) # 处理纯文本节点 else: str_text = str(current_sibling).strip() if str_text: content.append(str_text) current_sibling = current_sibling.next_sibling header_content_pairs[key] = ' '.join(content) print(header_content_pairs)
注意事项
- 如果Header仅存在于
div.conWrap容器内,记得限定查找范围:soup.select_one("div.conWrap").find_all(re.compile('^h[1-6]')),避免提取页面其他区域的Header - 若需保留HTML格式而非纯文本,可将
sibling.text替换为str(sibling),存储完整节点的HTML代码
内容的提问来源于stack exchange,提问作者Akshay Toranagatti
相关产品推荐
相关产品推荐

