BeautifulSoup/lxml提取script标签前后文本失效解决方案
提取script标签相邻裸文本的可落地方案
两种方案提取失败的核心原因:
- script前后的目标内容是没有被HTML标签包裹的裸文本节点
next_siblings、following-sibling默认遍历逻辑如果没有显式指定匹配文本节点,会直接跳过这类内容,只返回带标签的元素节点,自然拿不到结果
先给出测试用的目标HTML结构(和问题描述的页面对应):
<span class="fw-normal small"> UTC <script>/* 时区渲染逻辑 */</script> +05:30 </span> <script>/* 交易时间渲染逻辑 */</script> 20:28
方案1:BeautifulSoup 实现
遍历节点时显式识别NavigableString类型的纯文本节点,过滤掉script标签和无意义的空白换行即可:
from bs4 import BeautifulSoup, NavigableString, Tag # 替换为实际页面HTML源码 html = """<上述测试HTML片段>""" soup = BeautifulSoup(html, "html.parser") # 定位目标span容器 target_span = soup.select_one("span.fw-normal.small") # 提取span内script前后的时区文本 tz_parts = [] for node in target_span.contents: if isinstance(node, NavigableString): clean_text = node.strip() if clean_text: tz_parts.append(clean_text) tz_content = " ".join(tz_parts) # 输出:UTC +05:30 # 提取同级script后的交易时间文本 trade_time = "" for sibling in target_span.next_siblings: # 定位紧跟span的script标签 if isinstance(sibling, Tag) and sibling.name == "script": # 找script后第一个非空文本节点 for node in sibling.next_siblings: if isinstance(node, NavigableString): clean_text = node.strip() if clean_text: trade_time = clean_text break break print(tz_content, trade_time)
方案2:lxml etree + XPath 实现
XPath路径中显式指定匹配text()文本节点,用normalize-space()过滤空白内容:
from lxml import etree # 替换为实际页面HTML源码 html = """<上述测试HTML片段>""" tree = etree.HTML(html) # 提取span内script前后的时区文本 span_text_nodes = tree.xpath('//span[@class="fw-normal small"]/text()[normalize-space()]') tz_content = " ".join([t.strip() for t in span_text_nodes]) # 输出:UTC +05:30 # 提取同级script后的交易时间文本 trade_time = tree.xpath( '//span[@class="fw-normal small"]/following-sibling::script[1]/following-sibling::text()[normalize-space()][1]' )[0].strip() # 输出:20:28 print(tz_content, trade_time)
避坑提示:不要直接用
find_next_sibling()、find_next()这类默认只匹配元素标签的方法,这类方法不会返回裸文本节点;遍历过程中必须对文本做strip()处理,过滤掉代码缩进、换行产生的无意义空字符串。
内容的提问来源于stack exchange,提问作者vineet singh
相关产品推荐
相关产品推荐

