You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup/lxml提取script标签前后文本失效解决方案

提取script标签相邻裸文本的可落地方案

两种方案提取失败的核心原因:

  • script前后的目标内容是没有被HTML标签包裹的裸文本节点
  • next_siblings、following-sibling默认遍历逻辑如果没有显式指定匹配文本节点,会直接跳过这类内容,只返回带标签的元素节点,自然拿不到结果

先给出测试用的目标HTML结构(和问题描述的页面对应):

<span class="fw-normal small">
    UTC
    <script>/* 时区渲染逻辑 */</script>
    +05:30
</span>
<script>/* 交易时间渲染逻辑 */</script>
20:28

方案1:BeautifulSoup 实现

遍历节点时显式识别NavigableString类型的纯文本节点,过滤掉script标签和无意义的空白换行即可:

from bs4 import BeautifulSoup, NavigableString, Tag

# 替换为实际页面HTML源码
html = """<上述测试HTML片段>"""
soup = BeautifulSoup(html, "html.parser")

# 定位目标span容器
target_span = soup.select_one("span.fw-normal.small")

# 提取span内script前后的时区文本
tz_parts = []
for node in target_span.contents:
    if isinstance(node, NavigableString):
        clean_text = node.strip()
        if clean_text:
            tz_parts.append(clean_text)
tz_content = " ".join(tz_parts) # 输出:UTC +05:30

# 提取同级script后的交易时间文本
trade_time = ""
for sibling in target_span.next_siblings:
    # 定位紧跟span的script标签
    if isinstance(sibling, Tag) and sibling.name == "script":
        # 找script后第一个非空文本节点
        for node in sibling.next_siblings:
            if isinstance(node, NavigableString):
                clean_text = node.strip()
                if clean_text:
                    trade_time = clean_text
                    break
        break

print(tz_content, trade_time)

方案2:lxml etree + XPath 实现

XPath路径中显式指定匹配text()文本节点,用normalize-space()过滤空白内容:

from lxml import etree

# 替换为实际页面HTML源码
html = """<上述测试HTML片段>"""
tree = etree.HTML(html)

# 提取span内script前后的时区文本
span_text_nodes = tree.xpath('//span[@class="fw-normal small"]/text()[normalize-space()]')
tz_content = " ".join([t.strip() for t in span_text_nodes]) # 输出:UTC +05:30

# 提取同级script后的交易时间文本
trade_time = tree.xpath(
    '//span[@class="fw-normal small"]/following-sibling::script[1]/following-sibling::text()[normalize-space()][1]'
)[0].strip() # 输出:20:28

print(tz_content, trade_time)

避坑提示:不要直接用find_next_sibling()、find_next()这类默认只匹配元素标签的方法,这类方法不会返回裸文本节点;遍历过程中必须对文本做strip()处理,过滤掉代码缩进、换行产生的无意义空字符串。

内容的提问来源于stack exchange,提问作者vineet singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 05:09:20