You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python通过Soup提取script标签中的totalPage数值?

解决方案
  • 核心思路:先定位包含pageList的script标签,再用正则精准匹配totalPage对应的数值
  • 具体代码实现:
from bs4 import BeautifulSoup
import re

# 替换为你的目标HTML内容
html = """
<!-- 示例HTML结构 -->
<script>
var pageList = {
    totalPage: 12,
    currentPage: 1,
    pageSize: 10
};
</script>
"""

# 解析HTML
soup = BeautifulSoup(html, 'html.parser')

# 筛选包含pageList关键字的script标签
target_script = soup.find('script', string=re.compile(r'pageList'))

# 匹配totalPage的数值(兼容空格、引号等常见格式)
if target_script:
    match_result = re.search(r'pageList\s*=\s*{[^}]*totalPage\s*:\s*["\']?(\d+)["\']?', target_script.string)
    if match_result:
        total_page = int(match_result.group(1))
        print(total_page)  # 输出12
  • 关键细节说明:
    • 正则表达式pageList\s*=\s*{[^}]*totalPage\s*:\s*["\']?(\d+)["\']?的设计考虑了多种格式情况:
      • \s*匹配任意数量的空格(包括换行、制表符),避免因代码排版导致匹配失败
      • [^}]*匹配{到totalPage之间的任意内容(除了闭合大括号)
      • ["\']?匹配可选的单/双引号,兼容数值被引号包裹的情况
      • (\d+)捕获目标数字,作为最终提取结果

内容的提问来源于stack exchange,提问作者gndps

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 21:31:11