You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取Flourish页面script标签中的数据?

用BeautifulSoup提取Flourish嵌入页面的数据

核心思路是通过requests获取页面源码,用BeautifulSoup定位目标script标签,再用正则表达式提取_Flourish_data_column_names和_Flourish_data后的JSON格式数据,最后转成Python可处理的结构。

步骤及代码实现

1. 导入依赖库

import requests
from bs4 import BeautifulSoup
import re
import json

2. 提取主页面script[4]中的数据

对应你提供的Xpath:/html/body/script[4](注意Python列表索引从0开始,所以取第4个标签要索引3)

# 替换为目标网站的实际URL
main_url = "https://example.com/target-page"
# 模拟浏览器请求头,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 获取页面源码
response = requests.get(main_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 定位到第4个script标签
script_tag = soup.find_all("script")[3]
script_content = script_tag.string

# 提取列名数据
column_names_match = re.search(r'var _Flourish_data_column_names = (.*?);', script_content)
if column_names_match:
    column_names = json.loads(column_names_match.group(1))
    print("列名数据:", column_names)

# 提取核心数据
flourish_data_match = re.search(r'_Flourish_data = (.*?);', script_content)
if flourish_data_match:
    flourish_data = json.loads(flourish_data_match.group(1))
    print("Flourish核心数据:", flourish_data)

3. 提取iframe中的目标script数据

对应Xpath:/html/body/main/div[1]/div/div/div[2]/div/script,需要先获取iframe的链接再请求:

# 从主页面找到iframe标签
iframe = soup.find("iframe")
if iframe:
    iframe_url = iframe["src"]
    # 处理相对路径(如果iframe的src是相对地址,需要拼接主域名)
    # 示例:如果主域名是https://example.com,iframe src是/flourish-embed,就拼接成完整URL
    # iframe_url = f"{main_url.split('//')[0]}//{main_url.split('//')[1].split('/')[0]}{iframe_url}"
    
    # 请求iframe页面
    iframe_response = requests.get(iframe_url, headers=headers)
    iframe_soup = BeautifulSoup(iframe_response.text, "html.parser")
    
    # 根据Xpath层级定位目标div下的script标签
    target_div = iframe_soup.select_one("main > div:nth-of-type(1) > div > div > div:nth-of-type(2) > div")
    if target_div:
        iframe_script = target_div.find("script")
        if iframe_script:
            iframe_script_content = iframe_script.string
            
            # 提取列名
            column_names_match = re.search(r'var _Flourish_data_column_names = (.*?);', iframe_script_content)
            if column_names_match:
                column_names = json.loads(column_names_match.group(1))
                print("iframe中的列名数据:", column_names)
            
            # 提取核心数据
            flourish_data_match = re.search(r'_Flourish_data = (.*?);', iframe_script_content)
            if flourish_data_match:
                flourish_data = json.loads(flourish_data_match.group(1))
                print("iframe中的核心数据:", flourish_data)

注意事项

  • 如果正则匹配不到数据,检查script标签的内容结构,可能需要调整正则的结束符(比如有些数据后面是}而非分号,可把正则里的;改成})
  • 若网站有反爬机制,可添加Cookie到请求头,或使用代理IP
  • 页面结构可能变化,必要时通过script标签的id、type等属性精准定位,避免依赖索引或层级

内容的提问来源于stack exchange,提问作者Brandon Peffer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 17:10:11