You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求提供Python脚本:从HTML三字段Sankey图提取数据导出至CSV

提取桑基图指定关联数据并导出CSV的Python脚本

实现思路

这类桑基图通常会在HTML的<script>标签中存储节点(nodes)和关联链接(links)的JSON数据,我们需要:

  • 解析HTML提取桑基图的核心数据
  • 定位中间字段中名为bibliographic coupling的节点
  • 筛选该节点对应的左侧参考文献关联和右侧主题关联
  • 将匹配结果导出为CSV文件

Python 实现脚本

import requests
from bs4 import BeautifulSoup
import json
import csv

# 1. 获取HTML内容(本地文件或远程地址,自行替换)
# 本地文件方式
with open("Sankey-Diagram.html", "r", encoding="utf-8") as f:
    html_content = f.read()

# 远程地址方式(如果需要从网络获取,取消注释下面两行)
# url = "你的HTML文件地址"
# html_content = requests.get(url).text

# 2. 解析HTML,提取桑基图数据
soup = BeautifulSoup(html_content, "html.parser")
# 找到包含桑基图数据的script标签(根据实际内容调整选择器)
script_tags = soup.find_all("script")
sankey_data = None

for script in script_tags:
    if "nodes" in script.text and "links" in script.text:
        # 提取JSON格式的数据部分(需要根据实际HTML的代码结构调整截取逻辑)
        script_text = script.text.strip()
        # 匹配数据结构,补全JSON格式
        start_idx = script_text.find("nodes:")
        end_idx = script_text.rfind("}")
        data_str = script_text[start_idx:end_idx].replace("nodes:", '"nodes":').replace("links:", '"links":')
        data_str = "{" + data_str + "}"
        try:
            sankey_data = json.loads(data_str)
            break
        except json.JSONDecodeError:
            continue

if not sankey_data:
    print("未找到桑基图数据")
    exit()

nodes = sankey_data["nodes"]
links = sankey_data["links"]

# 3. 定位中间节点 'bibliographic coupling' 的ID
middle_node_id = None
for node in nodes:
    if node.get("name") == "bibliographic coupling":
        middle_node_id = node["id"]
        break

if not middle_node_id:
    print("未找到目标中间节点")
    exit()

# 4. 筛选左侧关联(指向中间节点的链接)和右侧关联(从中间节点出发的链接)
left_relation_ids = {link["source"] for link in links if link["target"] == middle_node_id}
right_relation_ids = {link["target"] for link in links if link["source"] == middle_node_id}

# 构建节点ID到名称的映射
node_id_map = {node["id"]: node["name"] for node in nodes}

# 5. 整理数据:左侧参考文献 -> 中间节点 -> 右侧主题
result_rows = []
for left_id in left_relation_ids:
    left_name = node_id_map.get(left_id, "未知节点")
    for right_id in right_relation_ids:
        right_name = node_id_map.get(right_id, "未知节点")
        result_rows.append([left_name, "bibliographic coupling", right_name])

# 6. 导出到CSV文件
output_file = "sankey_extracted_data.csv"
with open(output_file, "w", newline="", encoding="utf-8") as csvfile:
    writer = csv.writer(csvfile)
    writer.writerow(["左侧参考文献", "中间字段", "右侧关联主题"])
    writer.writerows(result_rows)

print(f"数据已成功导出到 {output_file}")

注意事项

  • 若HTML中桑基图数据的存储结构与脚本假设不同,需要调整script标签的筛选逻辑和数据截取规则
  • 确保安装了所需依赖包:pip install requests beautifulsoup4

内容的提问来源于stack exchange,提问作者Shadow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 10:13:03