请求提供Python脚本:从HTML三字段Sankey图提取数据导出至CSV
提取桑基图指定关联数据并导出CSV的Python脚本
实现思路
这类桑基图通常会在HTML的<script>标签中存储节点(nodes)和关联链接(links)的JSON数据,我们需要:
- 解析HTML提取桑基图的核心数据
- 定位中间字段中名为
bibliographic coupling的节点 - 筛选该节点对应的左侧参考文献关联和右侧主题关联
- 将匹配结果导出为CSV文件
Python 实现脚本
import requests from bs4 import BeautifulSoup import json import csv # 1. 获取HTML内容(本地文件或远程地址,自行替换) # 本地文件方式 with open("Sankey-Diagram.html", "r", encoding="utf-8") as f: html_content = f.read() # 远程地址方式(如果需要从网络获取,取消注释下面两行) # url = "你的HTML文件地址" # html_content = requests.get(url).text # 2. 解析HTML,提取桑基图数据 soup = BeautifulSoup(html_content, "html.parser") # 找到包含桑基图数据的script标签(根据实际内容调整选择器) script_tags = soup.find_all("script") sankey_data = None for script in script_tags: if "nodes" in script.text and "links" in script.text: # 提取JSON格式的数据部分(需要根据实际HTML的代码结构调整截取逻辑) script_text = script.text.strip() # 匹配数据结构,补全JSON格式 start_idx = script_text.find("nodes:") end_idx = script_text.rfind("}") data_str = script_text[start_idx:end_idx].replace("nodes:", '"nodes":').replace("links:", '"links":') data_str = "{" + data_str + "}" try: sankey_data = json.loads(data_str) break except json.JSONDecodeError: continue if not sankey_data: print("未找到桑基图数据") exit() nodes = sankey_data["nodes"] links = sankey_data["links"] # 3. 定位中间节点 'bibliographic coupling' 的ID middle_node_id = None for node in nodes: if node.get("name") == "bibliographic coupling": middle_node_id = node["id"] break if not middle_node_id: print("未找到目标中间节点") exit() # 4. 筛选左侧关联(指向中间节点的链接)和右侧关联(从中间节点出发的链接) left_relation_ids = {link["source"] for link in links if link["target"] == middle_node_id} right_relation_ids = {link["target"] for link in links if link["source"] == middle_node_id} # 构建节点ID到名称的映射 node_id_map = {node["id"]: node["name"] for node in nodes} # 5. 整理数据:左侧参考文献 -> 中间节点 -> 右侧主题 result_rows = [] for left_id in left_relation_ids: left_name = node_id_map.get(left_id, "未知节点") for right_id in right_relation_ids: right_name = node_id_map.get(right_id, "未知节点") result_rows.append([left_name, "bibliographic coupling", right_name]) # 6. 导出到CSV文件 output_file = "sankey_extracted_data.csv" with open(output_file, "w", newline="", encoding="utf-8") as csvfile: writer = csv.writer(csvfile) writer.writerow(["左侧参考文献", "中间字段", "右侧关联主题"]) writer.writerows(result_rows) print(f"数据已成功导出到 {output_file}")
注意事项
- 若HTML中桑基图数据的存储结构与脚本假设不同,需要调整
script标签的筛选逻辑和数据截取规则 - 确保安装了所需依赖包:
pip install requests beautifulsoup4
内容的提问来源于stack exchange,提问作者Shadow
相关产品推荐
相关产品推荐

