You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Plotly生成的HTML绘图文件中提取CSV数据?

从Plotly生成的HTML文件中提取数据生成CSV的方法

问题背景

你有一个通过Plotly生成的HTML绘图文件,希望从中提取数据生成与原DataFrame(如示例中的iris数据集)尽可能接近的CSV文件。示例中使用Plotly Express生成散点图并导出为HTML,数据以类似JavaScript字典的形式存储在HTML中。

解决方案

一、优先推荐:Python + Plotly官方工具解析

Plotly提供的plotly.io模块可直接读取HTML中的figure对象,无需手动解析JS代码,是最可靠的方法:

import plotly.io as pio
import pandas as pd

# 读取HTML中的第一个figure对象
fig = pio.read_html("plot.html")[0]

# 遍历所有trace,提取对应数据字段
df_list = []
for trace in fig.data:
    # 从trace中映射原DataFrame的字段:x对应sepal_width,y对应sepal_length,颜色对应petal_length
    trace_data = {
        "sepal_width": trace.x,
        "sepal_length": trace.y,
        "petal_length": trace.marker.color,
        "species": trace.name  # facet_col的分类信息会保存在trace的name中
    }
    df_list.append(pd.DataFrame(trace_data))

# 合并所有trace的数据
combined_df = pd.concat(df_list, ignore_index=True)

# 补充原数据中的species_id字段(通过species映射)
species_id_map = {"setosa": 1, "versicolor": 2, "virginica": 3}
combined_df["species_id"] = combined_df["species"].map(species_id_map)

# 若需要补充petal_width,可检查trace是否存储了该字段(示例中未在绘图中使用,需根据实际情况调整)

# 导出为CSV
combined_df.to_csv("extracted_iris_data.csv", index=False)

说明:

  • 该方法直接利用Plotly官方API解析HTML,兼容性最好,无需处理复杂JS语法。
  • 不同类型图表(如折线图、柱状图)的trace属性有差异,需根据实际绘图的映射关系调整字段提取逻辑。

二、备选方案:手动解析HTML中的JS数据

如果plotly.io无法读取HTML(比如HTML由旧版本Plotly生成),可通过正则表达式提取JS中的figure数据:

import re
import json
import pandas as pd

# 读取HTML文件内容
with open("plot.html", "r", encoding="utf-8") as f:
    html_content = f.read()

# 匹配HTML中的figure JSON数据(需根据实际HTML结构调整正则)
match = re.search(r'var\s+fig\s*=\s*({.*?});', html_content, re.DOTALL)
if not match:
    # 尝试另一种常见的存储位置
    match = re.search(r'window\.PLOTLY_ENV\.plotlyObjects\s*=\s*({.*?});', html_content, re.DOTALL)

if match:
    # 修复JS对象的语法差异(单引号转双引号,处理undefined)
    fig_json_str = match.group(1).replace("'", '"').replace("undefined", "null")
    fig_data = json.loads(fig_json_str)

    # 提取数据,逻辑同第一种方法
    df_list = []
    for trace in fig_data["data"]:
        df_list.append(pd.DataFrame({
            "sepal_width": trace["x"],
            "sepal_length": trace["y"],
            "petal_length": trace["marker"]["color"],
            "species": trace["name"]
        }))
    
    combined_df = pd.concat(df_list, ignore_index=True)
    combined_df["species_id"] = combined_df["species"].map({"setosa":1, "versicolor":2, "virginica":3})
    combined_df.to_csv("extracted_iris_data.csv", index=False)
else:
    print("未在HTML中找到可解析的figure数据")

三、非Python方案:浏览器开发者工具提取

如果不想写代码,可直接通过浏览器提取:

  • 打开目标HTML文件,按F12打开开发者工具
  • 切换到Console标签,输入Plotly.d3.select('.plotly').data()(或类似代码,根据图表结构调整),回车后即可看到所有trace数据
  • 右键点击数据结果,选择Copy object将数据复制为JSON格式
  • 用Excel、Google Sheets或本地CSV转换工具将JSON转为CSV,再手动补充缺失字段(如species_id)

内容的提问来源于stack exchange,提问作者user171780

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 23:01:14