You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从bltindex.com折线图爬取数据并转为pandas DataFrame?

提取谷歌表格折线图数据并转换为Pandas DataFrame

解决方案步骤

  1. 定位目标Script标签:由于页面的nonce属性是动态生成的,我们通过脚本内容中的关键词google.visualization.DataTable来定位包含数据的标签。
  2. 提取并格式化数据:使用正则表达式匹配出DataTable的构造参数(JavaScript对象格式),将其转换为标准JSON格式后解析。
  3. 转换为DataFrame:从解析后的字典中提取表头和数据行,整理后构建Pandas DataFrame。

完整代码实现

import requests
from bs4 import BeautifulSoup
import re
import pandas as pd
import json

# 目标谷歌表格图表页面URL
url = 'https://docs.google.com/spreadsheets/d/e/2PACX-1vQG9TYlv8_LpCvO7EI3Y3s8MoxQEfOHTd3-EqccN5PoeHcdxraxZC0y8UWFx_2NnogVIIuk1i-phvFe/pubchart?oid=813038046&format=interactive'
html = requests.get(url)
soup = BeautifulSoup(html.content, 'html.parser')

# 查找包含图表数据的script标签
target_script = None
for script in soup.find_all('script'):
    if 'google.visualization.DataTable' in script.text:
        target_script = script
        break

if not target_script:
    raise ValueError("未找到包含数据的目标脚本标签")

# 正则提取DataTable的参数内容
data_pattern = re.compile(r'new google\.visualization\.DataTable\((.*?)\);', re.DOTALL)
data_match = data_pattern.search(target_script.text)
if not data_match:
    raise ValueError("无法从脚本中提取数据结构")

# 将JavaScript对象转换为可解析的JSON格式
raw_data_str = data_match.group(1)
# 替换单引号为双引号,移除JSON不兼容的 trailing commas
cleaned_data_str = raw_data_str.replace("'", "\"")
cleaned_data_str = re.sub(r',\s*([\]}])', r'\1', cleaned_data_str)

# 解析JSON数据
data_dict = json.loads(cleaned_data_str)

# 提取表头和数据行
column_headers = [col['label'] for col in data_dict['cols']]
raw_rows = [row['c'] for row in data_dict['rows']]

# 处理每行数据,提取单元格实际值
processed_rows = []
for row in raw_rows:
    processed_row = [cell['v'] for cell in row]
    processed_rows.append(processed_row)

# 构建DataFrame并转换日期列
df = pd.DataFrame(processed_rows, columns=column_headers)
df[column_headers[0]] = pd.to_datetime(df[column_headers[0]])

# 输出结果示例
print(df.head())

关键说明

  • 正则匹配逻辑:谷歌可视化图表的DataTable构造参数是结构化的JavaScript对象,我们通过正则捕获该对象的完整内容,确保不会遗漏数据。
  • 格式兼容处理:JavaScript对象允许使用单引号和尾逗号,这不符合JSON标准,因此需要替换单引号为双引号,并移除尾逗号才能被json.loads解析。
  • 数据提取:DataTable的cols字段存储表头信息,rows字段存储每行数据,每个单元格的v属性是实际的数值/日期值,提取后即可直接用于构建DataFrame。

内容的提问来源于stack exchange,提问作者Mr. Ivan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 10:55:17