You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup无法抓取本地HTML文件表格数据求解决方案

How to Extract Cell Ranger Summary Table Data from web_summary.html

Let's fix your script to pull that summary table data correctly. The main issues with your original code are:

  1. You're targeting the wrong element to find the data script
  2. The JSON extraction logic doesn't handle Cell Ranger's actual data structure
  3. The path to the table rows uses incorrect naming (underscores vs camelCase)

Here are two reliable methods to get the data you want:

Method 1: Using BeautifulSoup (Refined)

Cell Ranger stores all summary data in a global window.__PRELOADED_STATE__ variable inside a script tag. We'll target that directly instead of relying on div attributes:

from bs4 import BeautifulSoup
import json
import pandas as pd

# Load the local HTML file
with open("web_summary.html", "r") as file:
    html_content = file.read()

soup = BeautifulSoup(html_content, "html.parser")

# Find the script tag containing the preloaded state data
script_tag = soup.find('script', text=lambda text: text and 'window.__PRELOADED_STATE__' in text)

if script_tag:
    # Extract the JSON string by stripping the variable declaration and trailing semicolon
    raw_json = script_tag.text.split('window.__PRELOADED_STATE__ = ')[1].rstrip(';')
    preloaded_data = json.loads(raw_json)
    
    # Navigate to the summary table rows (note camelCase instead of underscores)
    table_rows = preloaded_data['summary']['summaryTab']['table']['rows']
    
    # Convert to DataFrame and print
    summary_df = pd.DataFrame(table_rows, columns=['Metric', 'Value'])
    print(summary_df.to_string(index=False))
else:
    print("Could not locate the preloaded data script in the HTML file.")

Method 2: Using Regular Expressions (No BeautifulSoup Needed)

If you prefer a more lightweight approach, you can skip BeautifulSoup and use regex to pull the JSON directly from the file content:

import re
import json
import pandas as pd

with open("web_summary.html", "r") as file:
    html_content = file.read()

# Use regex to match the JSON object inside the preloaded state variable
match = re.search(r'window\.__PRELOADED_STATE__ = ({.*?});', html_content, re.DOTALL)

if match:
    json_data = json.loads(match.group(1))
    table_rows = json_data['summary']['summaryTab']['table']['rows']
    
    summary_df = pd.DataFrame(table_rows, columns=['Metric', 'Value'])
    print(summary_df.to_string(index=False))
else:
    print("Failed to extract the summary data from the HTML file.")

Troubleshooting Tip

If the path to the rows doesn't work for your specific Cell Ranger version, print the keys of the nested dictionaries to explore the structure:

print(preloaded_data['summary'].keys())  # Check what's under the summary key

Both methods will output a DataFrame matching the format you requested, with all the metrics and their corresponding values.

内容的提问来源于stack exchange,提问作者Rob John

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 17:06:51