使用BeautifulSoup无法抓取本地HTML文件表格数据求解决方案
Let's fix your script to pull that summary table data correctly. The main issues with your original code are:
- You're targeting the wrong element to find the data script
- The JSON extraction logic doesn't handle Cell Ranger's actual data structure
- The path to the table rows uses incorrect naming (underscores vs camelCase)
Here are two reliable methods to get the data you want:
Method 1: Using BeautifulSoup (Refined)
Cell Ranger stores all summary data in a global window.__PRELOADED_STATE__ variable inside a script tag. We'll target that directly instead of relying on div attributes:
from bs4 import BeautifulSoup import json import pandas as pd # Load the local HTML file with open("web_summary.html", "r") as file: html_content = file.read() soup = BeautifulSoup(html_content, "html.parser") # Find the script tag containing the preloaded state data script_tag = soup.find('script', text=lambda text: text and 'window.__PRELOADED_STATE__' in text) if script_tag: # Extract the JSON string by stripping the variable declaration and trailing semicolon raw_json = script_tag.text.split('window.__PRELOADED_STATE__ = ')[1].rstrip(';') preloaded_data = json.loads(raw_json) # Navigate to the summary table rows (note camelCase instead of underscores) table_rows = preloaded_data['summary']['summaryTab']['table']['rows'] # Convert to DataFrame and print summary_df = pd.DataFrame(table_rows, columns=['Metric', 'Value']) print(summary_df.to_string(index=False)) else: print("Could not locate the preloaded data script in the HTML file.")
Method 2: Using Regular Expressions (No BeautifulSoup Needed)
If you prefer a more lightweight approach, you can skip BeautifulSoup and use regex to pull the JSON directly from the file content:
import re import json import pandas as pd with open("web_summary.html", "r") as file: html_content = file.read() # Use regex to match the JSON object inside the preloaded state variable match = re.search(r'window\.__PRELOADED_STATE__ = ({.*?});', html_content, re.DOTALL) if match: json_data = json.loads(match.group(1)) table_rows = json_data['summary']['summaryTab']['table']['rows'] summary_df = pd.DataFrame(table_rows, columns=['Metric', 'Value']) print(summary_df.to_string(index=False)) else: print("Failed to extract the summary data from the HTML file.")
Troubleshooting Tip
If the path to the rows doesn't work for your specific Cell Ranger version, print the keys of the nested dictionaries to explore the structure:
print(preloaded_data['summary'].keys()) # Check what's under the summary key
Both methods will output a DataFrame matching the format you requested, with all the metrics and their corresponding values.
内容的提问来源于stack exchange,提问作者Rob John

