You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使pandas read_html等待页面表格加载完成后再读取?

Ah, I've run into this exact issue before! The problem here is that pandas.read_html() only parses the static HTML returned by the initial page request—it can't handle content that's loaded dynamically with JavaScript (like your delayed table). That's why you're seeing the "Still loading" placeholder instead of the actual data.

Luckily, there are two solid solutions to fix this:

1. Use Selenium to simulate a browser (wait for the table to load)

Selenium lets you spin up a real browser, wait for the JavaScript to render the table, then grab the fully loaded HTML to pass to read_html(). Here's how to do it:

First, make sure you have Selenium installed, plus the matching driver for your browser (e.g., ChromeDriver for Chrome):

pip install selenium

Then use this code:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import pandas as pd

# Initialize the browser driver (use Chrome here; adjust for Firefox/Edge if needed)
driver = webdriver.Chrome()
driver.get("YOUR_TARGET_URL")

# Wait for the table to finish loading (adjust the selector and timeout as needed)
# This example waits until the first data cell is no longer "Still loading"
wait = WebDriverWait(driver, 15)  # Wait up to 15 seconds
wait.until(
    EC.text_to_be_present_in_element(
        (By.CSS_SELECTOR, "table tr:nth-child(2) td:nth-child(1)"), 
        "Still loading"
    ),
    reverse=True
)

# Grab the fully rendered page source
page_html = driver.page_source

# Parse with pandas
dfs = pd.read_html(page_html)
target_table = dfs[0]  # Adjust index if there are multiple tables

# Clean up
driver.quit()

print(target_table)

2. Fetch the data directly from the backend API (faster alternative)

Often, delayed tables pull data from a hidden API endpoint via AJAX. You can skip the browser entirely by calling this API directly:

  1. Open your browser's DevTools (F12), go to the Network tab, and refresh the page.
  2. Look for XHR/Fetch requests that return JSON data matching your table content.
  3. Copy the request URL, headers, and any parameters.
  4. Use requests to pull the data and convert it to a DataFrame:
import requests
import pandas as pd

# Replace with your actual API endpoint and headers
api_url = "https://example.com/api/table-data"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

response = requests.get(api_url, headers=headers)
data = response.json()

# Convert the JSON data to a DataFrame (adjust based on the API's response structure)
df = pd.DataFrame(data["results"])
print(df)

Which method to choose?

  • Use Selenium if you can't find the API endpoint, or if the page has complex JS rendering logic.
  • Use the API method for better performance—it's faster and uses fewer resources than spinning up a browser.

内容的提问来源于stack exchange,提问作者iceagle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:33:40