You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 2.7+Selenium抓取JS生成表格遇阻:页面可见但page_source无内容

Troubleshooting JavaScript-Generated Table Scraping with Selenium 2.7

Hey there, let's break down why you're seeing that table in Chrome's DevTools but not in Selenium's page_source—it's a common gotcha with dynamic content, but we've got a few solid fixes to try.

First, a key clarification: Selenium's page_source doesn't always reflect the fully rendered DOM the way DevTools does. It’s often a snapshot of the initial page load plus some updates, but elements injected after complex JavaScript execution might not show up there. Instead of relying on the full page source, target the table directly.

Here are the steps to try, in order of likelihood:

1. Grab the table's HTML directly via the element itself

Skip the page_source entirely and fetch the table's rendered HTML straight from the DOM element:

# First, wait for the table to be present (you mentioned waiting isn't the issue, but just to be thorough)
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

wait = WebDriverWait(driver, 10)
table_element = wait.until(EC.presence_of_element_located((By.XPATH, "//table[contains(@id, 'your-table-id')]")))

# Now get the full HTML of the table
table_html = table_element.get_attribute('outerHTML')
print(table_html)

This should give you the exact HTML that's rendered in the browser, even if it doesn't show up in page_source.

2. Check for Shadow DOM encapsulation

If the table is wrapped in a Shadow DOM (a common pattern in modern web apps to isolate components), DevTools will show it but Selenium's standard element queries won't pick it up—and it won't appear in page_source. Use a JavaScript snippet to access the Shadow Root and fetch the table:

table_html = driver.execute_script("""
// Replace with the selector for the element that hosts the Shadow DOM
const shadowHost = document.querySelector('.shadow-host-class');
if (!shadowHost) return null;
const shadowRoot = shadowHost.shadowRoot;
const table = shadowRoot.querySelector('table');
return table ? table.outerHTML : 'Table not found in Shadow DOM';
""")
print(table_html)

3. Double-check for dynamically loaded iframes (again!)

You mentioned switching iframes didn't work, but sometimes iframes are injected after the initial page load—so you need to wait for them to exist before switching. Try this:

wait = WebDriverWait(driver, 15)
# Wait for the iframe to be present and switch to it
iframe = wait.until(EC.frame_to_be_available_and_switch_to_it((By.NAME, 'dynamic-iframe-name')))

# Now try locating the table inside the iframe
table_element = wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'table-class')))
print(table_element.get_attribute('outerHTML'))

# Don't forget to switch back to the main content when done
driver.switch_to.default_content()

4. Fetch the full rendered DOM via JavaScript

If all else fails, grab the entire rendered document's HTML using a JS command instead of page_source:

full_rendered_html = driver.execute_script("return document.documentElement.outerHTML")
# Now you can parse this HTML for your table, e.g., with BeautifulSoup
from bs4 import BeautifulSoup
soup = BeautifulSoup(full_rendered_html, 'html.parser')
table = soup.find('table', {'class': 'your-table-class'})
print(table)

One of these methods should get you that elusive table HTML. The most likely fix is directly targeting the element's outerHTML since page_source is often not up-to-date with late-stage dynamic content.

内容的提问来源于stack exchange,提问作者Guillaume

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:25:10