You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium网页表格爬取疑难问题求助

Troubleshooting Incomplete Table Scraping with Selenium & Pandas

Let’s break down your issues step by step—this is a common pain point with dynamic table scraping, so let’s unpack what’s going on.

Why You’re Only Getting Headers + NaN Rows

There are a few key reasons your current setup isn’t pulling full table data:

  • Insufficient wait conditions: You’re waiting for the outer div container to be visible, but that doesn’t guarantee the table rows inside have finished loading. The container might render empty first, then populate data via JavaScript later. Your page_source captures that empty state, so pandas only sees the header.
  • Misaligned element targeting: The div you’re selecting is a scroll wrapper, not the actual table content. Inside that div, there’s likely a nested <table> tag with the rows. When you pass the div’s HTML to pd.read_html, pandas might struggle to parse the nested structure correctly.
  • Dynamic data loading: The table rows are probably loaded asynchronously (via AJAX/fetch calls) after the initial page load. Selenium’s page_source captures the DOM at the moment you call it—if the data hasn’t finished fetching/rendering yet, it won’t be included.

Why This Scraping Task Is Tricky

Dynamic tables like this are tough for a few reasons:

  • Framework-generated classes: The long, hyphenated class names (e.g., QuoteHistoryTable__table__scroller) are almost certainly auto-generated by a frontend framework like React or Vue. These can change with site updates, making your selectors fragile.
  • Client-side rendering: The table data isn’t present in the initial HTML response—it’s built by JavaScript after the page loads. You can’t just scrape the raw HTML; you need to wait for the JS to finish its work.
  • Scroll-dependent loading: Some tables load rows incrementally as you scroll. If the table is long, only the first few rows might be in the DOM when you capture page_source.

Why You Can’t Reproduce DebanjanB’s Solution

Here are the most likely culprits for the discrepancy:

  • Page structure changes: Websites often tweak their frontend code—since DebanjanB provided the solution, the site might have updated class names, nesting, or how data is loaded. Your original selectors (or the ones in the solution) might no longer point to the right elements.
  • Environment mismatches: Selenium behavior can vary based on browser version, WebDriver version, and Selenium library version. If your setup (e.g., Chrome 120 vs. their Chrome 118) doesn’t match, certain wait conditions or locators might fail.
  • Session state differences: The site might require a logged-in session, specific cookies, or geographic restrictions to load the full table. If DebanjanB was logged in or had different session context, you’ll miss that data without replicating those conditions.
  • Wait condition gaps: The solution might have used more precise waits (e.g., waiting for a minimum number of table rows to be present) instead of just waiting for the container to be visible. If you’re using your original wait logic instead of their exact conditions, you’ll still capture the empty state.

Fixes to Try

Let’s adjust your code to address these issues:

  1. Wait for table rows instead of the container:

    # Wait for at least one data row to be present (adjust the XPath to match your rows)
    WebDriverWait(driver, 15).until(
        expected_conditions.presence_of_all_elements_located(
            (By.XPATH, "//div[@class='table-scroller ScrollableTable__table-scroller QuoteHistoryTable__table__scroller QuoteHistoryTable__QuoteHistoryTable__table__scroller']//tr[contains(@class, 'data-row')]")
        )
    )
    

    Replace 'data-row' with a class or attribute that identifies non-header rows in the table.

  2. Target the actual table tag:
    After capturing page_source, use BeautifulSoup to find the nested <table> instead of the wrapper div:

    soup = BeautifulSoup(source, "html5lib")
    # Find the table inside the scroll wrapper
    table = soup.find('div', {'class': 'table-scroller ScrollableTable__table-scroller QuoteHistoryTable__table__scroller QuoteHistoryTable__QuoteHistoryTable__table__scroller'}).find('table')
    df = pd.read_html(str(table), flavor='html5lib', header=0, thousands='.', decimal=',')
    
  3. Check for AJAX requests:
    Use your browser’s DevTools (Network tab) to look for API calls that return the table data. If you can find the request URL, you can skip Selenium entirely and fetch the data directly with requests—this is often more reliable.

  4. Verify your environment:
    Make sure your browser, WebDriver (e.g., chromedriver), and Selenium library versions are compatible. Check the Selenium documentation for version matching guidelines.

内容的提问来源于stack exchange,提问作者JamesHudson81

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:43:12