Selenium网页表格爬取疑难问题求助
Let’s break down your issues step by step—this is a common pain point with dynamic table scraping, so let’s unpack what’s going on.
Why You’re Only Getting Headers + NaN Rows
There are a few key reasons your current setup isn’t pulling full table data:
- Insufficient wait conditions: You’re waiting for the outer
divcontainer to be visible, but that doesn’t guarantee the table rows inside have finished loading. The container might render empty first, then populate data via JavaScript later. Yourpage_sourcecaptures that empty state, so pandas only sees the header. - Misaligned element targeting: The
divyou’re selecting is a scroll wrapper, not the actual table content. Inside that div, there’s likely a nested<table>tag with the rows. When you pass the div’s HTML topd.read_html, pandas might struggle to parse the nested structure correctly. - Dynamic data loading: The table rows are probably loaded asynchronously (via AJAX/fetch calls) after the initial page load. Selenium’s
page_sourcecaptures the DOM at the moment you call it—if the data hasn’t finished fetching/rendering yet, it won’t be included.
Why This Scraping Task Is Tricky
Dynamic tables like this are tough for a few reasons:
- Framework-generated classes: The long, hyphenated class names (e.g.,
QuoteHistoryTable__table__scroller) are almost certainly auto-generated by a frontend framework like React or Vue. These can change with site updates, making your selectors fragile. - Client-side rendering: The table data isn’t present in the initial HTML response—it’s built by JavaScript after the page loads. You can’t just scrape the raw HTML; you need to wait for the JS to finish its work.
- Scroll-dependent loading: Some tables load rows incrementally as you scroll. If the table is long, only the first few rows might be in the DOM when you capture
page_source.
Why You Can’t Reproduce DebanjanB’s Solution
Here are the most likely culprits for the discrepancy:
- Page structure changes: Websites often tweak their frontend code—since DebanjanB provided the solution, the site might have updated class names, nesting, or how data is loaded. Your original selectors (or the ones in the solution) might no longer point to the right elements.
- Environment mismatches: Selenium behavior can vary based on browser version, WebDriver version, and Selenium library version. If your setup (e.g., Chrome 120 vs. their Chrome 118) doesn’t match, certain wait conditions or locators might fail.
- Session state differences: The site might require a logged-in session, specific cookies, or geographic restrictions to load the full table. If DebanjanB was logged in or had different session context, you’ll miss that data without replicating those conditions.
- Wait condition gaps: The solution might have used more precise waits (e.g., waiting for a minimum number of table rows to be present) instead of just waiting for the container to be visible. If you’re using your original wait logic instead of their exact conditions, you’ll still capture the empty state.
Fixes to Try
Let’s adjust your code to address these issues:
Wait for table rows instead of the container:
# Wait for at least one data row to be present (adjust the XPath to match your rows) WebDriverWait(driver, 15).until( expected_conditions.presence_of_all_elements_located( (By.XPATH, "//div[@class='table-scroller ScrollableTable__table-scroller QuoteHistoryTable__table__scroller QuoteHistoryTable__QuoteHistoryTable__table__scroller']//tr[contains(@class, 'data-row')]") ) )Replace
'data-row'with a class or attribute that identifies non-header rows in the table.Target the actual table tag:
After capturingpage_source, use BeautifulSoup to find the nested<table>instead of the wrapper div:soup = BeautifulSoup(source, "html5lib") # Find the table inside the scroll wrapper table = soup.find('div', {'class': 'table-scroller ScrollableTable__table-scroller QuoteHistoryTable__table__scroller QuoteHistoryTable__QuoteHistoryTable__table__scroller'}).find('table') df = pd.read_html(str(table), flavor='html5lib', header=0, thousands='.', decimal=',')Check for AJAX requests:
Use your browser’s DevTools (Network tab) to look for API calls that return the table data. If you can find the request URL, you can skip Selenium entirely and fetch the data directly withrequests—this is often more reliable.Verify your environment:
Make sure your browser, WebDriver (e.g., chromedriver), and Selenium library versions are compatible. Check the Selenium documentation for version matching guidelines.
内容的提问来源于stack exchange,提问作者JamesHudson81

