使用Python Selenium获取网页表格数据报错:多表格下元素定位失败
Fixing Table Locator for BMF BOVESPA's "Ações em Circulação no Mercado"
Hey Ricardo, let's work through getting that target table loaded correctly. First, let's break down why your initial attempts didn't work:
Issues with Your Original Code
- You used
find_element_by_css_selector()but passed an XPath expression (//div[@id="div1"]) — CSS selectors use entirely different syntax (likediv#div1instead), so that was a syntax mismatch from the start. - Your second XPath had a typo:
div 1should bediv[1](ordiv:nth-child(1)), and even then, relying on fixed div indexes can be fragile if the page structure shifts unexpectedly.
Step-by-Step Fixes
1. First, Verify the Table's Actual Structure
Open the BMF BOVESPA page in your browser, hit F12 to launch DevTools, then use the element picker to locate the "Ações em Circulação no Mercado" table. Look for:
- A unique
idorclassattribute on the table itself - The heading element for the table (usually an
<h3>or<h4>) to use as a stable reference point
2. Use a More Reliable Locator
Here are two robust approaches:
Option A: Target via the Table Heading
This avoids fragile div indexes by using the table's visible title as a anchor:
from selenium.webdriver.common.by import By # First locate the heading text table_heading = browser.find_element(By.XPATH, '//h3[contains(text(), "Ações em Circulação no Mercado")]') # Then grab the table that follows this heading target_table = table_heading.find_element(By.XPATH, './following-sibling::div//table/tbody')
Option B: Wait for Dynamic Content (Critical for Loaded Pages)
If the table loads dynamically (after the initial page load), you need to wait for it to be present instead of trying to grab it immediately:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # Wait up to 10 seconds for the table to appear target_table = WebDriverWait(browser, 10).until( EC.presence_of_element_located( (By.XPATH, '//h3[contains(text(), "Ações em Circulação no Mercado")]/following-sibling::div//table/tbody') ) ) # Extract the table content content = target_table.get_attribute('innerHTML')
3. Test Your Locator First
Before putting it in code, test your XPath in DevTools' Console:
- Type
$x('//h3[contains(text(), "Ações em Circulação no Mercado")]')and press Enter — if it returns the heading element, your path is correct. - Adjust the XPath until it reliably picks out the table body.
内容的提问来源于stack exchange,提问作者Ricardo
相关产品推荐
相关产品推荐

