You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+Selenium无法定位维基百科页面全部表格的问题求助

问题排查与解决方案

核心问题分析

  1. XPath匹配错误:你编写的XPath中使用的"Población histórica de México"与页面实际表格的标题不符。第二个表格的标题是"Población histórica por entidad federativa",第三个表格的标题是"Proyecciones de población por entidad federativa",这直接导致Selenium无法定位到目标元素。
  2. 等待逻辑不足:仅等待页面标题加载完成,无法保证页面下方的表格完全渲染,部分表格可能存在加载延迟。

修正后的代码

# Libraries
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from webdriver_manager.chrome import ChromeDriverManager
import pandas as pd
import time

print("Starting script...")

# Initialize Chrome options
options = webdriver.ChromeOptions()
options.add_argument('--start-maximized')
options.add_argument('--disable-extensions')

service = Service(ChromeDriverManager().install())
driver = webdriver.Chrome(service=service, options=options)

try:
    driver.set_window_position(2000, 0)
    driver.maximize_window()
    time.sleep(5)

    driver.get('https://es.wikipedia.org/wiki/Anexo:Entidades_federativas_de_M%C3%A9xico_por_superficie,_poblaci%C3%B3n_y_densidad')

    # 等待所有表格加载完成,确保页面完全渲染
    WebDriverWait(driver, 30).until(
        EC.presence_of_all_elements_located((By.XPATH, '//table[contains(@class, "wikitable")]'))
    )

    print("Page loaded successfully")

    # 修正XPath,匹配实际表格标题
    first_table = driver.find_element(By.XPATH, '//table[contains(.//caption, "Entidades federativas de México por superficie, población y densidad")]')
    # 第二个表格:历史人口数据
    second_table = driver.find_element(By.XPATH, '//table[contains(.//caption, "Población histórica por entidad federativa")]')
    # 第三个表格:人口预测数据
    third_table = driver.find_element(By.XPATH, '//table[contains(.//caption, "Proyecciones de población por entidad federativa")]')

    print("All tables found")

    # First table extraction
    first_table_html = first_table.get_attribute('outerHTML')
    first_table_df = pd.read_html(first_table_html)[0]
    first_table_df = first_table_df.iloc[2:34, :]
    first_table_df.columns = first_table_df.iloc[0]
    first_table_df = first_table_df[1:]
    print("First table extracted successfully")

    # Second table extraction
    second_table_html = second_table.get_attribute('outerHTML')
    second_table_df = pd.read_html(second_table_html)[0]
    second_table_df.columns = ['Pos', 'Entidad', '2020', '2010', '2000', '1990', '1980', '1970', '1960', '1950', '1940', '1930', '1921', '1910']
    print("Second table extracted successfully")

    # Third table extraction
    third_table_html = third_table.get_attribute('outerHTML')
    third_table_df = pd.read_html(third_table_html)[0]
    third_table_df.columns = ['Pos', 'Entidad', '2010', '2015', '2020', '2025', '2030']
    print("Third table extracted successfully")

    # Save to Excel with each table on a different sheet
    with pd.ExcelWriter('mexico_population_data.xlsx') as writer:
        first_table_df.to_excel(writer, sheet_name='Superficie_Poblacion_Densidad', index=False)
        second_table_df.to_excel(writer, sheet_name='Poblacion_Historica', index=False)
        third_table_df.to_excel(writer, sheet_name='Poblacion_Futura', index=False)

    print("Data extraction and Excel file creation successful")

except Exception as e:
    print(f"An error occurred: {e}")

finally:
    # Close the browser after a delay to see the loaded page
    time.sleep(10)
    driver.quit()

额外优化说明

  • 移除了未使用的StringIO导入,精简代码结构。
  • 改用presence_of_all_elements_located等待所有表格加载,比仅等待页面标题更可靠。
  • 直接通过唯一的表格标题定位元素,避免了索引定位可能带来的错误。

内容的提问来源于stack exchange,提问作者Erik James Robles

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 14:37:02