Python+Selenium无法定位维基百科页面全部表格的问题求助
问题排查与解决方案
核心问题分析
- XPath匹配错误:你编写的XPath中使用的
"Población histórica de México"与页面实际表格的标题不符。第二个表格的标题是"Población histórica por entidad federativa",第三个表格的标题是"Proyecciones de población por entidad federativa",这直接导致Selenium无法定位到目标元素。 - 等待逻辑不足:仅等待页面标题加载完成,无法保证页面下方的表格完全渲染,部分表格可能存在加载延迟。
修正后的代码
# Libraries from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager import pandas as pd import time print("Starting script...") # Initialize Chrome options options = webdriver.ChromeOptions() options.add_argument('--start-maximized') options.add_argument('--disable-extensions') service = Service(ChromeDriverManager().install()) driver = webdriver.Chrome(service=service, options=options) try: driver.set_window_position(2000, 0) driver.maximize_window() time.sleep(5) driver.get('https://es.wikipedia.org/wiki/Anexo:Entidades_federativas_de_M%C3%A9xico_por_superficie,_poblaci%C3%B3n_y_densidad') # 等待所有表格加载完成,确保页面完全渲染 WebDriverWait(driver, 30).until( EC.presence_of_all_elements_located((By.XPATH, '//table[contains(@class, "wikitable")]')) ) print("Page loaded successfully") # 修正XPath,匹配实际表格标题 first_table = driver.find_element(By.XPATH, '//table[contains(.//caption, "Entidades federativas de México por superficie, población y densidad")]') # 第二个表格:历史人口数据 second_table = driver.find_element(By.XPATH, '//table[contains(.//caption, "Población histórica por entidad federativa")]') # 第三个表格:人口预测数据 third_table = driver.find_element(By.XPATH, '//table[contains(.//caption, "Proyecciones de población por entidad federativa")]') print("All tables found") # First table extraction first_table_html = first_table.get_attribute('outerHTML') first_table_df = pd.read_html(first_table_html)[0] first_table_df = first_table_df.iloc[2:34, :] first_table_df.columns = first_table_df.iloc[0] first_table_df = first_table_df[1:] print("First table extracted successfully") # Second table extraction second_table_html = second_table.get_attribute('outerHTML') second_table_df = pd.read_html(second_table_html)[0] second_table_df.columns = ['Pos', 'Entidad', '2020', '2010', '2000', '1990', '1980', '1970', '1960', '1950', '1940', '1930', '1921', '1910'] print("Second table extracted successfully") # Third table extraction third_table_html = third_table.get_attribute('outerHTML') third_table_df = pd.read_html(third_table_html)[0] third_table_df.columns = ['Pos', 'Entidad', '2010', '2015', '2020', '2025', '2030'] print("Third table extracted successfully") # Save to Excel with each table on a different sheet with pd.ExcelWriter('mexico_population_data.xlsx') as writer: first_table_df.to_excel(writer, sheet_name='Superficie_Poblacion_Densidad', index=False) second_table_df.to_excel(writer, sheet_name='Poblacion_Historica', index=False) third_table_df.to_excel(writer, sheet_name='Poblacion_Futura', index=False) print("Data extraction and Excel file creation successful") except Exception as e: print(f"An error occurred: {e}") finally: # Close the browser after a delay to see the loaded page time.sleep(10) driver.quit()
额外优化说明
- 移除了未使用的
StringIO导入,精简代码结构。 - 改用
presence_of_all_elements_located等待所有表格加载,比仅等待页面标题更可靠。 - 直接通过唯一的表格标题定位元素,避免了索引定位可能带来的错误。
内容的提问来源于stack exchange,提问作者Erik James Robles
相关产品推荐
相关产品推荐

