如何用Python爬取同一URL内的多页表格?Selenium爬取失效求助
问题解决思路
你的代码存在几个关键问题导致无法获取后续页面数据:
- 点击分页后没有重新获取页面源码并重新解析,仍然用的是第一页的soup对象
- XPath中的引号转义错误(
"是HTML转义,Selenium里直接用双引号或单引号即可) - 没有等待页面加载完成就解析,可能元素还没渲染
修正后的完整代码
import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup as soup # 初始化浏览器 driver = webdriver.Chrome() url = 'http://www.witcorp.co.th/th/product_chem.php?txtType=1#nogo' driver.get(url) # 存储所有页面的数据 all_products = [] # 循环处理5页 for page_num in range(1, 6): # 等待表格区域加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, 'showChemicalAll')) ) # 获取当前页面源码并解析 page_html = driver.page_source data = soup(page_html, 'html.parser') # 提取当前页的产品数据(根据实际表格结构调整选择器) current_page_products = data.find_all('div', {'class': 'table'}) all_products.extend(current_page_products) # 如果不是最后一页,点击下一页 if page_num < 5: # 定位下一页按钮,用页码文本定位更稳定 next_page_btn = driver.find_element(By.XPATH, f'//*[@id="showChemicalPage"]/ul/li[contains(@class, "page-item")][text()="{page_num + 1}"]') next_page_btn.click() # 关闭浏览器 driver.quit() # 处理数据(示例:转换为DataFrame,需根据实际页面结构调整字段提取逻辑) product_list = [] for product in all_products: # 示例:提取产品文本内容,实际按需拆分字段 items = product.get_text(strip=True, separator='|').split('|') product_list.append(items) df = pd.DataFrame(product_list) print(df)
关键修正点说明
- 重新获取页面源码:每次点击分页后必须重新调用
driver.page_source并生成新的soup对象,否则一直操作的是第一页的DOM - 修复XPath引号:把
"换成正常的双引号",Selenium能直接识别 - 添加显式等待:用
WebDriverWait等待表格区域加载完成,避免页面未渲染就解析 - 通用分页定位:通过页码文本定位下一页按钮,比硬编码li索引更稳定,适配页面结构变化
- 循环遍历所有页面:用for循环自动完成1-5页的切换和数据收集
内容的提问来源于stack exchange,提问作者Hju Gtt Sde
相关产品推荐
相关产品推荐

