You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium获取网页多页指定表格全部行数据的问题咨询

解决方案

核心问题排查

  • 获取表格行使用单数方法find_element_by_xpath,仅返回第一个匹配的<tr>,需改用复数方法find_elements_by_xpath获取当前页所有行
  • 遍历行时使用绝对路径错误定位到表头<th>,导致所有取值都是表头内容,需基于当前行节点使用相对路径定位对应<td>单元格
  • 存储数据的列表放在分页循环内部,每页加载后会清空历史数据,需移到分页循环外初始化

修正代码

from selenium import webdriver
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import pandas as pd
import time

webside = 'https://www.xxx.dk/find-arkitekt?display_view=block_3&field_company_region=All'

driver = webdriver.Chrome(executable_path=r'C:\Users\KristerJens\Downloads\chromedriver_win32\chromedriver')
driver.implicitly_wait(10)    
driver.get(webside)    
wait = WebDriverWait(driver,10)

# 处理cookie
driver.execute_script("arguments[0].click();", wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@class='agree-button eu-cookie-compliance-default-button']"))))

# 收集所有分页链接
pages = []
for page in driver.find_elements_by_xpath("//div[@class='pagination']//ul[@class='pager__items js-pager__items list-pages']//li"):
    pages.append(page.find_element_by_tag_name('a').get_attribute("href"))

# 初始化所有存储列表,放在分页循环外
virk=[]
post=[]
by=[]
web=[]
mail=[]    

# 遍历所有分页
for i in pages:
    driver.get(i)
    # 等待表格加载完成
    wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='architect-view-table']//table[@class='cols-5 responsive-enabled']//tbody//tr")))
    # 改用复数方法获取当前页所有tr
    options = driver.find_elements_by_xpath("//div[@class='architect-view-table']//table[@class='cols-5 responsive-enabled']//tbody//tr")
    
    # 遍历当前页每一行
    for opt in options:
        # 基于当前行opt使用相对路径定位对应单元格,td索引根据实际列顺序调整,从1开始计数
        virk.append(opt.find_element_by_xpath("./td[1]").text.strip())
        post.append(opt.find_element_by_xpath("./td[2]").text.strip())
        by.append(opt.find_element_by_xpath("./td[3]").text.strip())
        web.append(opt.find_element_by_xpath("./td[4]").text.strip())
        mail.append(opt.find_element_by_xpath("./td[5]").text.strip())

# 可选:导出为csv文件
df = pd.DataFrame({
    '公司名称': virk,
    '邮编': post,
    '城市': by,
    '官网': web,
    '邮箱': mail
})
df.to_csv('arkitekt_data.csv', index=False, encoding='utf-8-sig')

# 关闭浏览器
driver.quit()

注意事项

  • 代码中./td[N]的索引需根据你实际表格的字段顺序调整,从1开始计数
  • 如果你使用的是Selenium4版本,所有find_element_by_xxx方法需要替换为find_element(By.XXX, 路径)的写法,比如find_element_by_xpath("./td[1]")改为find_element(By.XPATH, "./td[1]")

内容的提问来源于stack exchange,提问作者awi1100

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 00:15:04