使用Python Selenium获取网页多页指定表格全部行数据的问题咨询
解决方案
核心问题排查
- 获取表格行使用单数方法
find_element_by_xpath,仅返回第一个匹配的<tr>,需改用复数方法find_elements_by_xpath获取当前页所有行 - 遍历行时使用绝对路径错误定位到表头
<th>,导致所有取值都是表头内容,需基于当前行节点使用相对路径定位对应<td>单元格 - 存储数据的列表放在分页循环内部,每页加载后会清空历史数据,需移到分页循环外初始化
修正代码
from selenium import webdriver from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import pandas as pd import time webside = 'https://www.xxx.dk/find-arkitekt?display_view=block_3&field_company_region=All' driver = webdriver.Chrome(executable_path=r'C:\Users\KristerJens\Downloads\chromedriver_win32\chromedriver') driver.implicitly_wait(10) driver.get(webside) wait = WebDriverWait(driver,10) # 处理cookie driver.execute_script("arguments[0].click();", wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@class='agree-button eu-cookie-compliance-default-button']")))) # 收集所有分页链接 pages = [] for page in driver.find_elements_by_xpath("//div[@class='pagination']//ul[@class='pager__items js-pager__items list-pages']//li"): pages.append(page.find_element_by_tag_name('a').get_attribute("href")) # 初始化所有存储列表,放在分页循环外 virk=[] post=[] by=[] web=[] mail=[] # 遍历所有分页 for i in pages: driver.get(i) # 等待表格加载完成 wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='architect-view-table']//table[@class='cols-5 responsive-enabled']//tbody//tr"))) # 改用复数方法获取当前页所有tr options = driver.find_elements_by_xpath("//div[@class='architect-view-table']//table[@class='cols-5 responsive-enabled']//tbody//tr") # 遍历当前页每一行 for opt in options: # 基于当前行opt使用相对路径定位对应单元格,td索引根据实际列顺序调整,从1开始计数 virk.append(opt.find_element_by_xpath("./td[1]").text.strip()) post.append(opt.find_element_by_xpath("./td[2]").text.strip()) by.append(opt.find_element_by_xpath("./td[3]").text.strip()) web.append(opt.find_element_by_xpath("./td[4]").text.strip()) mail.append(opt.find_element_by_xpath("./td[5]").text.strip()) # 可选:导出为csv文件 df = pd.DataFrame({ '公司名称': virk, '邮编': post, '城市': by, '官网': web, '邮箱': mail }) df.to_csv('arkitekt_data.csv', index=False, encoding='utf-8-sig') # 关闭浏览器 driver.quit()
注意事项
- 代码中
./td[N]的索引需根据你实际表格的字段顺序调整,从1开始计数 - 如果你使用的是Selenium4版本,所有
find_element_by_xxx方法需要替换为find_element(By.XXX, 路径)的写法,比如find_element_by_xpath("./td[1]")改为find_element(By.XPATH, "./td[1]")
内容的提问来源于stack exchange,提问作者awi1100
相关产品推荐
相关产品推荐

