使用Selenium+Pandas爬取网页表格遇TypeError问题求助
问题原因及解决方法
错误根源
你遇到的TypeError: 'WebElement' object is not iterable是因为误用了元素定位方法:
body.find_element(By.TAG_NAME,'tr')只会返回页面中第一个<tr>元素,是单个WebElement对象,无法被for循环遍历。- 同理,内层的
items.find_element(By.TAG_NAME,'td')也只会返回当前行的第一个<td>元素,无法遍历所有单元格。
必须使用**find_elements()(复数形式)方法**来获取所有匹配的元素列表,才能进行循环遍历。
修正后的可运行代码
考虑到目标网站存在动态加载,加入显式等待确保元素加载完成,避免空列表问题:
from selenium.webdriver.common.keys import Keys import pandas as pd from selenium.webdriver.common.by import By from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("https://www.investing.com/crypto/currencies") # 显式等待表格加载完成,最长等待10秒 table = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, 'table')) ) # 获取表头 thead = table.find_element(By.TAG_NAME, 'thead') header_cells = thead.find_elements(By.TAG_NAME, 'th') headers = [cell.text.strip() for cell in header_cells] # 获取表体所有行 tbody = table.find_element(By.TAG_NAME, 'tbody') rows = tbody.find_elements(By.TAG_NAME, 'tr') list_rows = [] for row in rows: # 获取当前行所有单元格 cells = row.find_elements(By.TAG_NAME, 'td') row_data = [cell.text.strip() for cell in cells] list_rows.append(row_data) # 转为DataFrame df = pd.DataFrame(list_rows, columns=headers) print(df) driver.close()
通用网页表格爬取模板
以下是适配绝大多数静态/半动态网页表格的通用代码,只需根据目标网页调整元素定位方式即可:
import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def crawl_table(url, table_locator, header_locator, row_locator, cell_locator): driver = webdriver.Chrome() driver.get(url) # 等待表格加载 table = WebDriverWait(driver, 15).until( EC.presence_of_element_located(table_locator) ) # 提取表头 headers = [h.text.strip() for h in table.find_elements(*header_locator)] # 提取所有行数据 rows = table.find_elements(*row_locator) data = [] for row in rows: cells = row.find_elements(*cell_locator) row_data = [c.text.strip() for c in cells] data.append(row_data) driver.close() return pd.DataFrame(data, columns=headers) # 示例调用(以你的目标网站为例) if __name__ == "__main__": target_url = "https://www.investing.com/crypto/currencies" # 定位参数:(定位方式, 定位值) table_loc = (By.TAG_NAME, 'table') header_loc = (By.TAG_NAME, 'th') row_loc = (By.TAG_NAME, 'tr') cell_loc = (By.TAG_NAME, 'td') result_df = crawl_table(target_url, table_loc, header_loc, row_loc, cell_loc) print(result_df)
使用说明
- 根据目标网页的HTML结构,修改
table_locator、header_locator等定位参数,比如可以用By.ID、By.CLASS_NAME等方式定位元素。 - 若表格是完全动态加载(如滚动加载更多行),需额外添加滚动加载的逻辑,比如循环执行
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")并等待新行加载。
内容的提问来源于stack exchange,提问作者Mathematics physics Biology
相关产品推荐
相关产品推荐

