You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium+Pandas爬取网页表格遇TypeError问题求助

问题原因及解决方法

错误根源

你遇到的TypeError: 'WebElement' object is not iterable是因为误用了元素定位方法:

  • body.find_element(By.TAG_NAME,'tr')只会返回页面中第一个<tr>元素,是单个WebElement对象,无法被for循环遍历。
  • 同理,内层的items.find_element(By.TAG_NAME,'td')也只会返回当前行的第一个<td>元素,无法遍历所有单元格。

必须使用**find_elements()(复数形式)方法**来获取所有匹配的元素列表,才能进行循环遍历。

修正后的可运行代码

考虑到目标网站存在动态加载,加入显式等待确保元素加载完成,避免空列表问题:

from selenium.webdriver.common.keys import Keys
import pandas as pd
from selenium.webdriver.common.by import By
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
driver.get("https://www.investing.com/crypto/currencies")

# 显式等待表格加载完成,最长等待10秒
table = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.TAG_NAME, 'table'))
)

# 获取表头
thead = table.find_element(By.TAG_NAME, 'thead')
header_cells = thead.find_elements(By.TAG_NAME, 'th')
headers = [cell.text.strip() for cell in header_cells]

# 获取表体所有行
tbody = table.find_element(By.TAG_NAME, 'tbody')
rows = tbody.find_elements(By.TAG_NAME, 'tr')

list_rows = []
for row in rows:
    # 获取当前行所有单元格
    cells = row.find_elements(By.TAG_NAME, 'td')
    row_data = [cell.text.strip() for cell in cells]
    list_rows.append(row_data)

# 转为DataFrame
df = pd.DataFrame(list_rows, columns=headers)
print(df)

driver.close()

通用网页表格爬取模板

以下是适配绝大多数静态/半动态网页表格的通用代码,只需根据目标网页调整元素定位方式即可:

import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def crawl_table(url, table_locator, header_locator, row_locator, cell_locator):
    driver = webdriver.Chrome()
    driver.get(url)
    
    # 等待表格加载
    table = WebDriverWait(driver, 15).until(
        EC.presence_of_element_located(table_locator)
    )
    
    # 提取表头
    headers = [h.text.strip() for h in table.find_elements(*header_locator)]
    
    # 提取所有行数据
    rows = table.find_elements(*row_locator)
    data = []
    for row in rows:
        cells = row.find_elements(*cell_locator)
        row_data = [c.text.strip() for c in cells]
        data.append(row_data)
    
    driver.close()
    return pd.DataFrame(data, columns=headers)

# 示例调用(以你的目标网站为例)
if __name__ == "__main__":
    target_url = "https://www.investing.com/crypto/currencies"
    # 定位参数:(定位方式, 定位值)
    table_loc = (By.TAG_NAME, 'table')
    header_loc = (By.TAG_NAME, 'th')
    row_loc = (By.TAG_NAME, 'tr')
    cell_loc = (By.TAG_NAME, 'td')
    
    result_df = crawl_table(target_url, table_loc, header_loc, row_loc, cell_loc)
    print(result_df)

使用说明

  • 根据目标网页的HTML结构,修改table_locator、header_locator等定位参数,比如可以用By.ID、By.CLASS_NAME等方式定位元素。
  • 若表格是完全动态加载(如滚动加载更多行),需额外添加滚动加载的逻辑,比如循环执行driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")并等待新行加载。

内容的提问来源于stack exchange,提问作者Mathematics physics Biology

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 15:45:39