You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何爬取FDA网站带分页表格的前3页食品饮料类召回数据(下一页按钮无明确HTML标识)

如何爬取FDA网站带分页表格的前3页食品饮料类召回数据(下一页按钮无明确HTML标识)

我看了你的代码,发现两个核心问题导致你只能抓取第一页的数据:一是下一页按钮的定位XPath完全写错了,二是你把max_pages设成了1,当然只会循环一次啦!另外用固定的time.sleep等待页面加载也不太可靠,容易因为页面加载慢导致操作失败。

下面是修改后的完整代码,我标注了所有关键修改点:

from selenium import webdriver
from selenium.webdriver.support.ui import Select
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup
import csv

# Base and target URLs
root = 'https://www.fda.gov'
website = f'{root}/safety/recalls-market-withdrawals-safety-alerts'

# Set up Selenium WebDriver
driver = webdriver.Chrome()
driver.get(website)

# 显式等待下拉菜单加载完成,代替固定sleep,更稳定
wait = WebDriverWait(driver, 10)
dropdown = Select(wait.until(EC.presence_of_element_located((By.ID, "edit-field-regulated-product-field"))))
dropdown.select_by_value("2323")  # 2323 corresponds to Food & Beverages

# Initialize data storage
recall_data = []
page_count = 0
max_pages = 3  # 修改:要爬3页就设为3

while page_count < max_pages:
    # Parse the page content
    soup = BeautifulSoup(driver.page_source, 'html.parser')

    # Locate the table
    table = soup.find('table', {'class': 'table'})
    if not table:
        print("Table not found on current page, stopping pagination.")
        break

    # Extract data from the current page
    rows = table.find_all('tr')[1:]  # Skip header row
    for row in rows:
        cols = row.find_all('td')
        if len(cols) > 1:
            recall_info = {
                'Date': cols[0].text.strip(),
                'Brand Names': cols[1].text.strip(),
                'Product Description': cols[2].text.strip(),
                'Product Type': cols[3].text.strip(),
                'Recall Reason Description': cols[4].text.strip(),
                'Company Name': cols[5].text.strip(),
                'Terminated Recall': cols[6].text.strip(),
            }
            recall_data.append(recall_info)

    # 爬完第3页后不需要再点下一页,直接退出循环
    if page_count == max_pages - 1:
        break

    # 修改:正确定位下一页按钮,FDA网站的下一页在分页组件的li.pager__item--next容器内
    try:
        next_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//li[contains(@class, 'pager__item--next')]/a")))
        next_button.click()
        page_count += 1
        # 等待页面刷新,确保新页面的表格加载完成
        wait.until(EC.staleness_of(table))
    except Exception as e:
        print(f"Next button not found or click failed: {str(e)}, ending pagination.")
        break

# Save data to CSV
csv_filename = 'recalls.csv'
csv_headers = [
    'Date', 
    'Brand Names', 
    'Product Description', 
    'Product Type', 
    'Recall Reason Description', 
    'Company Name', 
    'Terminated Recall'
]

with open(csv_filename, 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=csv_headers)
    writer.writeheader()
    writer.writerows(recall_data)

print(f"Data has been saved to {csv_filename}")

# Close the driver
driver.quit()

关键修改说明:

  1. 调整max_pages值:从1改成3,确保循环会执行3次,覆盖前3页数据
  2. 修复下一页按钮定位:原代码的XPath完全不符合FDA网站的实际DOM结构,现在的定位是通过分页组件的专属类找到下一页按钮,更可靠
  3. 替换固定等待为显式等待:用WebDriverWait等待元素加载完成/可点击,比time.sleep更灵活,能适应不同的页面加载速度
  4. 增加最后一页判断:爬完第3页后直接退出循环,避免无效的下一页点击尝试
  5. 清理冗余代码:删掉了你代码里重复的URL行,让代码更整洁

如果之后遇到分页组件的DOM结构变化,你可以用浏览器开发者工具(F12)查看下一页按钮的实际标签和类名,再调整XPath即可。

备注:内容来源于stack exchange,提问作者jw1622a

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 16:10:28