如何爬取FDA网站带分页表格的前3页食品饮料类召回数据(下一页按钮无明确HTML标识)
如何爬取FDA网站带分页表格的前3页食品饮料类召回数据(下一页按钮无明确HTML标识)
我看了你的代码,发现两个核心问题导致你只能抓取第一页的数据:一是下一页按钮的定位XPath完全写错了,二是你把max_pages设成了1,当然只会循环一次啦!另外用固定的time.sleep等待页面加载也不太可靠,容易因为页面加载慢导致操作失败。
下面是修改后的完整代码,我标注了所有关键修改点:
from selenium import webdriver from selenium.webdriver.support.ui import Select from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from bs4 import BeautifulSoup import csv # Base and target URLs root = 'https://www.fda.gov' website = f'{root}/safety/recalls-market-withdrawals-safety-alerts' # Set up Selenium WebDriver driver = webdriver.Chrome() driver.get(website) # 显式等待下拉菜单加载完成,代替固定sleep,更稳定 wait = WebDriverWait(driver, 10) dropdown = Select(wait.until(EC.presence_of_element_located((By.ID, "edit-field-regulated-product-field")))) dropdown.select_by_value("2323") # 2323 corresponds to Food & Beverages # Initialize data storage recall_data = [] page_count = 0 max_pages = 3 # 修改:要爬3页就设为3 while page_count < max_pages: # Parse the page content soup = BeautifulSoup(driver.page_source, 'html.parser') # Locate the table table = soup.find('table', {'class': 'table'}) if not table: print("Table not found on current page, stopping pagination.") break # Extract data from the current page rows = table.find_all('tr')[1:] # Skip header row for row in rows: cols = row.find_all('td') if len(cols) > 1: recall_info = { 'Date': cols[0].text.strip(), 'Brand Names': cols[1].text.strip(), 'Product Description': cols[2].text.strip(), 'Product Type': cols[3].text.strip(), 'Recall Reason Description': cols[4].text.strip(), 'Company Name': cols[5].text.strip(), 'Terminated Recall': cols[6].text.strip(), } recall_data.append(recall_info) # 爬完第3页后不需要再点下一页,直接退出循环 if page_count == max_pages - 1: break # 修改:正确定位下一页按钮,FDA网站的下一页在分页组件的li.pager__item--next容器内 try: next_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//li[contains(@class, 'pager__item--next')]/a"))) next_button.click() page_count += 1 # 等待页面刷新,确保新页面的表格加载完成 wait.until(EC.staleness_of(table)) except Exception as e: print(f"Next button not found or click failed: {str(e)}, ending pagination.") break # Save data to CSV csv_filename = 'recalls.csv' csv_headers = [ 'Date', 'Brand Names', 'Product Description', 'Product Type', 'Recall Reason Description', 'Company Name', 'Terminated Recall' ] with open(csv_filename, 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=csv_headers) writer.writeheader() writer.writerows(recall_data) print(f"Data has been saved to {csv_filename}") # Close the driver driver.quit()
关键修改说明:
- 调整
max_pages值:从1改成3,确保循环会执行3次,覆盖前3页数据 - 修复下一页按钮定位:原代码的XPath完全不符合FDA网站的实际DOM结构,现在的定位是通过分页组件的专属类找到下一页按钮,更可靠
- 替换固定等待为显式等待:用
WebDriverWait等待元素加载完成/可点击,比time.sleep更灵活,能适应不同的页面加载速度 - 增加最后一页判断:爬完第3页后直接退出循环,避免无效的下一页点击尝试
- 清理冗余代码:删掉了你代码里重复的URL行,让代码更整洁
如果之后遇到分页组件的DOM结构变化,你可以用浏览器开发者工具(F12)查看下一页按钮的实际标签和类名,再调整XPath即可。
备注:内容来源于stack exchange,提问作者jw1622a
相关产品推荐
相关产品推荐

