Selenium点击分页后Beautiful Soup无法加载新页面问题求助
问题分析与解决方案
你的代码存在两个核心问题:一是count变量无法跨函数更新,导致每次解析都从1开始输出;二是点击下一页后未确保页面内容完成更新,BS4仍解析到旧页面的缓存内容。
具体修复点:
修复count变量传递
Python中整数是不可变类型,函数内部修改count不会影响外部变量。需要让parse_html函数返回更新后的count,在调用时重新赋值。等待页面内容加载完成
点击下一页后,仅靠time.sleep()不可靠,应该等待表格内容发生变化(比如等待当前页第一行的公司名称与上一页不同),确保动态加载的新内容已渲染完成。正确使用传入的pagesource参数
函数定义了pagesource参数但未使用,改为直接使用传入的参数,避免冗余。优化异常捕获
替换裸except为具体异常类型,避免隐藏代码中的其他错误。
修改后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from bs4 import BeautifulSoup as BeautifulSoup import time driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) def parse_html(pagesource, count): soup = BeautifulSoup(pagesource, 'html.parser') tables = soup.findChildren('table') my_table = tables[0] table_body = my_table.find('tbody') all_rows = table_body.find_all('tr') for row in all_rows: print(count) count += 1 try: path_body = row.find("td", class_="views-field-company-name") if not path_body: continue path = path_body.find("a")['href'] company_name = path_body.find("a").text.strip() print(company_name) issue_recepient_office = row.find("td", class_="views-field-field-building").string if issue_recepient_office: issue_recepient_office = issue_recepient_office.strip() detailed_description = row.find("td", class_="views-field-field-detailed-description-2").string detailed_description = detailed_description.strip() if detailed_description else "" except Exception as e: print(f"解析行时出错: {e}") continue return count # 返回更新后的count url = 'https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/compliance-actions-and-activities/warning-letters' driver.get(url) count = 1 # 记录第一页的第一个公司名称,用于后续判断页面是否更新 first_company = WebDriverWait(driver, 20).until( EC.presence_of_element_located((By.CSS_SELECTOR, ".views-field-company-name a")) ).text.strip() count = parse_html(driver.page_source, count) for i in range(0,3): # 点击下一页按钮 next_btn = WebDriverWait(driver, 30).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '#datatable_next a'))) next_btn.click() # 等待页面更新:直到第一个公司名称与上一页不同 WebDriverWait(driver, 30).until( lambda d: d.find_element(By.CSS_SELECTOR, ".views-field-company-name a").text.strip() != first_company ) # 更新当前页的第一个公司名称,用于下一次判断 first_company = driver.find_element(By.CSS_SELECTOR, ".views-field-company-name a").text.strip() # 解析新页面 count = parse_html(driver.page_source, count) driver.quit()
效果说明
修改后,每次点击下一页会等待页面内容完全更新,count变量也会持续递增,不会重复输出1-10,而是依次输出所有分页的序号和公司信息。
内容的提问来源于stack exchange,提问作者futurenext110
相关产品推荐
相关产品推荐

