You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium点击分页后Beautiful Soup无法加载新页面问题求助

问题分析与解决方案

你的代码存在两个核心问题:一是count变量无法跨函数更新,导致每次解析都从1开始输出;二是点击下一页后未确保页面内容完成更新,BS4仍解析到旧页面的缓存内容。

具体修复点:

  1. 修复count变量传递
    Python中整数是不可变类型,函数内部修改count不会影响外部变量。需要让parse_html函数返回更新后的count,在调用时重新赋值。

  2. 等待页面内容加载完成
    点击下一页后,仅靠time.sleep()不可靠,应该等待表格内容发生变化(比如等待当前页第一行的公司名称与上一页不同),确保动态加载的新内容已渲染完成。

  3. 正确使用传入的pagesource参数
    函数定义了pagesource参数但未使用,改为直接使用传入的参数,避免冗余。

  4. 优化异常捕获
    替换裸except为具体异常类型,避免隐藏代码中的其他错误。

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup as BeautifulSoup
import time

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))

def parse_html(pagesource, count):
    soup = BeautifulSoup(pagesource, 'html.parser')
    tables = soup.findChildren('table')
    my_table = tables[0]
    table_body = my_table.find('tbody')
    all_rows = table_body.find_all('tr')

    for row in all_rows:
        print(count)
        count += 1
        try:
            path_body = row.find("td", class_="views-field-company-name")
            if not path_body:
                continue
            path = path_body.find("a")['href']
            company_name = path_body.find("a").text.strip()
            print(company_name)

            issue_recepient_office = row.find("td", class_="views-field-field-building").string
            if issue_recepient_office:
                issue_recepient_office = issue_recepient_office.strip()

            detailed_description = row.find("td", class_="views-field-field-detailed-description-2").string
            detailed_description = detailed_description.strip() if detailed_description else ""
        except Exception as e:
            print(f"解析行时出错: {e}")
            continue
    return count  # 返回更新后的count

url = 'https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/compliance-actions-and-activities/warning-letters'

driver.get(url)
count = 1
# 记录第一页的第一个公司名称,用于后续判断页面是否更新
first_company = WebDriverWait(driver, 20).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, ".views-field-company-name a"))
).text.strip()
count = parse_html(driver.page_source, count)

for i in range(0,3):
    # 点击下一页按钮
    next_btn = WebDriverWait(driver, 30).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '#datatable_next a')))
    next_btn.click()
    
    # 等待页面更新:直到第一个公司名称与上一页不同
    WebDriverWait(driver, 30).until(
        lambda d: d.find_element(By.CSS_SELECTOR, ".views-field-company-name a").text.strip() != first_company
    )
    # 更新当前页的第一个公司名称,用于下一次判断
    first_company = driver.find_element(By.CSS_SELECTOR, ".views-field-company-name a").text.strip()
    
    # 解析新页面
    count = parse_html(driver.page_source, count)
        
driver.quit()

效果说明

修改后,每次点击下一页会等待页面内容完全更新,count变量也会持续递增,不会重复输出1-10,而是依次输出所有分页的序号和公司信息。

内容的提问来源于stack exchange,提问作者futurenext110

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 15:55:20