You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BS4+Selenium爬虫缩进异常:弹窗表格多页数据爬取失败

Fixing Your Pagination Scraper for formularylookup.com

Let's work through this step by step—your main issues stem from incorrect indentation (which causes missed pages or repeated data) and failing to re-parse the page after navigating to new content. Here's how to fix it:

Key Issues in Your Current Code

  1. You only parse the page once at the start, but after clicking to a new page, the HTML updates—you need to re-fetch and re-parse the page source every time you load a new page.
  2. Your indentation has the "click next page" step before scraping the current page, so you're skipping the first page entirely.
  3. Your loop stops short of the last page (using range(1, int(max_page)) excludes the final page number).

Corrected Code with Proper Indentation & Logic

from bs4 import BeautifulSoup
import time
import pandas as pd

# Assuming self.browser is your initialized Selenium WebDriver instance
total = []

# First, grab the total number of pages from the initial loaded page
initial_html = self.browser.page_source
soup = BeautifulSoup(initial_html, "lxml")
grid_container = soup.find("div", {"id":"lookupDetailsGrid"})
max_page = int(grid_container.find("a", {"title":"Go to the last page"})["data-page"])

# Loop through every page (1 to max_page, inclusive)
for page_num in range(1, max_page + 1):
    # Re-fresh and re-parse the current page's HTML (critical after navigation)
    current_page_html = self.browser.page_source
    current_soup = BeautifulSoup(current_page_html, "lxml")
    
    # Scrape all 50 rows on the current page
    for detail_row in current_soup.find_all("tr", {"class":"k-master-row"}):
        plan_name = detail_row.find("td", {"class":"col-plan"}).text.strip()
        total.append({"Plan": plan_name})
        print(f"Page {page_num}: {plan_name}")
    
    # Only click next page if we're not on the final page
    if page_num != max_page:
        self.browser.find_element_by_xpath('//*[@id="lookupDetailsGrid"]/div[3]/a[3]').click()
        time.sleep(3)  # Adjust this wait time if needed—ensure the page fully loads

# Convert collected data to DataFrame
df = pd.DataFrame(total)
print(f"Total records scraped: {len(df)}")
print(df)

What Changed & Why

  • Re-parsing on every iteration: We fetch fresh page_source and create a new BeautifulSoup object for each page. This ensures we're scraping the current page's content, not the initial one loaded at the start.
  • Indentation order: We scrape the current page first, then click to the next page (only if we're not on the last page). This fixes the "missing first page" issue.
  • Including the last page: Using range(1, max_page + 1) ensures we loop through all 27 pages (instead of stopping at 26).
  • Cleaner data: Added .strip() to remove extra whitespace/newlines from the plan names.

Bonus: More Reliable Waiting (Optional)

Instead of fixed time.sleep(), use Selenium's WebDriverWait to wait for elements to load—this makes your scraper more robust if page load times vary:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# Replace the click + sleep with this:
if page_num != max_page:
    # Wait for next page button to be clickable
    next_btn = WebDriverWait(self.browser, 10).until(
        EC.element_to_be_clickable((By.XPATH, '//*[@id="lookupDetailsGrid"]/div[3]/a[3]'))
    )
    next_btn.click()
    # Wait for the page content to update
    WebDriverWait(self.browser, 10).until(
        EC.staleness_of(grid_container)
    )

内容的提问来源于stack exchange,提问作者doomdaam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 13:23:11