BS4+Selenium爬虫缩进异常:弹窗表格多页数据爬取失败
Fixing Your Pagination Scraper for formularylookup.com
Let's work through this step by step—your main issues stem from incorrect indentation (which causes missed pages or repeated data) and failing to re-parse the page after navigating to new content. Here's how to fix it:
Key Issues in Your Current Code
- You only parse the page once at the start, but after clicking to a new page, the HTML updates—you need to re-fetch and re-parse the page source every time you load a new page.
- Your indentation has the "click next page" step before scraping the current page, so you're skipping the first page entirely.
- Your loop stops short of the last page (using
range(1, int(max_page))excludes the final page number).
Corrected Code with Proper Indentation & Logic
from bs4 import BeautifulSoup import time import pandas as pd # Assuming self.browser is your initialized Selenium WebDriver instance total = [] # First, grab the total number of pages from the initial loaded page initial_html = self.browser.page_source soup = BeautifulSoup(initial_html, "lxml") grid_container = soup.find("div", {"id":"lookupDetailsGrid"}) max_page = int(grid_container.find("a", {"title":"Go to the last page"})["data-page"]) # Loop through every page (1 to max_page, inclusive) for page_num in range(1, max_page + 1): # Re-fresh and re-parse the current page's HTML (critical after navigation) current_page_html = self.browser.page_source current_soup = BeautifulSoup(current_page_html, "lxml") # Scrape all 50 rows on the current page for detail_row in current_soup.find_all("tr", {"class":"k-master-row"}): plan_name = detail_row.find("td", {"class":"col-plan"}).text.strip() total.append({"Plan": plan_name}) print(f"Page {page_num}: {plan_name}") # Only click next page if we're not on the final page if page_num != max_page: self.browser.find_element_by_xpath('//*[@id="lookupDetailsGrid"]/div[3]/a[3]').click() time.sleep(3) # Adjust this wait time if needed—ensure the page fully loads # Convert collected data to DataFrame df = pd.DataFrame(total) print(f"Total records scraped: {len(df)}") print(df)
What Changed & Why
- Re-parsing on every iteration: We fetch fresh
page_sourceand create a newBeautifulSoupobject for each page. This ensures we're scraping the current page's content, not the initial one loaded at the start. - Indentation order: We scrape the current page first, then click to the next page (only if we're not on the last page). This fixes the "missing first page" issue.
- Including the last page: Using
range(1, max_page + 1)ensures we loop through all 27 pages (instead of stopping at 26). - Cleaner data: Added
.strip()to remove extra whitespace/newlines from the plan names.
Bonus: More Reliable Waiting (Optional)
Instead of fixed time.sleep(), use Selenium's WebDriverWait to wait for elements to load—this makes your scraper more robust if page load times vary:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # Replace the click + sleep with this: if page_num != max_page: # Wait for next page button to be clickable next_btn = WebDriverWait(self.browser, 10).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="lookupDetailsGrid"]/div[3]/a[3]')) ) next_btn.click() # Wait for the page content to update WebDriverWait(self.browser, 10).until( EC.staleness_of(grid_container) )
内容的提问来源于stack exchange,提问作者doomdaam
相关产品推荐
相关产品推荐

