You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BS4高效爬取多页面?寻求非硬编码的优雅实现方案

Elegant Multi-Page Scraping with Python + BeautifulSoup

Hey there! I totally get where you’re coming from—hardcoding page numbers feels clunky and isn’t scalable if the number of pages ever changes. Let’s walk through two clean, maintainable approaches to handle this without those rigid page parameters.

Approach 1: Loop Through Predictable Page Patterns

If your target URLs follow a simple, consistent structure (like https://example.com/data?page=1, https://example.com/data?page=2), you can define a dynamic range of pages and loop through them. This keeps your code flexible if you ever need to adjust the number of pages later.

Here’s a sample implementation that builds on your existing setup:

import requests
from bs4 import BeautifulSoup

def scrape_single_page(page_num):
    # Replace with your actual URL pattern
    url = f"https://your-target-site.com/desired-content?page={page_num}"
    response = requests.get(url)
    response.raise_for_status()  # Catch and handle HTTP errors early
    
    soup = BeautifulSoup(response.text, "html.parser")
    # Drop your existing data extraction logic here
    # Example: extracted_items = soup.find_all("article", class_="content-card")
    return extracted_items

# Collect data across pages 1-3
all_collected_data = []
for page in range(1, 4):  # Range is exclusive of the end value, so this covers 1,2,3
    page_data = scrape_single_page(page)
    all_collected_data.extend(page_data)
    print(f"Scraped page {page} successfully!")

# Do whatever you need with the combined data
print(f"Total items collected: {len(all_collected_data)}")

If the website has a pagination bar with a "Next" button, you can dynamically scrape the next page’s URL instead of relying on a fixed page range. This is smarter because it works even if the total number of pages changes over time (like if new content gets added).

Here’s how that might look:

import requests
from bs4 import BeautifulSoup

def scrape_page_and_get_next(url):
    response = requests.get(url)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")
    
    # Extract your target data here
    extracted_data = soup.find_all("div", class_="target-content")
    
    # Find the next page link (adjust the selector to match the site's HTML)
    next_page_btn = soup.find("a", string="Next")  # Or use a class like "pagination-next"
    if next_page_btn:
        next_page_url = next_page_btn["href"]
        # Convert relative links to full URLs if needed
        if not next_page_url.startswith("http"):
            next_page_url = f"https://your-target-site.com{next_page_url}"
        return extracted_data, next_page_url
    else:
        # No more pages to scrape
        return extracted_data, None

# Start with the first page URL
current_page_url = "https://your-target-site.com/desired-content?page=1"
all_data = []

while current_page_url:
    page_data, current_page_url = scrape_page_and_get_next(current_page_url)
    all_data.extend(page_data)
    if current_page_url:
        print(f"Moving to next page: {current_page_url}")
    else:
        print("No more pages left to scrape!")

print(f"Total data collected: {len(all_data)}")

Quick Efficiency Tips

  • Add a small delay (time.sleep(1)) between requests to avoid overwhelming the server (always respect the site’s robots.txt rules!).
  • Use requests.Session() to persist cookies across requests—this is faster and helps with sites that track sessions.
  • Wrap request logic in try/except blocks to handle network blips gracefully.

内容的提问来源于stack exchange,提问作者Nick Read

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:24:08