You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法用BeautifulSoup实现Indeed.com翻页爬取,求有效解决方案

Fixing Pagination for Your Indeed Job Scraper

Hey there, let's get that pagination working for your Indeed scraper! The core issue here is how you're handling page navigation—Indeed uses the start query parameter to paginate results (each page shows 10 jobs, so start=0 is page 1, start=10 is page 2, etc.). Let's walk through a revised version of your code that handles pagination properly, plus adds safeguards against getting blocked.

First, Let's Fix Missing Functions & Pagination Logic

I noticed your code references extract_location_from_result and extract_summary_from_result but doesn't define them—let's add those, and wrap everything in a loop that iterates through pages:

import requests
import bs4
from bs4 import BeautifulSoup
import pandas as pd
import time

# Base URL without pagination parameters
BASE_URL = "https://www.indeed.com/jobs?q=Amazon&l="

# Mimic a real browser with a User-Agent (avoids immediate blocks)
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# Define all extraction functions (including missing ones)
def extract_job_title_from_result(soup):
    jobs = []
    for div in soup.find_all(name="div", attrs={"class": "row"}):
        job_title = div.find(name="a", attrs={"data-tn-element": "jobTitle"})
        if job_title:
            jobs.append(job_title["title"])
    return jobs

def extract_company_from_result(soup):
    companies = []
    for div in soup.find_all(name="div", attrs={"class": "row"}):
        company = div.find(name="span", attrs={"class": "company"})
        if company:
            companies.append(company.text.strip())
        else:
            sec_try = div.find(name="span", attrs={"class": "result-link-source"})
            if sec_try:
                companies.append(sec_try.text.strip())
            else:
                companies.append("Unknown")
    return companies

def extract_location_from_result(soup):
    locations = []
    for div in soup.find_all(name="div", attrs={"class": "row"}):
        location = div.find(name="div", attrs={"class": "recJobLoc"})
        if location:
            locations.append(location["data-rc-loc"])
        else:
            locations.append("Unknown")
    return locations

def extract_summary_from_result(soup):
    summaries = []
    for div in soup.find_all(name="div", attrs={"class": "row"}):
        summary = div.find(name="div", attrs={"class": "summary"})
        if summary:
            summaries.append(summary.text.strip())
        else:
            summaries.append("No summary available")
    return summaries

# Initialize lists to store data across all pages
all_jobs = []
all_companies = []
all_locations = []
all_summaries = []

# Set how many pages you want to scrape (adjust as needed)
num_pages = 5

for page in range(num_pages):
    # Calculate the start parameter for each page (10 jobs per page)
    start = page * 10
    url = f"{BASE_URL}&start={start}"
    
    # Send request with headers to avoid blocks
    response = requests.get(url, headers=HEADERS)
    # Check if request was successful
    if response.status_code != 200:
        print(f"Failed to load page {page+1}: Status code {response.status_code}")
        break
    
    soup = BeautifulSoup(response.text, "html.parser")
    
    # Extract data from current page
    jobs = extract_job_title_from_result(soup)
    companies = extract_company_from_result(soup)
    locations = extract_location_from_result(soup)
    summaries = extract_summary_from_result(soup)
    
    # Append to global lists
    all_jobs.extend(jobs)
    all_companies.extend(companies)
    all_locations.extend(locations)
    all_summaries.extend(summaries)
    
    # Add delay to avoid overwhelming the server
    time.sleep(2)
    print(f"Scraped page {page+1} successfully")

# Create final DataFrame
df = pd.DataFrame({
    "job_title": all_jobs,
    "company_name": all_companies,
    "location": all_locations,
    "summary": all_summaries
})

# Optional: Save to CSV
df.to_csv("amazon_jobs.csv", index=False)
print(f"Scraping complete! Total jobs collected: {len(df)}")

Key Changes That Make This Work:

  • Pagination Handling: We use the start parameter in the URL, incrementing by 10 each page (matches Indeed's 10-jobs-per-page layout).
  • Request Headers: Added a User-Agent to mimic a real browser, which avoids immediate blocks from Indeed's anti-scraping tools.
  • Error Checking: Verifies the request returns a 200 (success) status before proceeding.
  • Rate Limiting: Added time.sleep(2) between requests to prevent hitting the server too quickly, reducing the risk of being blocked.
  • Robust Extraction: Added missing location/summary functions, plus fallbacks like "Unknown" for cases where data is missing.

Important Notes:

  • Indeed regularly updates its page structure—if selectors stop working, inspect the page source and update class names/attributes in your functions.
  • Don't scrape too many pages too fast—this can get your IP blocked. Adjust num_pages and time.sleep duration as needed.
  • Always check Indeed's robots.txt to ensure you're allowed to scrape the data you're targeting.

内容的提问来源于stack exchange,提问作者Lesnar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:09:16