无法用BeautifulSoup实现Indeed.com翻页爬取,求有效解决方案
Fixing Pagination for Your Indeed Job Scraper
Hey there, let's get that pagination working for your Indeed scraper! The core issue here is how you're handling page navigation—Indeed uses the start query parameter to paginate results (each page shows 10 jobs, so start=0 is page 1, start=10 is page 2, etc.). Let's walk through a revised version of your code that handles pagination properly, plus adds safeguards against getting blocked.
First, Let's Fix Missing Functions & Pagination Logic
I noticed your code references extract_location_from_result and extract_summary_from_result but doesn't define them—let's add those, and wrap everything in a loop that iterates through pages:
import requests import bs4 from bs4 import BeautifulSoup import pandas as pd import time # Base URL without pagination parameters BASE_URL = "https://www.indeed.com/jobs?q=Amazon&l=" # Mimic a real browser with a User-Agent (avoids immediate blocks) HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # Define all extraction functions (including missing ones) def extract_job_title_from_result(soup): jobs = [] for div in soup.find_all(name="div", attrs={"class": "row"}): job_title = div.find(name="a", attrs={"data-tn-element": "jobTitle"}) if job_title: jobs.append(job_title["title"]) return jobs def extract_company_from_result(soup): companies = [] for div in soup.find_all(name="div", attrs={"class": "row"}): company = div.find(name="span", attrs={"class": "company"}) if company: companies.append(company.text.strip()) else: sec_try = div.find(name="span", attrs={"class": "result-link-source"}) if sec_try: companies.append(sec_try.text.strip()) else: companies.append("Unknown") return companies def extract_location_from_result(soup): locations = [] for div in soup.find_all(name="div", attrs={"class": "row"}): location = div.find(name="div", attrs={"class": "recJobLoc"}) if location: locations.append(location["data-rc-loc"]) else: locations.append("Unknown") return locations def extract_summary_from_result(soup): summaries = [] for div in soup.find_all(name="div", attrs={"class": "row"}): summary = div.find(name="div", attrs={"class": "summary"}) if summary: summaries.append(summary.text.strip()) else: summaries.append("No summary available") return summaries # Initialize lists to store data across all pages all_jobs = [] all_companies = [] all_locations = [] all_summaries = [] # Set how many pages you want to scrape (adjust as needed) num_pages = 5 for page in range(num_pages): # Calculate the start parameter for each page (10 jobs per page) start = page * 10 url = f"{BASE_URL}&start={start}" # Send request with headers to avoid blocks response = requests.get(url, headers=HEADERS) # Check if request was successful if response.status_code != 200: print(f"Failed to load page {page+1}: Status code {response.status_code}") break soup = BeautifulSoup(response.text, "html.parser") # Extract data from current page jobs = extract_job_title_from_result(soup) companies = extract_company_from_result(soup) locations = extract_location_from_result(soup) summaries = extract_summary_from_result(soup) # Append to global lists all_jobs.extend(jobs) all_companies.extend(companies) all_locations.extend(locations) all_summaries.extend(summaries) # Add delay to avoid overwhelming the server time.sleep(2) print(f"Scraped page {page+1} successfully") # Create final DataFrame df = pd.DataFrame({ "job_title": all_jobs, "company_name": all_companies, "location": all_locations, "summary": all_summaries }) # Optional: Save to CSV df.to_csv("amazon_jobs.csv", index=False) print(f"Scraping complete! Total jobs collected: {len(df)}")
Key Changes That Make This Work:
- Pagination Handling: We use the
startparameter in the URL, incrementing by 10 each page (matches Indeed's 10-jobs-per-page layout). - Request Headers: Added a
User-Agentto mimic a real browser, which avoids immediate blocks from Indeed's anti-scraping tools. - Error Checking: Verifies the request returns a 200 (success) status before proceeding.
- Rate Limiting: Added
time.sleep(2)between requests to prevent hitting the server too quickly, reducing the risk of being blocked. - Robust Extraction: Added missing location/summary functions, plus fallbacks like "Unknown" for cases where data is missing.
Important Notes:
- Indeed regularly updates its page structure—if selectors stop working, inspect the page source and update class names/attributes in your functions.
- Don't scrape too many pages too fast—this can get your IP blocked. Adjust
num_pagesandtime.sleepduration as needed. - Always check Indeed's robots.txt to ensure you're allowed to scrape the data you're targeting.
内容的提问来源于stack exchange,提问作者Lesnar
相关产品推荐
相关产品推荐

