使用BeautifulSoup爬取多页遇阻:目标网站分页无有效链接求助
Hey there, let's work through this pagination problem you're facing! That empty href on the page number links is a dead giveaway that the site is using client-side JavaScript to load subsequent pages instead of regular anchor links. Here are the most practical ways to get around this:
1. Dig for the Hidden Pagination API (Most Efficient)
The easiest way to bypass the empty links is to find the actual API endpoint that loads the event data when you click a page number. Here's how:
- Open your browser's DevTools (F12), switch to the Network tab, and make sure "XHR" or "Fetch" is selected.
- Go back to the events page and click the "2" page button. You'll see a new request pop up in the Network tab—this is the API call that fetches the second page's data.
- Check the request URL (it might look like
https://concreteplayground.com/auckland/events?page=2or something similar with a page parameter) and the response (it could be JSON or raw HTML). - Once you have that URL pattern, you can directly construct requests for each page (e.g.,
page=3,page=4) usingrequestsin Python, then parse the response with BeautifulSoup like you did for the first page.
2. Simulate Browser Clicks with Automation Tools
If you can't track down the API, you can mimic a real user's behavior using tools like Selenium or Playwright. This lets the browser handle the JavaScript loading for you:
Example with Selenium:
First, install the package and grab the appropriate browser driver (e.g., ChromeDriver for Chrome):
pip install selenium
Then use this code skeleton to crawl multiple pages:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # Initialize the browser driver = webdriver.Chrome() driver.get("https://concreteplayground.com/auckland/events") # Set how many pages you want to crawl max_pages = 5 for page in range(1, max_pages + 1): # Parse the current page's content soup = BeautifulSoup(driver.page_source, "html.parser") # Add your existing data extraction logic here (e.g., scrape event titles, dates) events = soup.find_all("div", class_="event-item") for event in events: # Process each event (example: print title) title = event.find("h3") if title: print(title.get_text(strip=True)) # Try to click the next page button try: # Wait for the page number link to be clickable (adjust selector if needed) next_page_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, f'a.page-numbers:contains("{page + 1}")')) ) next_page_btn.click() # Add a small delay to let the page load time.sleep(2) # Wait for the page to reload (wait until old events are no longer present) WebDriverWait(driver, 10).until( EC.staleness_of(events[0]) ) except: # No more pages to load, exit the loop print("No more pages available.") break # Clean up driver.quit()
3. Test Manual URL Parameter Tweaks
Sometimes even with empty href attributes, the site still supports traditional query parameters. Try manually typing https://concreteplayground.com/auckland/events?page=2 into your browser's address bar—if it loads the second page, you're golden! Just loop through page numbers in your requests calls, appending ?page=N to the base URL each time.
Quick Notes to Avoid Blocking
- Add a small delay between requests (using
time.sleep(2)for example) to avoid triggering anti-scraping measures. - Set a realistic
User-Agentheader in yourrequestscalls to mimic a real browser.
内容的提问来源于stack exchange,提问作者Mohan K

