使用BeautifulSoup爬取Foursquare评论时输出数据量波动的原因排查
Hey there! Let's break down why you're seeing fluctuating numbers (sometimes 144, sometimes 56) when scraping reviews from Foursquare. The issue is almost certainly a mix of how Foursquare serves content and your current scraping approach, rather than just plain network issues. Here's the breakdown:
1. Dynamic/Lazy Loading (Most Likely Culprit)
Foursquare doesn't load all reviews in the initial HTML response. Instead, it uses lazy loading—only loading the first batch of reviews (like the 56 you're seeing sometimes) initially, then fetching more as the user scrolls down the page.
Your current code uses requests to grab the raw HTML, which only includes the reviews loaded on the first page load. The times you get all 144 reviews might be rare edge cases where the server accidentally returned fully rendered content, or maybe you ran the script right after a browser session that had already loaded all reviews (but that's not reliable long-term).
2. Server-Side Content Filtering/Caching
Foursquare might serve different content based on your request's headers. If your requests call uses the default, generic user-agent, the site might flag you as a bot and return a truncated set of reviews. Alternatively, their caching system might sometimes serve a cached partial page instead of the full content.
3. Network Issues (Least Likely)
While network instability could cause incomplete page downloads, this usually leads to parsing errors or zero results—not a consistent partial count. So this is probably not the main issue.
Fixes to Get Consistent Results
Option 1: Simulate a Browser to Handle Lazy Loading
Use tools like Selenium or Playwright to mimic a real user scrolling the page, which triggers the loading of all reviews. Here's a quick Selenium example:
from selenium import webdriver import time from bs4 import BeautifulSoup url = "https://foursquare.com/v/pantai-klebang/4c7c12d22d3ba14318e595d0" driver = webdriver.Chrome() # Ensure ChromeDriver is installed and in your PATH driver.get(url) # Scroll repeatedly until no new content loads last_scroll_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll to the bottom of the page driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for new reviews to load time.sleep(2) # Check if the page height stopped changing (no more content) new_scroll_height = driver.execute_script("return document.body.scrollHeight") if new_scroll_height == last_scroll_height: break last_scroll_height = new_scroll_height # Parse the fully loaded page soup = BeautifulSoup(driver.page_source, "html.parser") containers = soup.findAll("li", {"class": "tip"}) print(f"Successfully scraped {len(containers)} reviews") # Clean up the browser session driver.quit()
Option 2: Mimic Real Browser Headers
If you want to stick with requests, add realistic headers to avoid being flagged as a bot. This might help the server return more content (though it still won't handle lazy loading):
import requests from bs4 import BeautifulSoup url = "https://foursquare.com/v/pantai-klebang/4c7c12d22d3ba14318e595d0" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5", "Referer": "https://www.google.com/" } r = requests.get(url, headers=headers) soup = BeautifulSoup(r.content, "html.parser") containers = soup.findAll("li", {"class": "tip"}) print(f"Found {len(containers)} reviews")
Option 3: Find the Underlying API (Advanced)
Use your browser's DevTools to inspect network requests while scrolling reviews. You'll likely find an AJAX endpoint that Foursquare uses to load more reviews. You can call this endpoint directly with requests to fetch all reviews without simulating a browser—this is faster and more efficient.
内容的提问来源于stack exchange,提问作者ali

