You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何循环多URL爬取TripAdvisor的评分与评论?

Hey Jerry, nice work getting the first page's data and all those pagination URLs sorted—you're already halfway there! Let's walk through the most efficient, robust ways to loop through those URLs and scrape ratings/reviews in bulk.

Optimal Bulk Scraping Strategies for TripAdvisor

1. Go Async for Faster Throughput

Synchronous requests (waiting for one page to load before the next) will crawl slowly if you have dozens or hundreds of pages. Asynchronous requests let you fetch multiple pages at the same time, cutting down total runtime drastically.

For Python, the aiohttp library is the go-to here. Pair it with asyncio to manage concurrent tasks, and add a semaphore to control how many requests you send at once (don't flood the server!). Here's a simplified example:

import aiohttp
import asyncio
import random
from bs4 import BeautifulSoup

async def scrape_single_page(session, url):
    # Add random delay to mimic human behavior
    await asyncio.sleep(random.uniform(1.5, 3.5))
    
    # Rotate User-Agents if you want extra stealth (keep a list of real ones)
    headers = {
        "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    try:
        async with session.get(url, headers=headers) as resp:
            if resp.status != 200:
                print(f"⚠️ Failed to load {url} (status: {resp.status})")
                return []
            
            html = await resp.text()
            soup = BeautifulSoup(html, "html.parser")
            
            # Replace this with your existing extraction logic
            page_reviews = []
            for review in soup.select(".review-container"):
                # Extract rating (from the bubble class, e.g., "ui_bubble_rating_40" → 4.0)
                rating = int(review.select_one(".ui_bubble_rating")["class"][1].split("_")[1]) / 10
                # Extract review text
                comment = review.select_one(".partial_entry").get_text(strip=True)
                page_reviews.append({"rating": rating, "comment": comment})
            
            return page_reviews
    
    except Exception as e:
        print(f"❌ Error scraping {url}: {str(e)}")
        return []

async def bulk_scrape(pagination_urls):
    async with aiohttp.ClientSession() as session:
        # Limit concurrent requests to avoid triggering anti-scraping
        semaphore = asyncio.Semaphore(8)
        
        async def bounded_scrape(url):
            async with semaphore:
                return await scrape_single_page(session, url)
        
        # Create tasks for all URLs
        tasks = [bounded_scrape(url) for url in pagination_urls]
        # Run all tasks and collect results
        all_results = await asyncio.gather(*tasks)
        
        # Flatten the list of lists into one big list of reviews
        all_reviews = [item for sublist in all_results for item in sublist]
        
        # Save immediately (don't wait till the end!)
        # Example: Write to CSV
        import csv
        with open("tripadvisor_reviews.csv", "w", newline="", encoding="utf-8") as f:
            writer = csv.DictWriter(f, fieldnames=["rating", "comment"])
            writer.writeheader()
            writer.writerows(all_reviews)
        
        print(f"✅ Done! Scraped {len(all_reviews)} total reviews.")

if __name__ == "__main__":
    # Replace with your actual list of pagination URLs
    your_pagination_urls = ["https://www.tripadvisor.com/...page1", "...page2", "...page3"]
    asyncio.run(bulk_scrape(your_pagination_urls))

2. Beat TripAdvisor's Anti-Scraping Measures

TripAdvisor is strict about bots—skip these steps and you'll get blocked fast:

  • Rotate User-Agents: Keep a list of real browser UA strings and pick one randomly per request.
  • Random Delays: Never use fixed wait times; random intervals between 1-4 seconds look more human.
  • Limit Concurrency: Start with 5-10 concurrent requests, adjust if you get 403/503 errors.
  • Proxy Rotation (for large-scale scraping): If you're crawling hundreds of pages, use residential proxies to avoid IP bans.
  • Respect robots.txt: Check TripAdvisor's robots.txt to make sure you're allowed to scrape the pages you're targeting.

3. Add Retries for Flaky Requests

Network blips or temporary blocks happen. Use a library like tenacity to automatically retry failed requests:

from tenacity import retry, stop_after_attempt, wait_random_exponential

@retry(stop=stop_after_attempt(3), wait=wait_random_exponential(multiplier=1, max=10))
async def scrape_single_page(session, url):
    # Your existing scraping logic here

This will retry failed requests up to 3 times, with increasing wait intervals.

4. Persist Data Incrementally

Don't wait until all pages are scraped to save data—if your program crashes, you'll lose everything. Instead:

  • Append each page's reviews to a CSV file immediately after scraping.
  • For larger datasets, use a lightweight database like SQLite to store entries as you go.

5. Log Everything

Use Python's logging module to track successes, failures, and errors. This makes debugging a breeze if something goes wrong:

import logging
logging.basicConfig(filename="scraper.log", level=logging.INFO, format="%(asctime)s - %(message)s")

# Inside your scrape function:
logging.info(f"Successfully scraped {len(page_reviews)} reviews from {url}")
logging.error(f"Failed to scrape {url}: {str(e)}")

Wrapping up: Start small with a handful of pages to test your logic and anti-scraping measures, then scale up. Adjust concurrency and delays based on how TripAdvisor responds—if you get blocked, dial back the speed.

内容的提问来源于stack exchange,提问作者Jerry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:02:47