使用Scrapy抓取letgo无限滚动网站的异步数据问题
Scraping Letgo US Used Items: Handling Async JSON Pagination
Hey there! Looks like you're working on a cool personal project scraping local used items from Letgo's US marketplace—nice work. I’ve tackled similar async pagination setups before, so let’s walk through how to nail that JSON API since the video you referenced missed some key details.
1. Breakdown of the API Parameters
First, let’s clarify what each part of that API endpoint does:
country_code=US: This locks your requests to the US region, so you don’t accidentally pull items from other countries.offset: This is your pagination marker. The initial load usesoffset=0to grab the first 30 items. For subsequent scrolls, you’ll increment this by the number of items returned in the last request (you mentioned 15 per scroll after the first, so offset goes 30 → 45 → 60, etc.).quadkey: This is the critical geolocation parameter—it ensures you’re fetching items from your target local area. You can grab this value by inspecting the initial network request in your browser’s DevTools when you load the Letgo page for your region.
Example API call:https://search-products-pwa.letgo.com/api/products?country_code=US&offset=0&quadkey=03...
2. Pagination Workflow to Follow
- Start with the initial batch: Fire a request with
offset=0to get the first 30 items. - Loop for additional items: Keep making requests, incrementing the
offseteach time by the count of items from the previous response. - Know when to stop: You’ll want to halt the loop when the API returns an empty
productsarray, or when you’ve fetched all items indicated by thetotalfield (if it’s included in the JSON response).
3. Pro Tips to Avoid Getting Blocked
Letgo has anti-scraping measures, so make sure you mimic real user behavior:
- Use realistic headers: Copy headers like
User-Agent,Accept, andRefererfrom your browser’s network tab (when you inspect the API call) and include them in your requests. This makes your scraper look like a regular browser. - Add delays: Don’t spam the API—throw in a 1-2 second delay between requests. This prevents you from triggering rate limits.
- Use a session: Tools like
requests.Session()in Python let you maintain cookies across requests, which helps you appear more like a genuine user.
4. Quick Example Code (Python)
Here’s a simple snippet to get you started with fetching and parsing the items:
import requests import time # Set up session and headers to mimic a browser session = requests.Session() headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "application/json, text/plain, */*" } # Replace with your target quadkey target_quadkey = "your-local-quadkey" offset = 0 all_scraped_items = [] while True: api_url = f"https://search-products-pwa.letgo.com/api/products?country_code=US&offset={offset}&quadkey={target_quadkey}" response = session.get(api_url, headers=headers) # Handle possible errors (e.g., 403 Forbidden) if response.status_code != 200: print(f"Request failed with status code {response.status_code}") break data = response.json() current_products = data.get("products", []) # Stop if no more items are returned if not current_products: print("No more items to scrape!") break all_scraped_items.extend(current_products) # Update offset for next page offset += len(current_products) print(f"Scraped {len(current_products)} items. Total so far: {len(all_scraped_items)}") # Add delay to avoid blocking time.sleep(1.5) # Process your scraped items (example: print titles and prices) for item in all_scraped_items: print(f"Title: {item['title']} | Price: ${item['price']}")
内容的提问来源于stack exchange,提问作者Keenan Burke-Pitts
相关产品推荐
相关产品推荐

