Scrapy中如何处理旧记录?亚马逊商品下架检测咨询
Hey there! Let's work through this expired product detection problem you're dealing with for your Amazon subcategory scraping. Your two initial ideas have clear tradeoffs, so let's break down a better approach that keeps your metadata intact while reliably tracking when items get taken down.
The core idea is to maintain a permanent record of each product's status in a database, so you can compare new crawls against this stored data to spot changes and log timestamps—no more losing metadata or relying on one-off temporary comparisons.
Step-by-Step Implementation
Use a Unique Product Identifier
Stick with Amazon'sASINfor each item—it's a unique, stable identifier that's way more reliable than product titles (which can change). This ensures you'll always match the same product across crawls.Store Product Data & State Persistently
Set up a simple database (SQLite works great for small-scale projects, PostgreSQL for larger ones, or even Redis for faster lookups) with fields like:asin(primary key)- Product metadata (title, price, category, etc.)
last_seen_active(timestamp of the last crawl where the product was available)current_status(enum:active/inactive)marked_inactive_at(timestamp when we first detected it was gone)
Compare New Crawls to Stored Data
After each crawl, run these checks:- Pull all
activeASINs from your database. - Compare to the ASINs from your latest crawl:
- New ASINs not in the database: Add them as
active, setlast_seen_activeto the current timestamp. - ASINs in the database but missing from the crawl: Mark them
inactiveand setmarked_inactive_atto now (but add a retry step first—see below!). - ASINs present in both: Update any changed metadata and refresh
last_seen_active.
- New ASINs not in the database: Add them as
- Pull all
Why This Beats Your Original Ideas
- Better than full re-scraping + data wipe: You keep every bit of historical metadata (like past prices, descriptions) and have a clear record of when items went out of stock—no data loss at all.
- Better than temporary two-crawl comparisons: You don't need to store duplicate crawl data, and you get precise timestamps for when items disappear, not just a "they're gone now" signal.
Extra Tips to Avoid Headaches
- Add a Retry Check
If an ASIN is missing from one crawl, don't immediately mark it inactive. Wait a few hours and run a targeted crawl for just that ASIN—sometimes Amazon's listings are temporarily unavailable due to caching or anti-scraping measures, not actually removed. - Incremental Crawls
For large subcategories, use Amazon's sorting filters (like "Newest Arrivals") to only crawl recently updated items instead of the entire category every time. This saves time and reduces the load on your scraper. - Archive Old Inactive Items
After a few months, move items that have been inactive for a long time to an archive table. This keeps your main database fast and clean while still preserving all historical data.
内容的提问来源于stack exchange,提问作者Josh

