亚马逊图书页面排名爬取代码优化方案咨询
Hey Sara, great question—your current approach works for the specific page right now, but hardcoding index slices like [24:32] and relying on fixed list positions is super fragile. Amazon frequently tweaks their page HTML, and even a tiny change could break your scraper. Let's refactor this to use text-based targeting and regex extraction, which is way more robust.
What's Wrong with the Original Code?
- Fixed index slices:
soup_rank_detail[1][24:32]assumes the rank is always at that exact character position, which will fail if Amazon adds/removes text before the rank. - Dependent on list order:
soup_rank[1]assumes the rank block is the second item in that list—again, easy to break with layout changes.
Optimized Solution
Here's a revised version that uses text matching to find the rank section and regex to extract the actual rank value:
import re from urllib.request import Request, urlopen from bs4 import BeautifulSoup def scrap_rank_amz(link): # Fetch page content with basic error handling url = link request = Request(url, headers={"User-agent": "Mozilla/5.0"}) try: html = urlopen(request) except Exception as e: print(f"Failed to load page: {str(e)}") return None # Parse HTML soup = BeautifulSoup(html, "html.parser") # Locate the Best Sellers Rank section by text content rank_text = None # Iterate through all detail bullets to find the one with rank info for bullet in soup.find_all(class_="a-list-item"): clean_text = bullet.get_text(strip=True) if "Best Sellers Rank" in clean_text: rank_text = clean_text break if not rank_text: print("Could not locate Best Sellers Rank section on the page") return None # Extract the rank using regex (handles numbers with commas) # Matches patterns like "#12,345 in Books" rank_match = re.search(r"#([\d,]+) in", rank_text) if rank_match: # Remove commas and return as a clean string (or convert to int if needed) return rank_match.group(1).replace(",", "") else: print("Failed to extract rank from the section text") return None
Key Improvements
- Text-based targeting: Instead of relying on class positions, we look for the exact phrase "Best Sellers Rank"—this is far less likely to change than HTML structure.
- Regex extraction: The regex
#([\d,]+) inmatches any rank format (with or without commas) regardless of its position in the text, eliminating the need for fragile slicing. - Error handling: Added try/except blocks for page requests and checks for missing rank sections, so your function won't crash unexpectedly.
Testing with Your Example Link
If you run scrap_rank_amz("https://www.amazon.com/Moonshine-Magic-Southern-Charms-Mystery-ebook/dp/B078SZLXB3"), it should return the clean rank number (e.g., something like "12345" without commas or extra text).
Quick Note
Remember that Amazon has anti-scraping measures—make sure to add delays between requests and respect their robots.txt rules. For production use, you might want to consider using Amazon's official Product Advertising API if you have access, as it's a more reliable and compliant way to get this data.
内容的提问来源于stack exchange,提问作者Sara U.

