如何通过正则表达式匹配爬取Alexa技能商店指定前缀的技能页面链接?
Got it, let's walk through exactly how to solve this. You want to target all skill pages that follow the structure https://www.alexaskillstore.com/[Category]/[SkillID] (like your example Business-Leadership-Series/B078LNGS5T), then pull comments from each. Here's a practical, step-by-step approach using BeautifulSoup and regex (plus some extra tips to avoid common pitfalls):
Step 1: Set Up Your Tools
First, make sure you have the necessary libraries installed. If not, run this in your terminal:
pip install requests beautifulsoup4
Step 2: Define the URL Pattern & Fetch Valid Skill Links
We'll use a regex pattern to spot valid skill pages, and handle both absolute and relative links (since some site links might skip the full domain). Here's the code:
import requests from bs4 import BeautifulSoup import re from urllib.parse import urljoin # Base domain of the skill store base_domain = "https://www.alexaskillstore.com" # Regex pattern to match skill pages: domain → category → unique skill ID # Breakdown: # ^ = start of string (ensures we don't match partial URLs) # https://www\.alexaskillstore\.com/ = exact domain (dots escaped for regex) # [^/]+ = one or more non-slash characters (matches the category name) # / = separator between category and skill ID # [^/]+ = one or more non-slash characters (matches the unique skill ID) # $ = end of string (ensures no extra path segments) skill_url_regex = re.compile(r'^https://www\.alexaskillstore\.com/[^/]+/[^/]+$') # Start by fetching a page that lists skills (e.g., homepage or a specific category page) start_page_url = base_domain # Or use a category URL like "/Business-Leadership-Series" response = requests.get(start_page_url, headers={"User-Agent": "Mozilla/5.0"}) response.raise_for_status() # Catch HTTP errors like 404 or 500 soup = BeautifulSoup(response.text, "html.parser") # Use a set to collect unique skill URLs (avoids duplicate pages) unique_skill_urls = set() for link in soup.find_all("a", href=True): # Convert relative links (e.g., "/Business-Leadership-Series/B078LNGS5T") to absolute URLs full_url = urljoin(base_domain, link["href"]) # Check if the URL matches our skill page pattern if skill_url_regex.match(full_url): unique_skill_urls.add(full_url) print(f"Found {len(unique_skill_urls)} unique skill pages to scrape!")
Step 3: Scrape Comments from Each Skill Page
Now that you have all valid skill URLs, loop through them to extract comments. You'll need to adjust the HTML selectors based on the actual structure of the site's comment section (use your browser's dev tools to inspect elements):
import time # Add delay between requests to avoid overwhelming the server for skill_url in unique_skill_urls: print(f"\nScraping comments from: {skill_url}") skill_response = requests.get(skill_url, headers={"User-Agent": "Mozilla/5.0"}) skill_soup = BeautifulSoup(skill_response.text, "html.parser") # Example selectors (replace these with the actual classes/IDs from the site) comment_containers = skill_soup.find_all("div", class_="review-item") if not comment_containers: print("No comments found for this skill.") time.sleep(1) # Add 1-second delay before next request continue for idx, comment in enumerate(comment_containers, 1): # Extract comment details author = comment.find("span", class_="review-author").get_text(strip=True) date = comment.find("span", class_="review-date").get_text(strip=True) text = comment.find("p", class_="review-text").get_text(strip=True) print(f"Comment {idx}:") print(f" Author: {author}") print(f" Date: {date}") print(f" Text: {text}\n") time.sleep(1) # Add delay between skill pages
Key Tips to Avoid Headaches
- Respect
robots.txt: Checkhttps://www.alexaskillstore.com/robots.txtfirst to make sure scraping is allowed for the pages you're targeting. - Refine the Regex: If you only want to target specific categories or skill ID formats (e.g., IDs starting with
B), tweak the pattern. For example:# Match only skills in the Business-Leadership-Series category targeted_regex = re.compile(r'^https://www\.alexaskillstore\.com/Business-Leadership-Series/[^/]+$') - Handle Pagination: If skills or comments are spread across multiple pages, detect and follow "Next" links using similar regex or selector logic.
内容的提问来源于stack exchange,提问作者Aadarsh Gupta

