Python无报错但Weedmaps爬虫返回空结果问题排查求助
Let's break down why your crawler is returning [[], [], [], [], [], [], [], [], []] even though your get_page_link and product_data functions work in isolation. Here are the key issues and actionable fixes:
1. URL Protocol & Redirect Mismatch
In your main function, you're using http://www.weedmaps.com for search requests. Weedmaps enforces HTTPS, so this request gets redirected to the HTTPS version—but the redirected page often has a different structure than what your selector expects, or the redirect process strips out the content you're trying to scrape.
Fix: Update the base URL in main to use HTTPS and drop the www subdomain (Weedmaps redirects from www to the root domain automatically anyway):
urlz = f'https://weedmaps.com/search?entryType=home%20page%20product%20card&filter%5BboundingRadius%5D={distance}mi&page={x}'
2. Fragile Dynamic Class Names
Your CSS selectors rely on auto-generated hashed class names like styles__NameRatingWrap-j5iyiv-15.eaLQmf and styled-components__Price-sc-1fbw3xt-15.lbyswm. Weedmaps uses tools like styled-components that generate these class names, which change frequently (even between page loads or site deployments). When you tested functions individually, the class names were valid—but by the time you ran main, they likely changed, or the search page uses different dynamic classes than product pages.
Fix: Replace dynamic class names with stable, structure-based selectors:
- For
get_page_link, target links that explicitly point to product pages (they start with/products/):def get_page_link(url): r = requests.get(url, headers=agent, timeout=3) sp = BeautifulSoup(r.text, 'lxml') baseurl = "https://weedmaps.com" # Target all links pointing to product pages, deduplicate to avoid repeats links = sp.select('a[href^="/products/"]') unique_links = list({link.attrs['href'] for link in links}) return [baseurl + link for link in unique_links] - For
product_data, use flexible selectors and add safety checks to avoidAttributeErrorif elements are missing:
Note: Thedef product_data(url): r = requests.get(url, headers=agent, timeout=3) sp = BeautifulSoup(r.text, 'lxml') product = { 'Title': sp.select_one('h1').text.strip() if sp.select_one('h1') else 'N/A', 'Brand': sp.select_one('div[class*="ProductCategoryBrand"] a').text.strip() if sp.select_one('div[class*="ProductCategoryBrand"] a') else 'N/A', 'Price': sp.select_one('div[class*="Price"]').text.strip() if sp.select_one('div[class*="Price"]') else 'N/A', 'Pick_up_location': sp.select_one('span:contains("Pickup")').parent.text.strip() if sp.select_one('span:contains("Pickup")') else 'N/A', 'Obj_type': sp.select_one('div[class*="ProductCategoryBrand"]').text.split('•')[0].strip() if sp.select_one('div[class*="ProductCategoryBrand"]') else 'N/A', } return product:contains()selector works with BeautifulSoup'shtml.parser; if you stick withlxml, you can filter elements manually with a loop.
3. Dynamic Content Loading (JS-Rendered Pages)
Weedmaps' search results are loaded dynamically via JavaScript. The requests library only fetches the initial static HTML, which doesn't include product cards—so get_page_link can't find any links when run in main, even if your selector is correct. This explains why individual functions worked (you likely tested them on fully loaded product pages, not the search page).
Fix: Use a browser automation tool like Selenium to render the page fully before scraping:
Here's a modified get_page_link using Selenium:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def get_page_link(url): # Ensure ChromeDriver is in your system PATH driver = webdriver.Chrome() driver.get(url) # Wait 10 seconds for product cards to load WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'a[href^="/products/"]')) ) sp = BeautifulSoup(driver.page_source, 'lxml') driver.quit() baseurl = "https://weedmaps.com" links = sp.select('a[href^="/products/"]') unique_links = list({link.attrs['href'] for link in links}) return [baseurl + link for link in unique_links]
Final Debug Checks
- Add debug prints in
mainto verify how many links are being found:def main(): distance = input("Miles: ") results = [] for x in range(1,10): urlz = f'https://weedmaps.com/search?entryType=home%20page%20product%20card&filter%5BboundingRadius%5D={distance}mi&page={x}' urls = get_page_link(urlz) print(f"Page {x}: Found {len(urls)} product links") # Debug line productinfo = [product_data(url) for url in urls] results.append(productinfo) time.sleep(2) # Add delay to avoid rate limiting return results - Weedmaps has rate limiting—adding a 2-second delay between requests will help you avoid getting blocked.
内容的提问来源于stack exchange,提问作者SomeWhiteGuy

