You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python无报错但Weedmaps爬虫返回空结果问题排查求助

Troubleshooting Your Weedmaps Crawler (Empty Results Issue)

Let's break down why your crawler is returning [[], [], [], [], [], [], [], [], []] even though your get_page_link and product_data functions work in isolation. Here are the key issues and actionable fixes:

1. URL Protocol & Redirect Mismatch

In your main function, you're using http://www.weedmaps.com for search requests. Weedmaps enforces HTTPS, so this request gets redirected to the HTTPS version—but the redirected page often has a different structure than what your selector expects, or the redirect process strips out the content you're trying to scrape.

Fix: Update the base URL in main to use HTTPS and drop the www subdomain (Weedmaps redirects from www to the root domain automatically anyway):

urlz = f'https://weedmaps.com/search?entryType=home%20page%20product%20card&filter%5BboundingRadius%5D={distance}mi&page={x}'

2. Fragile Dynamic Class Names

Your CSS selectors rely on auto-generated hashed class names like styles__NameRatingWrap-j5iyiv-15.eaLQmf and styled-components__Price-sc-1fbw3xt-15.lbyswm. Weedmaps uses tools like styled-components that generate these class names, which change frequently (even between page loads or site deployments). When you tested functions individually, the class names were valid—but by the time you ran main, they likely changed, or the search page uses different dynamic classes than product pages.

Fix: Replace dynamic class names with stable, structure-based selectors:

  • For get_page_link, target links that explicitly point to product pages (they start with /products/):
    def get_page_link(url):
        r = requests.get(url, headers=agent, timeout=3)
        sp = BeautifulSoup(r.text, 'lxml')
        baseurl = "https://weedmaps.com"
        # Target all links pointing to product pages, deduplicate to avoid repeats
        links = sp.select('a[href^="/products/"]')
        unique_links = list({link.attrs['href'] for link in links})
        return [baseurl + link for link in unique_links]
    
  • For product_data, use flexible selectors and add safety checks to avoid AttributeError if elements are missing:
    def product_data(url):
        r = requests.get(url, headers=agent, timeout=3)
        sp = BeautifulSoup(r.text, 'lxml')
        product = {
            'Title': sp.select_one('h1').text.strip() if sp.select_one('h1') else 'N/A',
            'Brand': sp.select_one('div[class*="ProductCategoryBrand"] a').text.strip() if sp.select_one('div[class*="ProductCategoryBrand"] a') else 'N/A',
            'Price': sp.select_one('div[class*="Price"]').text.strip() if sp.select_one('div[class*="Price"]') else 'N/A',
            'Pick_up_location': sp.select_one('span:contains("Pickup")').parent.text.strip() if sp.select_one('span:contains("Pickup")') else 'N/A',
            'Obj_type': sp.select_one('div[class*="ProductCategoryBrand"]').text.split('•')[0].strip() if sp.select_one('div[class*="ProductCategoryBrand"]') else 'N/A',
        }
        return product
    
    Note: The :contains() selector works with BeautifulSoup's html.parser; if you stick with lxml, you can filter elements manually with a loop.

3. Dynamic Content Loading (JS-Rendered Pages)

Weedmaps' search results are loaded dynamically via JavaScript. The requests library only fetches the initial static HTML, which doesn't include product cards—so get_page_link can't find any links when run in main, even if your selector is correct. This explains why individual functions worked (you likely tested them on fully loaded product pages, not the search page).

Fix: Use a browser automation tool like Selenium to render the page fully before scraping:
Here's a modified get_page_link using Selenium:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def get_page_link(url):
    # Ensure ChromeDriver is in your system PATH
    driver = webdriver.Chrome()
    driver.get(url)
    # Wait 10 seconds for product cards to load
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, 'a[href^="/products/"]'))
    )
    sp = BeautifulSoup(driver.page_source, 'lxml')
    driver.quit()
    baseurl = "https://weedmaps.com"
    links = sp.select('a[href^="/products/"]')
    unique_links = list({link.attrs['href'] for link in links})
    return [baseurl + link for link in unique_links]

Final Debug Checks

  • Add debug prints in main to verify how many links are being found:
    def main():
        distance = input("Miles: ")
        results = []
        for x in range(1,10):
            urlz = f'https://weedmaps.com/search?entryType=home%20page%20product%20card&filter%5BboundingRadius%5D={distance}mi&page={x}'
            urls = get_page_link(urlz)
            print(f"Page {x}: Found {len(urls)} product links") # Debug line
            productinfo = [product_data(url) for url in urls]
            results.append(productinfo)
            time.sleep(2) # Add delay to avoid rate limiting
        return results
    
  • Weedmaps has rate limiting—adding a 2-second delay between requests will help you avoid getting blocked.

内容的提问来源于stack exchange,提问作者SomeWhiteGuy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 16:52:53