You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python2.7+Firefox无头爬虫遇Failed to decode response from marionette错误求助

Hey there, let’s work through this problem together. You mentioned you’re using a Python 2.7 script with headless Firefox on OS X 10.10 to scrape Amazon reviews (which require JavaScript to load), and you haven’t found a solution after searching here and on Google. Let’s break down the most likely issues and fixes for your setup:

1. Check Firefox & OS X 10.10 Compatibility

OS X 10.10 (Yosemite) is pretty outdated, and newer Firefox versions won’t run on it. The last Firefox release that supports Yosemite is Firefox 78 ESR. You can confirm your Firefox version by running this in your terminal:

firefox -v

If you’re on a newer version, roll back to 78 ESR—this is a critical first step, since mismatched browser/OS versions cause all kinds of headless mode failures.

2. Match Selenium & GeckoDriver Versions

Selenium 3.x is the latest version compatible with Python 2.7, and you need a GeckoDriver version that pairs with Firefox 78 ESR. Use GeckoDriver v0.27.0 (newer versions dropped support for older Firefox builds).

Here’s how to set up headless mode correctly in your script:

from selenium import webdriver
from selenium.webdriver.firefox.options import Options

# Configure headless options
firefox_options = Options()
firefox_options.add_argument("-headless")
# Add these flags to avoid common headless mode glitches
firefox_options.add_argument("-disable-gpu")
firefox_options.add_argument("-no-sandbox")
# Mimic a real browser's user-agent to avoid being blocked
firefox_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/49.0.2623.112 Safari/537.36")

# Initialize the driver (replace the path with your GeckoDriver location)
driver = webdriver.Firefox(firefox_options=firefox_options, executable_path="/usr/local/bin/geckodriver")
3. Beat Amazon’s Anti-Scraping Measures

Amazon actively blocks scrapers, even with headless browsers. Try these tweaks:

  • Wait for elements to load properly instead of using hard time.sleep() calls—this ensures you’re trying to scrape content only after it’s rendered:
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # Wait up to 10 seconds for the reviews section to appear
    reviews_container = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, "customer_review"))
    )
    
  • Add small delays between requests (2-5 seconds) to mimic human browsing.
  • If you’re getting blocked repeatedly, consider rotating user-agents or using a proxy service.
4. Debug Headless Mode Issues

If your script works in regular Firefox but not headless, use these tools to diagnose:

  • Take a screenshot to see what the headless browser is actually rendering:
    driver.save_screenshot("amazon_reviews.png")
    
  • Enable debug logging to catch hidden errors:
    import logging
    logging.basicConfig(level=logging.DEBUG)
    
5. Alternative: Try Headless Chrome (If Possible)

If Firefox continues to give you trouble, Chrome supports OS X 10.10 up to version 49. Pair it with ChromeDriver v2.21, and adjust your script to use Chrome’s headless mode—many users find this more stable on older OS versions.

内容的提问来源于stack exchange,提问作者Reily Bourne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:20:35