You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Owlerm网站内容失败,遇反爬拦截求助

How to Bypass Distil Anti-Scraping on Owler

Got it, let's break down why your current approach isn't working and how to fix it. The HTML response you're getting clearly shows Owler's Distil Networks anti-scraping system has flagged your request—this happens because plain requests calls don't execute JavaScript, and Distil actively checks for browser-like behavior that your script isn't mimicking.

Why Your Current Code Fails

Distil doesn't just look at the User-Agent header. It runs client-side JavaScript to verify things like:

  • Whether the request is coming from a real browser with a full JS engine
  • Session cookies set during initial page loads
  • Mouse movements, scroll behavior, and other human-like interactions
  • Fingerprinting details (like screen resolution, browser plugins, etc.)

Your requests script skips all of this, so Distil immediately redirects you to a captcha page.

Solutions to Try

1. Use a Headless Browser (Playwright/Selenium)

Headless browsers simulate real browsers, execute JS, and handle cookies/sessions automatically. Playwright is a modern, easier-to-use option compared to Selenium.

First, install Playwright and the required browser:

pip install playwright
playwright install chromium

Then use this script to scrape the page:

from playwright.sync_api import sync_playwright
import time
import random
from bs4 import BeautifulSoup

def scrape_owler(url):
    with sync_playwright() as p:
        # Launch Chromium (set headless=False to see the browser for debugging)
        browser = p.chromium.launch(
            headless=True,
            args=["--no-sandbox", "--disable-setuid-sandbox"]  # Fixes permission issues on some systems
        )
        # Randomize user agent to avoid detection
        user_agents = [
            "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36",
            "Mozilla/5.0 (Macintosh; Intel Mac OS X 14_1) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Safari/605.1.15",
            "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/119.0"
        ]
        page = browser.new_page(user_agent=random.choice(user_agents))
        
        # Navigate to the page and wait for network activity to settle
        page.goto(url, wait_until="networkidle")
        
        # Mimic human browsing with random delays
        time.sleep(random.uniform(2, 4))
        
        # Optional: Simulate scrolling to trigger lazy-loaded content
        page.mouse.wheel(0, random.randint(1000, 3000))
        time.sleep(random.uniform(1, 2))
        
        # Get the full page HTML
        html_content = page.content()
        
        browser.close()
        return html_content

# Scrape and parse the page
seed_url = "https://www.owler.com/location/new-york-companies?p=2"
page_html = scrape_owler(seed_url)
soup = BeautifulSoup(page_html, "lxml")

# Example: Extract company names (adjust selector based on actual page structure)
for company in soup.select(".company-listing__name"):
    print(company.get_text(strip=True))

2. Additional Anti-Detection Tweaks

If you still get blocked, try these:

  • Rotate Proxies: Use residential proxies (avoid datacenter proxies—they’re easier to flag) to switch IPs if your current one gets blacklisted.
  • Avoid Rapid Requests: Add longer, random pauses between repeated requests to mimic human browsing speed.
  • Disable Browser Fingerprinting: Use Playwright's context.add_init_script() to override fingerprinting-related JS properties (e.g., screen size, canvas rendering).

3. Legal and Ethical Notes

Before scraping, make sure you:

  • Review Owler's Terms of Service to confirm scraping is allowed.
  • Respect https://www.owler.com/robots.txt for restricted paths.
  • Don’t overload their servers—keep request rates low.

内容的提问来源于stack exchange,提问作者joel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:29:25