You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需Selenium:能否通过Chrome直接向Python脚本传输爬取数据?

Can I Use Chrome's Network XHR Data Directly in Python (Without Selenium)?

Absolutely! Ditching Selenium for this scenario is not just possible—it’s the far more efficient choice since you’ll cut down on resource usage and bandwidth by only fetching the raw data you need, instead of loading an entire browser instance.

Here’s how it works:

Those XHR/Fetch requests you see in Chrome’s DevTools are exactly what the browser sends to the server to get JSON/XML data. You can replicate these requests directly in Python using HTTP libraries, no browser middleman required.

Step 1: Capture the Request Details in Chrome

First, grab all the info you need from the Network panel:

  • Open Chrome DevTools (F12 or Ctrl+Shift+I), go to the Network tab.
  • Filter requests by XHR/Fetch to narrow down to the data you want.
  • Click on the target request, then check:
    • Request URL: The endpoint the browser is calling.
    • Request Method: GET, POST, etc.
    • Request Headers: Look for critical headers like User-Agent, Cookie, Authorization (these often prevent the server from blocking your request).
    • Payload (for POST requests): If the request sends data to the server, copy the form data or JSON body here.

Step 2: Replicate the Request in Python

Use a lightweight library like requests (install it first with pip install requests) to mimic the request. Here are examples for common cases:

Example 1: GET Request

import requests

# Replace with values from DevTools
url = "https://target-site.com/api/your-data-endpoint"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Cookie": "session_id=abc123; auth_token=xyz789"  # Copy from DevTools' Request Headers
}

# Send the request
response = requests.get(url, headers=headers)

# Parse the JSON response
data = response.json()

# Use the data however you need
print(data)

Example 2: POST Request with JSON Payload

import requests

url = "https://target-site.com/api/submit-data"
headers = {
    "User-Agent": "Your browser's User-Agent string",
    "Content-Type": "application/json"  # Often required for JSON payloads
}

# Copy the payload from DevTools' "Payload" tab
payload = {
    "filter": "recent",
    "limit": 20
}

response = requests.post(url, headers=headers, json=payload)
data = response.json()
print(data)

Key Notes to Avoid Blocking:

  • Don’t skip headers: Servers often block requests that don’t have a valid User-Agent or required auth headers (like cookies/tokens from your logged-in session).
  • Handle dynamic parameters: Some sites use dynamic values (like timestamps or signatures) in requests. If you encounter this, you’ll need to reverse-engineer how those values are generated (check the site’s JavaScript) and replicate that logic in Python.
  • Rate limiting: Be respectful of the server—don’t spam requests, or you might get IP-blocked.

To Answer Your Exact Questions:

  • Do I need Selenium or a browser tool? No, you can directly fetch the data using Python’s HTTP libraries.
  • How to transfer data from Chrome to Python? You don’t need to "transfer" it—instead, you use Chrome to inspect what the browser is doing, then replicate that behavior in Python to get the data directly from the server.

内容的提问来源于stack exchange,提问作者Raghavendra Phayde

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:09:05