You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中获取产品页XHR请求的URL?

Great question—let’s break down how to solve this since you already have a solid foundation with Selenium and BeautifulSoup. The core challenge here is linking each product page to its specific accessories XHR endpoint, but we can tackle this with a few targeted strategies:

Approach 1: Capture XHR Requests via Selenium's Network Logging

Since you’re already comfortable with Selenium, you can configure it to log network traffic and filter out the exact XHR request that loads the accessories data. This works because when you click the Accessories tab, the browser fires that XHR call—we just need to catch it.

Steps:

  1. Enable performance logging in your ChromeOptions to capture network requests.
  2. Load the product page and navigate to the Accessories tab (if it doesn’t load by default).
  3. Filter the logs to find the XHR request that returns the accessories list (look for the URL pattern you already identified, e.g., containing accessories or matching the structure you’ve seen).
  4. Extract the full XHR URL from the log entry.

Code Example:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
import json

# Configure Chrome to log network traffic
chrome_options = Options()
chrome_options.set_capability("goog:loggingPrefs", {"performance": "ALL"})

driver = webdriver.Chrome(options=chrome_options)
product_url = "https://www.phoenixcontact.com/online/portal/us/?uri=pxc-oc-itemdetail:pid=3248125&library=usen&pcck=P-15-11-08-02-05&tab=1&selectedCategory=ALL"

driver.get(product_url)
# Click the Accessories tab (adjust the selector to match the actual tab element on the site)
try:
    driver.find_element(By.CSS_SELECTOR, "[data-tab='accessories']").click()
except Exception as e:
    print(f"Error clicking accessories tab: {e}")

# Extract network logs
logs = driver.get_log("performance")
accessories_xhr_url = None

for log in logs:
    log_message = json.loads(log["message"])["message"]
    # Filter for XHR requests and match your target URL pattern
    if log_message["method"] == "Network.requestWillBeSent" and "accessories" in log_message["params"]["request"]["url"]:
        accessories_xhr_url = log_message["params"]["request"]["url"]
        break

if accessories_xhr_url:
    print(f"Found accessories XHR URL: {accessories_xhr_url}")
    # Now pass this URL to BeautifulSoup to scrape the data
else:
    print("Could not find the accessories XHR request")

driver.quit()
Approach 2: Extract API Parameters from Embedded JavaScript

Many modern sites store API configuration (like endpoint URLs and parameters) in embedded JavaScript variables (e.g., window.__INITIAL_STATE__ or similar). You can scrape the page source with BeautifulSoup, extract this JSON data, and pull out the accessories API endpoint directly.

Steps:

  1. Fetch the product page source with BeautifulSoup (or Selenium, if JS rendering is needed).
  2. Search for script tags that contain global state variables.
  3. Use regex to extract the JSON data from the script content.
  4. Parse the JSON to find the accessories endpoint and construct the full URL.

Code Example:

import requests
from bs4 import BeautifulSoup
import re
import json

product_url = "https://www.phoenixcontact.com/online/portal/us/?uri=pxc-oc-itemdetail:pid=3248125&library=usen&pcck=P-15-11-08-02-05&tab=1&selectedCategory=ALL"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}

response = requests.get(product_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# Look for script tags containing state data (adjust the regex pattern to match the site's variable name)
script_tags = soup.find_all("script")
accessories_api_url = None

for script in script_tags:
    if script.string and "window.__productContext" in script.string:  # Example variable name, adjust as needed
        # Extract the JSON part using regex
        match = re.search(r"window.__productContext\s*=\s*({.*?});", script.string, re.DOTALL)
        if match:
            product_data = json.loads(match.group(1))
            # Navigate the JSON structure to find the accessories endpoint (inspect the actual data to find the right path)
            accessories_api_url = product_data["accessoriesEndpoint"] + "?site=usen&itemsPerPage=100"
            break

if accessories_api_url:
    print(f"Constructed accessories API URL: {accessories_api_url}")
else:
    print("Could not find embedded product data")
Approach 3: Construct the XHR URL Directly from Product Page Parameters

This is the most efficient method if the XHR URL uses parameters that are already present in the product page URL. Looking at your example product page:

  • The product page URL contains pid=3248125 and pcck=P-15-11-08-02-05
  • Your discovered XHR URL likely uses these same parameters to fetch accessories for that product

Steps:

  1. Extract pid and pcck parameters from each product page URL (use urllib.parse to parse the query string).
  2. Construct the XHR URL using the base endpoint you’ve already identified, replacing the pid and pcck values, and setting site=usen and itemsPerPage=100 (or your desired count).

Code Example:

from urllib.parse import urlparse, parse_qs

product_url = "https://www.phoenixcontact.com/online/portal/us/?uri=pxc-oc-itemdetail:pid=3248125&library=usen&pcck=P-15-11-08-02-05&tab=1&selectedCategory=ALL"

# Parse the product URL to extract parameters
parsed_url = urlparse(product_url)
query_params = parse_qs(parsed_url.query)
# Extract pid from the uri parameter
pid = query_params["uri"][0].split("pid=")[1]
# Extract pcck directly
pcck = query_params["pcck"][0]

# Construct the XHR URL using the base endpoint you found
base_xhr_url = "https://www.phoenixcontact.com/online/portal/us/api/accessories/list"  # Replace with your actual base endpoint
accessories_xhr_url = f"{base_xhr_url}?pid={pid}&pcck={pcck}&site=usen&itemsPerPage=100"

print(f"Constructed XHR URL: {accessories_xhr_url}")

Pro Tip for Avoiding 403 Errors

Since you ran into 403s earlier:

  • Always set a realistic User-Agent header in your requests (whether using Selenium or requests).
  • For Selenium, use a stealth plugin (like selenium-stealth) to mimic a real browser and avoid bot detection.
  • Add random delays between requests to avoid overwhelming the server.

内容的提问来源于stack exchange,提问作者Skyler Cornaby

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 16:17:28