如何在Python中获取产品页XHR请求的URL?
Great question—let’s break down how to solve this since you already have a solid foundation with Selenium and BeautifulSoup. The core challenge here is linking each product page to its specific accessories XHR endpoint, but we can tackle this with a few targeted strategies:
Since you’re already comfortable with Selenium, you can configure it to log network traffic and filter out the exact XHR request that loads the accessories data. This works because when you click the Accessories tab, the browser fires that XHR call—we just need to catch it.
Steps:
- Enable performance logging in your ChromeOptions to capture network requests.
- Load the product page and navigate to the Accessories tab (if it doesn’t load by default).
- Filter the logs to find the XHR request that returns the accessories list (look for the URL pattern you already identified, e.g., containing
accessoriesor matching the structure you’ve seen). - Extract the full XHR URL from the log entry.
Code Example:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By import json # Configure Chrome to log network traffic chrome_options = Options() chrome_options.set_capability("goog:loggingPrefs", {"performance": "ALL"}) driver = webdriver.Chrome(options=chrome_options) product_url = "https://www.phoenixcontact.com/online/portal/us/?uri=pxc-oc-itemdetail:pid=3248125&library=usen&pcck=P-15-11-08-02-05&tab=1&selectedCategory=ALL" driver.get(product_url) # Click the Accessories tab (adjust the selector to match the actual tab element on the site) try: driver.find_element(By.CSS_SELECTOR, "[data-tab='accessories']").click() except Exception as e: print(f"Error clicking accessories tab: {e}") # Extract network logs logs = driver.get_log("performance") accessories_xhr_url = None for log in logs: log_message = json.loads(log["message"])["message"] # Filter for XHR requests and match your target URL pattern if log_message["method"] == "Network.requestWillBeSent" and "accessories" in log_message["params"]["request"]["url"]: accessories_xhr_url = log_message["params"]["request"]["url"] break if accessories_xhr_url: print(f"Found accessories XHR URL: {accessories_xhr_url}") # Now pass this URL to BeautifulSoup to scrape the data else: print("Could not find the accessories XHR request") driver.quit()
Many modern sites store API configuration (like endpoint URLs and parameters) in embedded JavaScript variables (e.g., window.__INITIAL_STATE__ or similar). You can scrape the page source with BeautifulSoup, extract this JSON data, and pull out the accessories API endpoint directly.
Steps:
- Fetch the product page source with BeautifulSoup (or Selenium, if JS rendering is needed).
- Search for script tags that contain global state variables.
- Use regex to extract the JSON data from the script content.
- Parse the JSON to find the accessories endpoint and construct the full URL.
Code Example:
import requests from bs4 import BeautifulSoup import re import json product_url = "https://www.phoenixcontact.com/online/portal/us/?uri=pxc-oc-itemdetail:pid=3248125&library=usen&pcck=P-15-11-08-02-05&tab=1&selectedCategory=ALL" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get(product_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # Look for script tags containing state data (adjust the regex pattern to match the site's variable name) script_tags = soup.find_all("script") accessories_api_url = None for script in script_tags: if script.string and "window.__productContext" in script.string: # Example variable name, adjust as needed # Extract the JSON part using regex match = re.search(r"window.__productContext\s*=\s*({.*?});", script.string, re.DOTALL) if match: product_data = json.loads(match.group(1)) # Navigate the JSON structure to find the accessories endpoint (inspect the actual data to find the right path) accessories_api_url = product_data["accessoriesEndpoint"] + "?site=usen&itemsPerPage=100" break if accessories_api_url: print(f"Constructed accessories API URL: {accessories_api_url}") else: print("Could not find embedded product data")
This is the most efficient method if the XHR URL uses parameters that are already present in the product page URL. Looking at your example product page:
- The product page URL contains
pid=3248125andpcck=P-15-11-08-02-05 - Your discovered XHR URL likely uses these same parameters to fetch accessories for that product
Steps:
- Extract
pidandpcckparameters from each product page URL (useurllib.parseto parse the query string). - Construct the XHR URL using the base endpoint you’ve already identified, replacing the
pidandpcckvalues, and settingsite=usenanditemsPerPage=100(or your desired count).
Code Example:
from urllib.parse import urlparse, parse_qs product_url = "https://www.phoenixcontact.com/online/portal/us/?uri=pxc-oc-itemdetail:pid=3248125&library=usen&pcck=P-15-11-08-02-05&tab=1&selectedCategory=ALL" # Parse the product URL to extract parameters parsed_url = urlparse(product_url) query_params = parse_qs(parsed_url.query) # Extract pid from the uri parameter pid = query_params["uri"][0].split("pid=")[1] # Extract pcck directly pcck = query_params["pcck"][0] # Construct the XHR URL using the base endpoint you found base_xhr_url = "https://www.phoenixcontact.com/online/portal/us/api/accessories/list" # Replace with your actual base endpoint accessories_xhr_url = f"{base_xhr_url}?pid={pid}&pcck={pcck}&site=usen&itemsPerPage=100" print(f"Constructed XHR URL: {accessories_xhr_url}")
Pro Tip for Avoiding 403 Errors
Since you ran into 403s earlier:
- Always set a realistic
User-Agentheader in your requests (whether using Selenium or requests). - For Selenium, use a stealth plugin (like
selenium-stealth) to mimic a real browser and avoid bot detection. - Add random delays between requests to avoid overwhelming the server.
内容的提问来源于stack exchange,提问作者Skyler Cornaby

