使用Python Requests获取瑞士交易所股票行业信息失败求助
Since the SIX Swiss Exchange page renders most content via JavaScript, using requests alone won't work—it only grabs the initial static HTML, not the dynamically loaded data. Here are three practical approaches to extract the ICB Supersector under the "Shares Details" tab:
1. Use Selenium (Browser Automation)
Selenium launches a real browser, executes the page's JavaScript, and lets you interact with the fully rendered content. It's a tried-and-true tool for scraping JS-heavy pages.
First, install the required packages:
pip install selenium
You’ll also need to download the ChromeDriver (or driver for your preferred browser) and add it to your system PATH.
Here’s a working script tailored to your use case:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize Chrome driver driver = webdriver.Chrome() try: # Navigate to the target stock page url = "https://www.six-swiss-exchange.com/shares/security_info_en.html?id=CH0012221716CHF4" driver.get(url) # Wait for the "Shares Details" section to load (adjust timeout if needed) wait = WebDriverWait(driver, 10) icb_element = wait.until( EC.presence_of_element_located( (By.XPATH, "//div[contains(@class, 'shares-details')]//dt[text()='ICB Supersector']/following-sibling::dd") ) ) # Extract and print the ICB Supersector value icb_supersector = icb_element.text.strip() print(f"ICB Supersector: {icb_supersector}") finally: # Make sure to close the browser when done driver.quit()
2. Use Playwright (Modern Browser Automation)
Playwright is a more streamlined alternative to Selenium—it handles driver installation automatically and offers better performance for dynamic pages. It’s great if you want to avoid manual driver setup.
Install Playwright and its browsers first:
pip install playwright playwright install
Sample script:
from playwright.sync_api import sync_playwright with sync_playwright() as p: # Launch Chrome (swap with 'firefox' or 'webkit' if preferred) browser = p.chromium.launch(headless=False) # Set headless=True for background execution page = browser.new_page() # Navigate to the page and wait for network activity to settle page.goto("https://www.six-swiss-exchange.com/shares/security_info_en.html?id=CH0012221716CHF4", wait_until="networkidle") # Extract ICB Supersector using XPath icb_supersector = page.locator("//div[contains(@class, 'shares-details')]//dt[text()='ICB Supersector']/following-sibling::dd").text_content().strip() print(f"ICB Supersector: {icb_supersector}") browser.close()
3. Directly Call the Underlying API (Most Efficient)
Instead of automating a browser, you can inspect the page’s network traffic to find the API endpoint that fetches the security data. This is faster and more reliable, as it skips the overhead of rendering a full page.
Here’s how to find it:
- Open the target page in Chrome/Firefox, hit F12 to open DevTools, and go to the Network tab.
- Refresh the page, filter requests by "XHR". You’ll spot a request that returns JSON with all the security details.
Once you have the API URL, use requests to fetch the data directly:
import requests # Replace with the actual API endpoint you found in DevTools api_url = "https://api.six-group.com/api/findata/security/v1/CH0012221716CHF4" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) data = response.json() # Extract ICB Supersector (adjust the JSON path based on the actual response structure) icb_supersector = data["security"]["classification"]["icb"]["supersector"]["name"] print(f"ICB Supersector: {icb_supersector}")
内容的提问来源于stack exchange,提问作者Eric

