无法定位表格id属性,如何用BeautifulSoup抓取港交所指定表格?
Hey there! I get it—targeting elements when IDs don't play nice can be tricky, but let's break this down step by step to grab that total turnover, market cap, and other data from the HKEX page.
Step 1: Set Up Your Tools
First, make sure you have requests and beautifulsoup4 installed. If not, run this in your terminal:
pip install requests beautifulsoup4
Step 2: Fetch the Page Content
HKEX might block plain requests, so we'll add a basic user-agent header to mimic a real browser:
import requests from bs4 import BeautifulSoup # Define the target URL and browser-like headers url = "https://www.hkex.com.hk/Mutual-Market/Stock-Connect/Statistics/Hong-Kong-and-Mainland-Market-Highlights?sc_lang=en#select3=0&select2=2&select1=28" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # Fetch the page and handle potential errors response = requests.get(url, headers=headers) response.raise_for_status() # This will throw an error if the request fails
Step 3: Locate the Container Div and Table
Since the table is nested inside that specific div, we'll first target the div using a combo of its class and ID (more precise than relying on just one attribute), then grab the table inside it:
# Parse the HTML content soup = BeautifulSoup(response.text, 'html.parser') # Find the exact container div target_div = soup.find('div', {'class': 'table-container fixed-freeze-tb-parent', 'id': 'Tbl__0'}) # Grab the table inside the div (exit if div isn't found) if target_div: table = target_div.find('table') else: print("Couldn't locate the target div—double-check the class/ID values!") exit()
Step 4: Extract Total Turnover, Market Cap, and Other Metrics
Now we'll loop through the table rows to find the data points you need. Since you're using the English version of the page, we'll match labels like "Total Turnover" or "Total Market Capitalization":
# Initialize variables to store your desired data total_turnover = None total_market_cap = None # Loop through each row in the table for row in table.find_all('tr'): cells = row.find_all('td') # Make sure the row has at least 2 cells (label + value) if len(cells) >= 2: label = cells[0].get_text(strip=True) value = cells[1].get_text(strip=True) # Match the labels to extract values if label == "Total Turnover": total_turnover = value elif label == "Total Market Capitalization": total_market_cap = value # Print out the results print(f"Total Turnover: {total_turnover}") print(f"Total Market Capitalization: {total_market_cap}")
Quick Tips to Avoid Headaches
- Dynamic Content Check: If you get
Nonefor your values, the table might be loaded with JavaScript. In that case, you'll need a tool likeseleniumto render the page fully before scraping. Start with the above code though—many HKEX stats pages are statically rendered. - Exact Text Matching: Double-check the label text by copying it directly from the page (hidden spaces or special characters can break matches).
- Be Respectful: Don't spam requests to HKEX—add a small delay (like
time.sleep(2)) if scraping multiple pages to avoid getting blocked.
内容的提问来源于stack exchange,提问作者willbillion

