切换指数筛选后无法抓取网页表格数据的技术求助
The page uses JavaScript to load data dynamically when you change the dropdown selection. The requests library only fetches the initial static HTML of the page, which doesn’t include the updated table content that loads after interacting with the dropdown. That’s why your soup isn’t picking up the new rows.
Solution 1: Scrape the Underlying API (More Efficient)
Most dynamic sites like this rely on a hidden API to fetch data in the background when you interact with elements. Here’s how to find and use it:
Locate the API Endpoint:
- Open your browser’s Developer Tools (F12), switch to the Network tab.
- Select a different option from the "Select Indices or Sub Indices" dropdown.
- Look for an XHR/Fetch request that appears immediately after your selection (it might be named something like
get_index_data). - Note the request method (POST/GET) and parameters (like
index_id) it sends.
Code to Use the API Directly:
Assuming you found the endpoint ishttp://nepalstock.com/indices/get_index_dataand it accepts a POST request with anindex_idparameter:import requests import pandas as pd from bs4 import BeautifulSoup # Replace with the actual index ID (find this from your browser's dev tools) target_index_id = 2 # Example value, adjust to your desired index # Send request to the API endpoint response = requests.post( "http://nepalstock.com/indices/get_index_data", data={"index_id": target_index_id} ) # Parse the returned table HTML soup = BeautifulSoup(response.content, "html.parser") table_rows = [] # Extract data from each row (skip empty/header rows if needed) for row in soup.find_all("tr"): cells = row.find_all("td") if cells: row_data = [cell.text.strip() for cell in cells] table_rows.append(row_data) # Convert to DataFrame df = pd.DataFrame(table_rows) print(df)
Solution 2: Use Selenium to Simulate Browser Interaction
If you can’t locate the API endpoint, use Selenium to mimic user behavior and load the dynamic content:
Install Dependencies:
pip install seleniumYou’ll also need to download a browser driver (e.g., ChromeDriver) and add it to your system PATH.
Code Example:
from selenium import webdriver from selenium.webdriver.support.ui import Select from selenium.webdriver.common.by import By import pandas as pd import time # Initialize browser driver = webdriver.Chrome() # Use Firefox/Edge driver if preferred driver.get("http://nepalstock.com/indices") # Wait for page to fully load time.sleep(2) # Find dropdown and select your desired index dropdown = Select(driver.find_element(By.NAME, "index_id")) # Replace with actual dropdown name/ID dropdown.select_by_visible_text("NEPSE Index") # Or use select_by_value() with index ID # Wait for table to update time.sleep(2) # Parse updated page source soup = BeautifulSoup(driver.page_source, "html.parser") table_rows = [] for row in soup.find_all("tr"): cells = row.find_all("td") if cells: row_data = [cell.text.strip() for cell in cells] table_rows.append(row_data) df = pd.DataFrame(table_rows) print(df) # Clean up: close browser driver.quit()
Pro Tips:
- Avoid Blocking: The site may flag repeated requests. Add small delays between requests or include browser-like headers in your
requestscalls. - Element Selectors: Use your browser’s dev tools to double-check the correct IDs/names for the dropdown and table elements (they might change over time).
- Pagination: If the table has multiple pages, either find the API’s pagination parameters or use Selenium to click the "Next" button and scrape each page.
内容的提问来源于stack exchange,提问作者Dimanjan

