Python新手求助:使用BeautifulSoup无法抓取网站表格数据
我是Python新手,想要使用BeautifulSoup抓取指定网站的表格数据,但编写的代码无法正常运行,以下是我的代码,恳请各位提供技术帮助:
import requests from bs4 import BeautifulSoup url = 'https://www.tfex.co.th/en/products/equity/set50-index-options/S50U24C850/historical-trading' r = requests.get(url) soup = BeautifulSoup(r.text, 'html.parser') price_table = soup.find('div',class_ = 'historical-trading-table-SET50IO') print(price_table) for j in price_table.find_all('tbody'): rows = j.find_all('tr') for row in rows: date = row.find_all('td', class_ = 'table-b-table-default b-table-sticky-column td-class mxn-date fs-14px default-font-family text-deep-gray text-nowrap column-sticky ps-7')[1].text open_price = row.find_all('td', class_ = 'text-end mxn-value pe-4')[2].text high_price = row.find_all('td', class_ = 'text-end mxn-value pe-4')[3].text low_price = row.find_all('td', class_ = 'text-end mxn-value pe-4')[4].text close_price = row.find_all('td', class_ = 'text-end mxn-value pe-4')[5].text sp_close_price = row.find_all('td', class_ = 'text-end mxn-value pe-4')[6].text chg_price = row.find_all('td', class_ = 'text-end mxn-value pe-4')[7].text chg_percentage = row.find_all('td', class_ = 'text-end mxn-value pe-4')[8].text volume_contracts = row.find_all('td', class_ = 'text-end mxn-value pe-4')[9].text OI_contracts = row.find_all('td', class_ = 'text-end mxn-value pe-4')[10].text print(date, open_price, high_price, low_price, close_price, sp_close_price, chg_price, chg_percentage, colume_contracts, OI_contracts)
解决BeautifulSoup抓取TFEX历史交易表格的问题
原代码核心问题
- 动态内容未加载:目标网站的表格数据是通过JavaScript动态渲染的,
requests.get()只能获取静态HTML源码,无法拿到JS加载后的表格内容,导致price_table返回None,后续循环直接报错。 - 元素定位冗余脆弱:使用过多class组合定位元素,网站前端更新class名称后代码会直接失效;同时依赖固定索引取值,容易出现越界错误。
- 变量拼写错误:最后print语句中的
colume_contracts应为volume_contracts,会引发未定义变量错误。
修正方案(用Selenium处理动态加载)
由于目标网站依赖JS渲染数据,推荐用Selenium模拟浏览器加载页面,再提取表格内容。
安装依赖
pip install selenium beautifulsoup4
同时需下载对应浏览器的驱动(如ChromeDriver),确保其路径配置正确。
修正后的代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time url = 'https://www.tfex.co.th/en/products/equity/set50-index-options/S50U24C850/historical-trading' # 配置无头浏览器(可选,不想打开浏览器窗口则启用) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') # 初始化浏览器 driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待页面加载完成(根据网络情况调整时间) time.sleep(3) # 获取渲染后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 定位表格容器 price_table = soup.find('div', class_='historical-trading-table-SET50IO') if price_table: # 找到表格的tbody tbody = price_table.find('tbody') if tbody: rows = tbody.find_all('tr') for row in rows: # 提取当前行的所有td元素 tds = row.find_all('td') if len(tds) >= 11: date = tds[1].get_text(strip=True) open_price = tds[2].get_text(strip=True) high_price = tds[3].get_text(strip=True) low_price = tds[4].get_text(strip=True) close_price = tds[5].get_text(strip=True) sp_close_price = tds[6].get_text(strip=True) chg_price = tds[7].get_text(strip=True) chg_percentage = tds[8].get_text(strip=True) volume_contracts = tds[9].get_text(strip=True) oi_contracts = tds[10].get_text(strip=True) print(date, open_price, high_price, low_price, close_price, sp_close_price, chg_price, chg_percentage, volume_contracts, oi_contracts) else: print("未找到表格容器") # 关闭浏览器 driver.quit()
代码说明
- Selenium模拟浏览器:通过浏览器驱动加载页面,确保获取到JS渲染后的完整内容。
- 简化元素定位:直接提取行内所有td元素,通过索引取值,同时增加长度判断避免越界。
- 添加错误处理:检查表格和tbody是否存在,避免空值报错。
- 修复变量拼写错误:修正了
colume_contracts的拼写问题。
内容的提问来源于stack exchange,提问作者Alan Ma
相关产品推荐
相关产品推荐

