You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取SGX债券数据时无法完整获取表格内容

问题

爬取新加坡交易所(SGX)零售固定收益证券页面的债券信息时,使用Selenium仅能获取表格前13行数据,从第14行开始所有字段均为None。观察发现网页HTML标签、类名结构一致,但超过第13行后,find_elements方法返回空值。

现有代码
a = driver.find_elements(By.TAG_NAME,'sgx-table-row')

combined=[]

for num in range(len(a)):
    combined.append([])
counter=0
for item in a:

    ticker = item.find_elements(By.TAG_NAME,'a')
    name = item.find_elements(By.TAG_NAME,'sgx-table-cell-text')
    price1 = item.find_elements(By.TAG_NAME,'sgx-table-cell-number')
    
    for item in ticker:
        if len(item.text) != 0:
            combined[counter].append(item.text)
        else:
            pass
    for item in name:
        if len(item.text) !=0:
           
            combined[counter].append(item.text)
        else:
            pass
    for item in price1:
        if len(item.text) != 0:
            
            combined[counter].append(item.text)
        else:
            pass
    counter+=1


df = pd.DataFrame(combined)
print(df)
输出结果
N518100E 230201  CMHS   99.000      99  0.827   98.173     ﹣     ﹣     0   
1   N519100A 240201  LSHS   97.000      97  0.945   96.055     ﹣     ﹣     0   
2   N520100A 251101  QGES        ﹣       ﹣  0.111        ﹣     ﹣     ﹣     0   
3   N521100V 261101  IRRS        ﹣       ﹣      0        ﹣     ﹣     ﹣     0   
4   NA12100N 420401  PH1S  110.000     110  0.842  109.158     ﹣     ﹣     0   
5   NA16100H 460301  BJGS  108.000     108  1.069  106.931     ﹣     ﹣     0   
6   NA20100F 500301  ZL8S  108.000     108  0.729  107.271     ﹣     ﹣     0   
7   NA21200W 511001  ZFGS   87.000      87      0       87     ﹣     ﹣     0   
8   NX13100H 230701  R1MS  101.500   101.5  0.157  101.343     ﹣     ﹣     0   
9   NX15100Z 250601  AFUS   99.701  99.701  0.331    99.37     ﹣     ﹣     0   
10  NX16100F 260601  BJHS  102.000     102  0.296  101.704     ﹣     ﹣     0   
11  NX18100A 280501  CMGS   90.000      90  0.585   89.415     ﹣     ﹣     0   
12  NX21100N 310701  RXYS        ﹣       ﹣  0.093        ﹣     ﹣     ﹣     0   
13  NY07100X 220901  7PMS  101.380  101.38  1.214  100.166     ﹣     ﹣     0   
14             None  None     None    None   None     None  None  None  None   
15             None  None     None    None   None     None  None  None  None   
16             None  None     None    None   None     None  None  None  None   
17             None  None     None    None   None     None  None  None  None   
18             None  None     None    None   None     None  None  None  None   
19             None  None     None    None   None     None  None  None  None   
20             None  None     None    None   None     None  None  None  None   
21             None  None     None    None   None     None  None  None  None   
22             None  None  
解决方案

问题根源

该页面采用滚动加载机制,未出现在浏览器视口内的表格行不会被完全渲染,导致Selenium无法获取这些行内元素的文本内容。

修复步骤

  1. 滚动到目标行:遍历每一行时,先将该行滚动到视口内,确保元素可见。
  2. 显式等待:等待行内元素加载完成后再获取数据,避免因元素未渲染导致的空值。
  3. 优化数据收集逻辑:处理元素为空的情况,用占位符替代None,保证数据结构一致。

修改后的代码

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

# 假设driver已初始化并打开目标页面
rows = driver.find_elements(By.TAG_NAME, 'sgx-table-row')
combined = []

for row in rows:
    # 滚动到当前行,确保元素在视口内
    driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", row)
    # 等待行内核心元素加载完成
    WebDriverWait(driver, 5).until(
        EC.presence_of_all_elements_located((By.TAG_NAME, 'sgx-table-cell-text'))
    )
    
    row_data = []
    # 收集ticker链接文本
    tickers = row.find_elements(By.TAG_NAME, 'a')
    for ticker in tickers:
        row_data.append(ticker.text.strip() if ticker.text else '-')
    # 收集文本类单元格数据
    text_cells = row.find_elements(By.TAG_NAME, 'sgx-table-cell-text')
    for cell in text_cells:
        row_data.append(cell.text.strip() if cell.text else '-')
    # 收集数值类单元格数据
    num_cells = row.find_elements(By.TAG_NAME, 'sgx-table-cell-number')
    for cell in num_cells:
        row_data.append(cell.text.strip() if cell.text else '-')
    
    combined.append(row_data)

df = pd.DataFrame(combined)
print(df)

额外优化建议

如果表格行数较多,可先一次性滚动到页面底部触发所有行加载,再批量获取元素,提升爬取效率:

# 滚动到页面底部,加载所有表格行
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
# 等待所有行加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_all_elements_located((By.TAG_NAME, 'sgx-table-row'))
)
# 后续执行遍历收集数据的逻辑

内容的提问来源于stack exchange,提问作者smexy123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 17:24:29