You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python抓取msamb网站HTML表格数据?解决提取与输出问题

解决商品表格数据抓取与格式化输出问题

问题诊断

你目前的代码核心问题是错误地将下拉选项的文本(商品名称)当作HTML传给BeautifulSoup解析,而不是选择商品后获取页面上的表格HTML——这就导致你根本拿不到目标表格的数据。下面是修正后的完整解决方案:

修正后的完整代码

from selenium import webdriver
from selenium.webdriver.support.ui import Select
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup
import pandas as pd
from datetime import datetime

# 初始化浏览器驱动(注意:Chrome版本需与chromedriver匹配)
driver = webdriver.Chrome(executable_path='G:/data/depend/chromedriver.exe')
driver.get('https://www.msamb.com/ApmcDetail/ArrivalPriceInfo/')

# 获取页面显示的当日日期(如需历史数据,可扩展添加日期选择逻辑)
current_date = datetime.now().strftime('%Y-%m-%d')

# 初始化存储所有数据的列表
all_data = []

# 获取商品下拉选择框
commodity_select = Select(driver.find_element(By.ID, "CommoditiesId"))

# 遍历所有商品选项(跳过第一个默认占位选项)
for index, option in enumerate(commodity_select.options):
    if index == 0:
        continue
    
    commodity_name = option.text.strip()
    commodity_value = option.get_attribute('value')
    
    # 选择当前商品
    commodity_select.select_by_value(commodity_value)
    
    # 等待表格加载完成(避免页面未加载完就解析)
    try:
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, "tblData"))
        )
    except Exception as e:
        print(f"加载{commodity_name}表格时超时,跳过该商品")
        continue
    
    # 获取当前页面HTML,解析目标表格
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html.parser')
    table = soup.find('table', id='tblData')
    
    if not table:
        print(f"{commodity_name}无对应数据,跳过")
        continue
    
    # 遍历表格行(跳过表头)
    rows = table.find_all('tr')[1:]
    for row in rows:
        tds = row.find_all('td')
        # 确保列数足够,避免索引错误
        if len(tds) >= 7:
            apmc = tds[0].text.strip()
            variety = tds[1].text.strip()
            unit = tds[2].text.strip()
            quantity = tds[3].text.strip()
            lrate = tds[4].text.strip()
            hrate = tds[5].text.strip()
            modal = tds[6].text.strip()
            
            # 组装符合要求的行数据
            data_row = {
                'Date': current_date,
                'Commodity': commodity_name,
                'APMC': apmc,
                'Variety': variety,
                'Unit': unit,
                'Quantity': quantity,
                'Lrate': lrate,
                'Hrate': hrate,
                'Modal': modal
            }
            all_data.append(data_row)

# 转换为DataFrame并生成~|~分隔的输出
df = pd.DataFrame(all_data)
output_str = df.to_csv(sep='~|~', index=False)
print(output_str)

# 可选:保存到本地文件
df.to_csv('commodity_prices.txt', sep='~|~', index=False, encoding='utf-8')

# 关闭浏览器
driver.quit()

关键改进点说明

  • 等待表格加载:用WebDriverWait等待表格元素出现,解决页面异步加载导致的解析失败问题。
  • 正确解析页面HTML:每次选择商品后,获取当前页面的完整HTML再解析表格,而不是错误地解析选项文本。
  • 异常处理:添加超时和空表格判断,避免单个商品加载失败导致整个程序崩溃。
  • 数据关联:自动绑定日期和商品名称到每一行数据,完全匹配你需要的输出格式。
  • 便捷格式化:利用pandas的to_csv直接生成~|~分隔的格式,既可以打印也能保存到文件。

注意事项

  • 确保Chrome浏览器版本与chromedriver.exe版本一致,否则会启动失败。
  • 若遇到反爬限制,可添加time.sleep(2)之类的等待,或配置浏览器的用户代理。
  • 部分商品可能没有数据,代码会自动跳过,避免报错。

内容的提问来源于stack exchange,提问作者Vijay_Shinde

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 14:52:27