You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求协助:从美股涨幅页提取前5最高涨幅存入数组并打印

嘿,我来帮你搞定这个抓取WSJ涨幅前5股票数据的问题!你现在的代码会输出所有行列,核心问题是没有对涨幅数据进行排序和截取前5。下面是完整的解决方案,用Python的requests和BeautifulSoup实现:

完整实现代码
import requests
from bs4 import BeautifulSoup

# 目标页面URL
target_url = "http://www.wsj.com/mdc/public/page/2_3021-gainnyse-gainer.html"

try:
    # 模拟浏览器请求,避免被反爬拦截
    browser_headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    response = requests.get(target_url, headers=browser_headers)
    response.raise_for_status()  # 检查请求是否成功

    # 解析HTML页面
    soup = BeautifulSoup(response.text, "html.parser")

    # 定位涨幅表格(WSJ页面的表格有固定的class标识)
    gain_table = soup.find("table", class_="mdcTable")
    if not gain_table:
        print("抱歉,未找到目标涨幅表格")
        exit()

    # 提取表头和行数据
    table_headers = [th.get_text(strip=True) for th in gain_table.find_all("th")]
    stock_data = []

    # 遍历表格行(跳过表头行)
    for row in gain_table.find_all("tr")[1:]:
        columns = [td.get_text(strip=True) for td in row.find_all("td")]
        if not columns:
            continue  # 跳过空行

        # 处理涨幅数据:把百分比字符串转为浮点数,方便排序
        try:
            change_pct_idx = table_headers.index("% Chg")
            change_pct = float(columns[change_pct_idx].replace("%", ""))
        except (ValueError, IndexError):
            continue  # 跳过数据格式异常的行

        # 整理单条股票数据
        stock_info = {
            "股票名称": columns[table_headers.index("Name")],
            "股票代码": columns[table_headers.index("Symbol")],
            "涨幅(%)": change_pct,
            "最新价格": columns[table_headers.index("Last")],
            "成交量": columns[table_headers.index("Volume")]
        }
        stock_data.append(stock_info)

    # 按涨幅从高到低排序
    sorted_stocks = sorted(stock_data, key=lambda x: x["涨幅(%)"], reverse=True)

    # 截取前5条存入数组
    top5_gainers = sorted_stocks[:5]

    # 打印结果
    print("📈 涨幅最高的前5条股票数据:")
    for rank, stock in enumerate(top5_gainers, 1):
        print(f"{rank}. {stock['股票名称']} ({stock['股票代码']}) | 涨幅: {stock['涨幅(%)']}% | 最新价: {stock['最新价格']} | 成交量: {stock['成交量']}")

except requests.exceptions.RequestException as e:
    print(f"请求页面时出错:{str(e)}")
关键步骤说明
  • 模拟浏览器请求:添加User-Agent请求头,避免被WSJ的反爬机制直接拦截
  • 数据清洗与转换:把涨幅的百分比字符串转为浮点数,这样才能进行数值排序
  • 排序与筛选:用sorted()函数按涨幅降序排列,再通过切片[:5]获取前5条数据存入top5_gainers数组
  • 异常处理:处理请求失败、表格找不到、数据格式异常等情况,让代码更健壮

你可以直接运行这段代码,它会自动抓取页面数据、筛选出涨幅最高的前5条并打印出来~

内容的提问来源于stack exchange,提问作者Sam Maloney

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:39:09