You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Pandas无法抓取雅虎财经财务报表?求解决方案

解决Yahoo Finance港股1928.HK财务报表抓取问题

Hey there, let's break down why your pandas methods aren't pulling the right financial data from Yahoo Finance for 1928.HK, and fix this up properly.

为什么你的原有方法失效?

The core issue here is that Yahoo Finance loads its financial statements dynamically with JavaScript. Pandas' built-in functions only fetch the static initial HTML of the page—they can't wait for JavaScript to render the actual financial tables you want. Here's why each method fell short:

  • pd.read_html(): Looks for static HTML tables in the initial page source, but the financials don't exist there yet (they load after the page finishes rendering).
  • pd.read_json(): The financial data isn't exposed as a direct, publicly accessible JSON endpoint that this function can target.
  • pd.read_table(): It's just pulling random static text/tables from the initial HTML, which has nothing to do with the actual financial report.

解决方案1:用Selenium抓取动态渲染内容

Selenium controls a real browser, so it can wait for JavaScript to fully load the financial tables before you scrape them. This is the most reliable approach for dynamic sites like Yahoo Finance.

步骤:

  1. First, install Selenium and a browser driver (like ChromeDriver—make sure it matches your Chrome version):
    pip install selenium
    
  2. Use this code to grab the financials:
    from selenium import webdriver
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    import pandas as pd
    
    # Initialize Chrome driver (adjust path if your driver isn't in system PATH)
    driver = webdriver.Chrome()
    url = "https://finance.yahoo.com/quote/1928.HK/financials?p=1928.HK"
    driver.get(url)
    
    # Wait up to 10 seconds for the financial tables to load
    wait = WebDriverWait(driver, 10)
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "table")))
    
    # Extract all tables from the fully rendered page
    tables = pd.read_html(driver.page_source)
    # The first table is usually the Income Statement—adjust the index if needed
    financial_df = tables[0]
    
    # Save to Excel
    financial_df.to_excel("1928HK_financials.xlsx", index=False)
    
    # Clean up: close the browser
    driver.quit()
    
    # Preview the data
    print(financial_df.head())
    

解决方案2:用requests_html(轻量无浏览器)

If you don't want to use a full browser, requests_html can render JavaScript in a headless environment—it's lighter than Selenium and works well for this use case.

步骤:

  1. Install the library first:
    pip install requests-html
    
  2. Run this code:
    from requests_html import HTMLSession
    import pandas as pd
    
    # Create a session to handle the request
    session = HTMLSession()
    url = "https://finance.yahoo.com/quote/1928.HK/financials?p=1928.HK"
    
    # Fetch the page and render JavaScript (wait 2 seconds for loading)
    r = session.get(url)
    r.html.render(sleep=2)
    
    # Extract all tables from the rendered page
    tables = pd.read_html(r.html.html)
    financial_df = tables[0]
    
    # Save to Excel
    financial_df.to_excel("1928HK_financials.xlsx", index=False)
    
    # Preview the data
    print(financial_df.head())
    

小提示

  • Yahoo Finance has rate limits, so don't scrape too frequently (add delays if you're scraping multiple tickers).
  • The table index (like tables[0]) might change if Yahoo updates their site layout—check len(tables) to see how many tables are loaded, and print each one to find the right financial statement.
  • Keep your libraries updated to avoid compatibility issues.

内容的提问来源于stack exchange,提问作者Arthur Law

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:39:13