如何使用BeautifulSoup和Requests抓取单URL多标签页下TradingView的本周收益数据?
Hey there! Let's break down why your current code isn't working and fix it so you can grab that "This Week" earnings data.
The Root of the Problem
TradingView loads earnings data dynamically with JavaScript. When you use requests.get(), you only fetch the initial static HTML of the page—which includes the "Today" tab's content (or sometimes none of the table rows, since they load after the page renders). The "This Week" tab's data isn't present in that initial response; it only loads when you click the tab via JavaScript. That's why your find_all() call returns an empty list.
Solution 1: Use TradingView's Scanner API (Most Efficient)
Instead of parsing HTML, you can directly call the API that TradingView uses to load earnings data. This is faster and gives you structured JSON data. Here's how to do it:
- Find the API Endpoint: Open your browser's DevTools (F12), go to the Network > XHR tab, then click the "This Week" tab on the page. You'll see a POST request to an API like
https://scanner.tradingview.com/america/scan. - Replicate the Request: Use
requeststo send a POST request with the right filters for "This Week" earnings.
Example Code
import requests headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Referer': 'https://www.tradingview.com/markets/stocks-usa/earnings/' } # API endpoint for TradingView's scanner api_url = 'https://scanner.tradingview.com/america/scan' # Payload defining filters for "This Week" earnings payload = { "filter": [ {"left": "earningsDate", "operation": "this_week"} ], "options": { "lang": "en" }, "markets": ["america"], "symbols": { "query": {"types": ["stock"]}, "tickers": [] }, "columns": [ "name", "close", "earningsDate", "earningsActual", "earningsEstimate", "earningsSurprisePct", "volume" ] } # Send the request and parse the JSON response response = requests.post(api_url, json=payload, headers=headers) data = response.json() # Extract and print the data for item in data['data']: ticker = item['s'] details = item['d'] print(f"Ticker: {ticker}") print(f"Company Name: {details[0]}") print(f"Earnings Date: {details[2]}") print(f"Actual EPS: {details[3]}") print(f"Estimated EPS: {details[4]}") print(f"Surprise %: {details[5]}") print("---")
Solution 2: Use Selenium to Simulate Browser Interaction (Beginner-Friendly)
If you prefer to mimic human interaction with the page, use Selenium to load the browser, click the "This Week" tab, and then parse the rendered HTML.
Steps
- Install Selenium:
pip install selenium - Download the appropriate driver for your browser (e.g., ChromeDriver for Chrome) and add it to your PATH.
Example Code
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # Initialize Chrome driver (specify driver path if not in PATH) driver = webdriver.Chrome() driver.get('https://www.tradingview.com/markets/stocks-usa/earnings/') try: # Wait for the "This Week" tab to be clickable, then click it this_week_tab = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//span[text()="This Week"]')) ) this_week_tab.click() # Wait for the table rows to load (adjust time if needed) time.sleep(3) # Alternatively use WebDriverWait for rows to appear # Get the fully rendered page source page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # Extract table rows rows = soup.find_all('tr', {'class':'tv-data-table__row tv-data-table__stroke tv-screener-table__result-row'}) for row in rows: cells = row.find_all('td') if cells: ticker = cells[0].get_text(strip=True) company = cells[1].get_text(strip=True) earnings_date = cells[2].get_text(strip=True) print(f"{ticker} | {company} | {earnings_date}") finally: # Close the browser driver.quit()
Why Your Original Code Failed
Your initial code uses requests, which can't execute JavaScript. The static HTML returned by requests.get() doesn't include the "This Week" tab's table rows—those are added dynamically after the page loads and the tab is clicked. Hence, soup.find_all() returns an empty list.
内容的提问来源于stack exchange,提问作者gibster

