如何用Python处理URL用于网页爬取?批量爬取Finviz股票数据
Hey there! Let's tackle your two questions one by one—they tie into each other really well, so we'll build from the basics up to your specific stock scraping task.
Web scraping with Python usually boils down to three core steps: handling the URL, sending a request to fetch the page, and parsing the content you need. Here's a breakdown of how to work with URLs specifically:
- Pick an HTTP request library: The most popular choices are
requests(super user-friendly, needs installation) andurllib(built into Python, no extra setup). For most cases,requestsis the way to go for simplicity. - Build or modify URLs: You can use string formatting to plug values into URLs, or use
urllib.parseto handle query parameters cleanly (this will come in handy for your stock code task later). - Handle URL encoding: If your URL has special characters (like spaces or symbols), use
urllib.parse.quote()to encode them properly—this ensures the server understands your request.
Here's a quick example of fetching a page with requests:
import requests # Basic URL handling url = "https://example.com" response = requests.get(url) # Check if the request worked (status code 200 means success) if response.status_code == 200: page_content = response.text # Now you can parse this content with tools like BeautifulSoup else: print(f"Request failed with status code {response.status_code}")
Your task is perfect for combining URL parameter handling with web scraping. Let's walk through this with actionable code:
First, let's assume your tweets array looks like this:
tweets = ["aapl", "cdti", "ovas", "mbot"]
Step 1: Generate the correct URLs for each stock
You have two straightforward ways to replace the t parameter in the Finviz URL:
- F-strings: Simple and readable for this specific use case.
urllib.parse.urlencode: More robust if you ever need to add or modify multiple parameters later.
Here's how to implement both:
Option 1: F-strings (quick and easy)
base_url = "https://finviz.com/quote.ashx?t={}" stock_urls = [base_url.format(stock) for stock in tweets]
Option 2: Using urllib.parse (cleaner for parameter management)
from urllib.parse import urlencode, urlunparse base_parts = ("https", "finviz.com", "/quote.ashx", "", "", "") stock_urls = [] for stock in tweets: params = {"t": stock} # Encode parameters and build the full URL full_url = urlunparse(base_parts[:4] + (urlencode(params),) + base_parts[5:]) stock_urls.append(full_url)
Step 2: Fetch each page and extract the table
Finviz's quote pages have a key stock data table. We can use pandas to pull tables directly from HTML (super fast!) or BeautifulSoup for more granular control over parsing.
Using pandas (fastest way for tables):
import pandas as pd import requests # Add a user-agent to avoid being blocked—many sites flag default requests as bots headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"} stock_data = {} for stock in tweets: url = f"https://finviz.com/quote.ashx?t={stock}" # Fetch the page with our headers response = requests.get(url, headers=headers) if response.status_code == 200: # Read all tables from the page; Finviz's main data table is usually the first one tables = pd.read_html(response.text) # Store the table in a dictionary using the stock code as the key stock_data[stock] = tables[0] else: print(f"Failed to fetch data for {stock} (status code: {response.status_code})") # Access individual stock tables like this: print(stock_data["aapl"])
Using BeautifulSoup (more control over parsing):
from bs4 import BeautifulSoup import requests headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"} stock_data = {} for stock in tweets: url = f"https://finviz.com/quote.ashx?t={stock}" response = requests.get(url, headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.text, "html.parser") # Target Finviz's snapshot table using its class name table = soup.find("table", class_="snapshot-table2") rows = table.find_all("tr") # Extract key-value pairs from the table rows stock_info = {} for row in rows: cols = row.find_all("td") for i in range(0, len(cols), 2): key = cols[i].text.strip() value = cols[i+1].text.strip() stock_info[key] = value stock_data[stock] = stock_info else: print(f"Failed to fetch data for {stock} (status code: {response.status_code})") # Access individual stock data like this: print(stock_data["cdti"])
A quick heads-up: Always include a User-Agent header in your requests—many websites block requests that don't have one, as they look like automated bots. Also, make sure to review Finviz's terms of service before scraping to ensure you're allowed to fetch their data programmatically.
内容的提问来源于stack exchange,提问作者Taylor Rhodes

