使用BeautifulSoup爬取富达官网嵌入跳转链接的美股板块数据求助
问题排查与解决方案
核心问题原因
- 你发送的请求没有携带浏览器标识头,目标站点识别为爬虫后直接拦截了请求,首页返回内容为空或非目标页面,导致你提取的跳转链接列表
links_list是空的,后续遍历逻辑没有执行,自然没有任何输出。 - 你使用的元素选择器
a.heading1和目标站点实际页面结构不匹配,也会导致无法正确提取链接。
修正后可运行代码
import requests import time from bs4 import BeautifulSoup # 增加请求头,模拟浏览器访问 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } url = "https://eresearch.fidelity.com/eresearch/goto/markets_sectors/landing.jhtml" req = requests.get(url, headers=headers) soup = BeautifulSoup(req.content, "html.parser") links_list = [] # 调整选择器匹配板块跳转链接,若页面更新可自行审查元素调整规则 next_page_links = soup.select("div.sector-name a") for link in next_page_links: next_page = "https://eresearch.fidelity.com" + link.get("href") links_list.append(next_page) result = [] for item in links_list: # 增加请求间隔,避免触发反爬限制 time.sleep(1) resp = requests.get(item, headers=headers) soup2 = BeautifulSoup(resp.content, 'html.parser') # 以下字段提取规则可根据实际页面结构调整 sector_name = soup2.select_one("h1.sector-title").get_text(strip=True) change_rate = soup2.select_one("div.performance-change").get_text(strip=True) market_cap = soup2.select_one("div.market-cap-value").get_text(strip=True) weight = soup2.select_one("div.market-weight-value").get_text(strip=True) sector_item = (sector_name, change_rate, market_cap, weight, item) result.append(sector_item) print(sector_item) # 最终目标元组列表存储在result变量中
调整说明
如果提取字段的规则不匹配,你可以打开目标页面按F12打开开发者工具,审查对应元素的类名/id,调整代码中的选择器即可。
内容的提问来源于stack exchange,提问作者Guacaka283
相关产品推荐
相关产品推荐

