Python新手求助:如何用BeautifulSoup爬取Reuters的Beta指标数值
如何获取路透社AAPL关键指标中的Beta值?
Hey there! I see you're just starting out with Python web scraping and ran into a snag trying to pull the Beta value from Reuters' key metrics page for AAPL. Let's break down what's going wrong with your original code and fix it step by step.
原代码的问题点
Your current approach has a couple of issues that are leading to the -1 result:
- You're using
find()on thethelement, which only returns the first matching element instead of all metric rows. - Trying to
find('Beta')directly on thekey_nameelement won't work—this method looks for a tag named<Beta>, not text content that includes "Beta". - You're looping over
key_name, which is a single element, not a list, so that loop isn't doing what you expect.
修改后的代码
Here's a revised version of your code that correctly targets the Beta value:
import requests from bs4 import BeautifulSoup url = 'https://www.reuters.com/companies/AAPL.OQ/key-metrics' page = requests.get(url) if page.status_code == requests.codes.ok: bs = BeautifulSoup(page.text, 'lxml') # 抓取所有包含指标数据的行 all_metric_rows = bs.find_all('tr', class_='data') # 遍历每一行寻找Beta条目 for row in all_metric_rows: # 从表头单元格提取指标名称 metric_name = row.find('th', class_='MarketsTable-label-_JI6s').get_text(strip=True) # 检查当前行是否对应Beta指标 if 'Beta' in metric_name: # 从对应数据单元格的span标签中提取数值 beta_value = row.find('td').find('span').get_text(strip=True) print(f"AAPL Beta Value: {beta_value}") break # 找到目标后退出循环,无需遍历剩余行 else: print(f"页面加载失败,状态码: {page.status_code}")
代码解释
find_all('tr', class_='data'): 抓取指标表格中的每一行,这样我们可以逐一检查每一行是否是Beta指标。get_text(strip=True): 从表头单元格提取干净的文本内容(去除多余空格),方便我们判断是否包含"Beta"关键词。- 定位数值: 找到目标行后,我们在
<td>数据单元格内找到存放数值的<span>标签,再提取其中的文本内容。
额外提示
要注意像路透社这类网站可能会随时调整HTML结构,如果之后代码失效,你需要用浏览器的开发者工具(右键>检查)重新查看页面结构,更新选择器。另外,部分网站有反爬机制,如果遇到访问被拦截的情况,可以在requests.get()中添加User-Agent请求头,模拟真实浏览器的访问。
内容的提问来源于stack exchange,提问作者Dlintz
相关产品推荐
相关产品推荐

