You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Yahoo Finance公司治理评分报错求助:AttributeError问题

解决Yahoo Finance公司治理评分爬取的AttributeError问题

我是Python新手,尝试爬取Yahoo Finance中AMP的公司治理板块评分(页面链接:https://ca.finance.yahoo.com/quote/AMP/profile?p=AMP),运行代码时出现AttributeError错误。以下是我的爬取代码、报错信息和目标页面的HTML源码片段:

爬取代码

import requests
from bs4 import BeautifulSoup

# URL of the webpage to be scraped
url = "https://ca.finance.yahoo.com/quote/AMP/profile?p=AMP"

# Send a GET request to the URL and store the response in a variable
response = requests.get(url)

# Parse the HTML content of the response using BeautifulSoup
soup = BeautifulSoup(response.content, 'html.parser')

# Find the section containing the governance scores
governance_section = soup.find('section', {'class': 'Mt(30px) corporate-governance-container'})

# Find the required elements using their text content and extract their values
scores = governance_section.find('span', text='The pillar scores are')
score_text = scores[0].text.strip() if scores else ''
audit_score = score_text[score_text.find('Audit: ') + 6 : score_text.find(';', score_text.find('Audit: '))].strip()
shareholder_score = score_text[score_text.find('Shareholder Rights: ') + 20 : score_text.find(';', score_text.find('Shareholder Rights: '))].strip()
compensation_score = score_text[score_text.find('Compensation: ') + 14 :].strip()

# Print the extracted information
print("Corporate Governance Score:", governance_section.find('span', {'class': 'Va(m) Fw(600) D(ib) Lh(23px)'}).text.strip())
print("Audit and Risk Oversight Score:", audit_score)
print("Shareholders' Rights Score:", shareholder_score)
print("Compensation Score:", compensation_score)
print("Board Structure Score:", board_score)

报错信息

AttributeError                            Traceback (most recent call last)
<ipython-input-15-140e65f0615d> in <cell line: 17>()
     15 
     16 # Find the required elements using their text content and extract their values
---&gt; 17 scores = governance_section.find('span', text='The pillar scores are')
     18 score_text = scores[0].text.strip() if scores else ''
     19 audit_score = score_text[score_text.find('Audit: ') + 6 : score_text.find(';', score_text.find('Audit: '))].strip()

AttributeError: 'NoneType' object has no attribute 'find'

目标页面HTML源码片段

Minneapolis, Minnesota.</p></section><section class="Mt(30px) corporate-governance-container"><h2 class="Fz(m) Lh(1) Fw(b) Mt(0) Mb(18px)"><span>Corporate Governance</span></h2><div><p class="Fz(s)"><span>Ameriprise Financial, Inc.’s ISS Governance QualityScore as of <span>April 1, 2023</span> is 5.</span>  <span>The pillar scores are Audit: 3; Board: 7; Shareholder Rights: 5; Compensation: 6.</span></p><div class="Mt(20px)"><span>Corporate governance scores courtesy of</span> <a href="https://issgovernance.com/quickscore" target="_blank" rel="noopener noreferrer" title="Institutional Shareholder Services (ISS)">Institutional Shareholder Services (ISS)</a>.  <span>Scores indicate decile rank relative to index or region. A decile score of 1 indicates lower governance risk, while a 10 indicates higher governance risk.</span></div></div></section></section></div></div><script>if (window.performance) {window.performance.mark && window.performance.mark('Col1-0-Profile');window.performance.measure && window.performance.measure('Col1-0-ProfileDone','PageStart','Col1-0-Profile');}</script></div><div>

问题原因分析

  1. 反爬拦截导致页面获取失败:Yahoo Finance会检测请求来源,直接用requests.get()发送请求未携带浏览器标识,服务器返回非目标页面内容,导致soup.find()找不到指定section,governance_section变为None,后续调用find()触发AttributeError。
  2. 代码逻辑缺陷:
    • scores = governance_section.find(...)返回单个元素而非列表,使用scores[0]会报错;
    • 未定义board_score变量,最后打印时会触发NameError;
    • 用字符串索引提取分数的方式不够健壮,页面文本格式变化就会失效。

修复方案

1. 添加请求头模拟浏览器

给requests.get()添加User-Agent头,让服务器识别为浏览器请求。

2. 增加空值检查

使用governance_section前先判断是否存在,避免None调用方法。

3. 修正分数提取逻辑

用正则表达式提取分数,比字符串索引更健壮,能适应小范围文本格式变化。

4. 补充缺失变量

定义board_score变量,避免打印时报错。

修复后的完整代码

import requests
from bs4 import BeautifulSoup
import re

# URL of the webpage to be scraped
url = "https://ca.finance.yahoo.com/quote/AMP/profile?p=AMP"

# 添加请求头,模拟浏览器请求
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

# Send a GET request to the URL with headers
response = requests.get(url, headers=headers)
response.raise_for_status()  # 检查请求是否成功

# Parse the HTML content of the response using BeautifulSoup
soup = BeautifulSoup(response.content, 'html.parser')

# Find the section containing the governance scores
governance_section = soup.find('section', {'class': 'Mt(30px) corporate-governance-container'})

if governance_section:
    # 提取总治理评分
    total_score_elem = governance_section.find('span', text=re.compile(r'ISS Governance QualityScore as of'))
    total_score = re.search(r'is (\d+)', total_score_elem.text.strip()).group(1) if total_score_elem else 'N/A'
    
    # 提取分支柱分数
    score_elem = governance_section.find('span', text=re.compile(r'The pillar scores are'))
    score_text = score_elem.text.strip() if score_elem else ''
    
    # 用正则提取各个分数
    score_matches = re.findall(r'(Audit|Board|Shareholder Rights|Compensation): (\d+)', score_text)
    score_dict = dict(score_matches)
    
    audit_score = score_dict.get('Audit', 'N/A')
    board_score = score_dict.get('Board', 'N/A')
    shareholder_score = score_dict.get('Shareholder Rights', 'N/A')
    compensation_score = score_dict.get('Compensation', 'N/A')
    
    # Print the extracted information
    print("Corporate Governance Score:", total_score)
    print("Audit and Risk Oversight Score:", audit_score)
    print("Shareholders' Rights Score:", shareholder_score)
    print("Compensation Score:", compensation_score)
    print("Board Structure Score:", board_score)
else:
    print("无法找到公司治理板块,请检查页面结构或请求是否被拦截。")

关键说明

  • 请求头:User-Agent可根据自身使用的浏览器调整,确保请求不被拦截;
  • 正则提取:用正则匹配分数,比字符串索引更健壮,能适应小的文本格式变化;
  • 空值检查:先判断元素是否存在,避免None调用方法导致的错误;
  • 异常处理:response.raise_for_status()会在请求失败(如403、500)时抛出异常,方便排查问题。

内容的提问来源于stack exchange,提问作者Dua.S Ramzan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 13:47:51