使用Beautiful Soup爬取网页遇None及AttributeError问题求助
问题原因与解决办法
核心问题
你的代码找不到目标元素,返回None,本质是两个原因:
- 页面结构变更:OpenWeatherMap的页面元素类名已更新,
heading类不再对应温度标签。 - 反爬拦截:直接用
requests.get()请求会被网站识别为非浏览器请求,返回的页面内容不完整,导致无法定位元素。
解决步骤
验证请求返回内容
先打印page变量,查看是否是完整的页面HTML。如果内容不全,说明被反爬拦截,需要添加请求头模拟浏览器。添加请求头绕过反爬
给请求加上User-Agent,伪装成浏览器访问,示例代码:headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } page = requests.get(url, headers=headers).text定位正确的元素选择器
打开目标页面,用浏览器开发者工具(F12)查看温度元素的实际class或标签:- 当前页面中,温度值通常在
current-temp类的元素内,你可以替换代码中的选择器:temperature = doc.find(class_="current-temp") - 如果需要提取纯文本,可使用
temperature.get_text(strip=True)。
- 当前页面中,温度值通常在
增加错误判断
避免因找不到元素触发报错,添加判断逻辑:if temperature: print(temperature.get_text(strip=True)) else: print("未找到温度元素,请检查页面结构或请求头")
完整修改后的代码
from bs4 import BeautifulSoup import requests url = "https://openweathermap.org/city/4930956" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } page = requests.get(url, headers=headers).text doc = BeautifulSoup(page, "html.parser") temperature = doc.find(class_="current-temp") if temperature: print(temperature.get_text(strip=True)) else: print("未找到温度元素,请检查页面结构或请求头")
内容的提问来源于stack exchange,提问作者NOG POG
相关产品推荐
相关产品推荐

