为何我的HTML解析器无法输出目标油价数值?
解决BeautifulSoup爬取油价返回None的问题
问题背景
我给编程作业的L/100KM油耗计算器加了每100公里油价计算功能,想用BeautifulSoup4从指定网站爬取实时油价并监控更新。已经找到了对应数值的CSS选择器,但运行代码后终端输出Initial number: none,没拿到预期的油价数值,后续的监控逻辑也没法正常执行。我的代码如下:
import requests from bs4 import BeautifulSoup import time # URL of the website to monitor url = 'https://nbeub.ca/index.php?page=current-petroleum-prices-2' # Function to fetch the number from the website def fetch_number(): response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # Adjust the selector to find the specific number number = soup.select_one('body > table > tbody > tr:nth-child(5) > td > table > tbody > tr > td > table > tbody > tr > td:nth-child(3) > table > tbody > tr:nth-child(3) > td:nth-child(2)') return str(number) # Main monitoring function def monitor(): last_number = fetch_number() print(f"Initial number: {last_number}") while True: time.sleep(2592000) # Wait for 30 days before checking again current_number = fetch_number() if current_number != last_number: print(f"Number updated: {current_number}") last_number = current_number # Start monitoring monitor()
问题原因分析
- CSS选择器冗余且不准确:选择器里包含多层
tbody,但很多网站的表格源码里并没有手动添加tbody,是浏览器渲染时自动补全的,直接用带tbody的选择器会匹配不到元素,导致select_one返回None。 - 未模拟浏览器请求:直接用
requests.get请求可能被网站识别为爬虫,返回的内容和浏览器看到的不一致,也会导致元素匹配失败。 - 返回值处理不当:即使找到元素,直接返回
str(number)会把整个标签对象转成字符串,而我们需要的是标签里的文本内容;如果没找到元素,返回str(None)就会输出none。
修复后的代码
import requests from bs4 import BeautifulSoup import time url = 'https://nbeub.ca/index.php?page=current-petroleum-prices-2' def fetch_number(): # 添加请求头模拟浏览器,避免被拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } try: response = requests.get(url, headers=headers) response.raise_for_status() # 检查请求是否成功 soup = BeautifulSoup(response.text, 'html.parser') # 简化选择器,去掉多余的tbody,定位到目标单元格 target_td = soup.select_one('table table table tr:nth-child(3) td:nth-child(2)') if target_td: # 提取文本并去除多余空格 price = target_td.get_text(strip=True) return price else: return "未找到油价数据" except Exception as e: return f"请求失败: {str(e)}" def monitor(): last_number = fetch_number() print(f"Initial number: {last_number}") # 测试时用短时间间隔,确认正常后再改回30天 check_interval = 10 # 2592000 while True: time.sleep(check_interval) current_number = fetch_number() if current_number != last_number: print(f"Number updated: {current_number}") last_number = current_number monitor()
关键修改点
- 添加请求头:用
headers模拟浏览器请求,降低被反爬拦截的概率。 - 简化CSS选择器:去掉自动生成的
tbody,用多层table直接定位,更贴合网页实际源码结构。 - 文本提取与错误处理:先判断元素是否存在,再用
get_text(strip=True)提取纯文本;添加try-except捕获请求异常,避免程序直接崩溃。 - 调整测试间隔:把30天的等待时间改成10秒,方便测试功能是否正常,确认没问题后再改回原间隔。
内容的提问来源于stack exchange,提问作者VXV
相关产品推荐
相关产品推荐

