Python网页爬虫无法通过CSS选择器定位元素获取电价
爬取sahko.tk当前电价的解决方案
问题分析
该网站的当前电价(芬兰语“Hinta nyt”下方)是静态渲染的,可直接通过BeautifulSoup定位,核心是找到正确的CSS选择器。
修正后的代码
import requests from bs4 import BeautifulSoup url = "https://sahko.tk/" # 定位当前电价的CSS选择器 element_selector = ".current-price" response = requests.get(url) response.encoding = 'utf-8' # 避免芬兰语特殊字符乱码 soup = BeautifulSoup(response.text, "html.parser") elements = soup.select(element_selector) if len(elements) == 0: print("No element found with selector '%s'" % element_selector) else: element_text = elements[0].text.strip() print(f"当前电价:{element_text}")
关键说明
- 正确的选择器:
.current-price是当前电价元素的类选择器,对应页面中显示价格的<span>标签 - 字符编码处理:添加
response.encoding = 'utf-8'确保芬兰语特殊字符正常解析 - 选择器方法:使用
soup.select()比find_all()更适配CSS选择器语法,定位更直观
备选定位方式(通过标题关联)
如果担心类名变更,也可以通过“Hinta nyt”标题找到相邻的价格元素:
import requests from bs4 import BeautifulSoup url = "https://sahko.tk/" response = requests.get(url) response.encoding = 'utf-8' soup = BeautifulSoup(response.text, "html.parser") # 先找到“Hinta nyt”标题,再获取其下一个兄弟元素 hinta_title = soup.find('h2', string='Hinta nyt') if hinta_title: price_element = hinta_title.find_next_sibling('span') if price_element: print(f"当前电价:{price_element.text.strip()}") else: print("未找到价格元素") else: print("未找到'Hinta nyt'标题")
内容的提问来源于stack exchange,提问作者Jho0
相关产品推荐
相关产品推荐

