无法从指定网站获取data-score值,求解决方案
解决Kununu网站data-score抓取的KeyError问题
问题原因分析
- 语法错误:原代码中
soup.find(...) ('data-score')是错误写法,正确的属性访问应该用.get()方法或字典索引,且字典索引在属性不存在时会直接抛出KeyError。 - 动态类名不可靠:你使用的类名(如
index__stars__nfK6S)带有随机后缀,这类动态生成的类名会随网站更新变化,无法稳定定位元素。 - 反爬/动态渲染限制:直接用
requests.get可能无法获取JS渲染后的真实页面内容,或者被网站识别为爬虫拦截。
解决方案
方案1:优化Requests+BeautifulSoup(静态页面场景)
通过属性选择器定位元素,添加请求头模拟浏览器,同时增加异常处理:
from bs4 import BeautifulSoup as bs import requests url = 'https://www.kununu.com/de/pan-dacom-networking4/bewertung/40726463-005e-45d9-af11-37e7afbd5110' # 模拟浏览器请求头,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) soup = bs(response.text, 'html.parser') # 直接定位所有带data-score属性的span元素,避开动态类名 score_spans = soup.find_all('span', attrs={'data-score': True}) # 获取第二个评分值 if len(score_spans) >= 2: second_score = score_spans[1].get('data-score')[0:3] print(second_score) else: print("未找到第二个带有data-score属性的span元素")
方案2:使用Selenium处理动态渲染页面
如果页面内容是JS动态生成的,requests无法获取完整内容,需要用浏览器自动化工具加载页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options url = 'https://www.kununu.com/de/pan-dacom-networking4/bewertung/40726463-005e-45d9-af11-37e7afbd5110' # 配置Chrome无头模式(后台运行) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 通过CSS选择器定位所有带data-score的span score_spans = driver.find_elements(By.CSS_SELECTOR, 'span[data-score]') if len(score_spans) >= 2: second_score = score_spans[1].get_attribute('data-score')[0:3] print(second_score) else: print("未找到第二个带有data-score属性的span元素") driver.quit()
关键注意事项
- 用
get('data-score')替代字典索引['data-score'],前者在属性不存在时返回None,不会直接抛出异常,更安全。 - 避免使用带随机后缀的动态类名,优先选择
data-*属性、固定ID或稳定的父元素结构定位。 - 若网站反爬严格,可进一步添加Cookie、代理IP等请求参数,或延长Selenium的页面等待时间确保内容加载完成。
内容的提问来源于stack exchange,提问作者Masud Al Nahid
相关产品推荐
相关产品推荐

