如何解析米其林餐厅列表Div节点中的data-*属性数据?
提取HTML元素中data-*属性的方法
既然你已经能定位到目标.js-favorite-restaurant元素,接下来提取这些data-*属性很简单,以下是不同场景下的实现方式:
一、浏览器控制台/前端JavaScript提取
直接利用DOM元素的dataset属性,它会自动将data-开头的属性转成驼峰命名的键,也可以用getAttribute()直接获取完整属性名:
// 获取目标元素 const restaurantEl = document.querySelector('.js-favorite-restaurant'); // 1. 获取所有data属性(返回一个DOMStringMap对象) const allRestaurantData = restaurantEl.dataset; console.log(allRestaurantData.restaurantName); // 输出 "Joji" console.log(allRestaurantData.cookingType); // 输出 "Japanese" // 2. 单独获取某个属性,用getAttribute更直接 const city = restaurantEl.getAttribute('data-dtm-city'); console.log(city); // 输出 "New York"
二、Python爬虫(BeautifulSoup)
用BeautifulSoup解析HTML后,通过元素的get()方法或attrs字典提取:
from bs4 import BeautifulSoup # 假设你已经拿到目标HTML片段或整个页面源码 soup = BeautifulSoup(your_html_content, 'html.parser') restaurantEl = soup.select_one('.js-favorite-restaurant') # 提取单个属性 restaurantName = restaurantEl.get('data-restaurant-name') cookingType = restaurantEl.get('data-cooking-type') # 提取所有data属性(过滤出以data-开头的键值对) allDataAttrs = {key: value for key, value in restaurantEl.attrs.items() if key.startswith('data-')} print(allDataAttrs)
三、Python爬虫(Scrapy框架)
在Scrapy中可以直接通过CSS选择器的::attr()语法提取目标属性:
def parse(self, response): # 遍历所有符合选择器的餐厅元素 for restaurant in response.css('.js-favorite-restaurant'): yield { 'restaurant_name': restaurant.css('::attr(data-restaurant-name)').get(), 'cooking_type': restaurant.css('::attr(data-cooking-type)').get(), 'city': restaurant.css('::attr(data-dtm-city)').get(), 'district': restaurant.css('::attr(data-dtm-district)').get(), # 按需添加其他data属性提取逻辑 }
注意点
- JavaScript的
dataset会自动将data-后的连字符命名转为驼峰(比如data-restaurant-name→dataset.restaurantName),而爬虫工具中直接使用完整的属性名即可。 - 如果部分
data-属性值为空(比如data-dtm-chef),提取时会返回空字符串或None,可以按需处理。
内容的提问来源于stack exchange,提问作者scipio1551
相关产品推荐
相关产品推荐

