如何在Python中基于字典值的频率属性随机抽取元素
基于frequency加权随机抽取的正确实现
你遇到的TypeError: '<' not supported between instances of 'float' and 'Timedelta',是因为调用random.choices()时,误把包含Timedelta类型的子字典当成了权重或待抽取元素,导致函数内部数值比较时出现类型不匹配。
要实现基于frequency的加权抽取,核心是把待抽取的位置名称和对应的frequency数值分别提取出来,再正确传入random.choices()。
具体实现步骤
假设你的原始数据存在变量activity_dict中,以下是针对单个活动(比如Kitchen_Activity)的抽取代码,也可以扩展到所有活动:
- 提取目标活动的位置与对应权重
先把指定活动下的位置名称(字典键)和对应的frequency值(子字典的frequency字段)拆分出来:
import random from datetime import timedelta # 模拟你的原始数据 activity_dict = { 'Kitchen_Activity': { 'near the bathroom sink': {'frequency': 0, 'average duration': 0, 'standard deviation': 0}, 'near the fridge': {'frequency': 0.2631578947368421, 'average duration': timedelta(seconds=8.2), 'standard deviation': timedelta(seconds=8.2885)}, 'near the stove': {'frequency': 0.2631578947368421, 'average duration': timedelta(seconds=4.2), 'standard deviation': timedelta(seconds=0.8366)}, 'on the bed': {'frequency': 0, 'average duration': 0, 'standard deviation': 0}, 'near the shower': {'frequency': 0, 'average duration': 0, 'standard deviation': 0}, 'at the kitchen entrance from the hallway': {'frequency': 0.10526315789473684, 'average duration': timedelta(seconds=5), 'standard deviation': timedelta(seconds=1.4142)}, 'at the bedroom entrance': {'frequency': 0, 'average duration': 0, 'standard deviation': 0} }, 'Read': {}, 'Sleep': {} } # 提取Kitchen_Activity的位置和权重 target_activity = 'Kitchen_Activity' locations = list(activity_dict[target_activity].keys()) weights = [item['frequency'] for item in activity_dict[target_activity].values()]
- 调用random.choices完成加权抽取
直接传入位置列表和权重列表,k参数指定你要抽取的数量:
# 抽取1个位置 selected_location = random.choices(locations, weights=weights, k=1)[0] print(selected_location) # 抽取多个位置(比如5个) selected_locations = random.choices(locations, weights=weights, k=5) print(selected_locations)
注意事项
- 对于
frequency=0的位置,权重为0,几乎不会被抽中(符合你“低频元素极少被抽中”的需求),如果完全不想保留这些位置,可以在提取时过滤:# 过滤frequency>0的位置和权重 filtered_items = [(loc, data['frequency']) for loc, data in activity_dict[target_activity].items() if data['frequency'] > 0] locations = [item[0] for item in filtered_items] weights = [item[1] for item in filtered_items] - 如果要对所有活动(比如Read、Sleep)做抽取,只需要循环遍历
activity_dict的键,重复上述步骤即可。
内容的提问来源于stack exchange,提问作者accipigna
相关产品推荐
相关产品推荐

