如何修复CSV屏幕耗时统计中的‘unhashable type: list’错误?
解决统计用户屏幕耗时时的「unhashable type: list」错误
问题背景
嘿,我看到你在基于CSV统计用户各屏幕耗时的时候碰到了unhashable type: list的错误,这是个Python字典使用中很常见的类型问题,我来帮你理清楚问题所在并给出修正方案。
你的原始代码
fn3 = 'screenviewclean.1.csv' f2 = open(fn3,"r") filereader = csv.reader(f2) mhist = {} i = 0 for line in filereader: i+=1 sec= line [10] IpAddress = line [2] timeStamp = line [6] time = timeStamp[11:13]+ timeStamp[13:19] Screen_View = line [7] if i>1: if IpAddress in mhist.keys(): mhist[IpAddress].append(Screen_View) else: mhist[IpAddress] = [Screen_View] #print (mhist) chist = {} for ip, screen in mhist.items(): k = screen #print (k) if k in chist.keys(): chist[k].append (time) else: chist[k] = [time] print (chist)
你的CSV样本数据
Unnamed: 0,lastLoggedVersion,IpAddress,deviceId,deviceOS,userId,timeStamp,screenName,userType,doc.id,seconds 0,1.6.0.1,192.168.0.77,7612F62D-E392-4269-B49B-4F1214AA3888,iOS13.6.1,5U1XW8wkoqUPCTGhC1ni9Whinvt1,2020-11-13 22:28:55.029000+00:00,StudentProfile,student,00mrvPyS9Y2Al9iTN1vw,1231534.547 2,1.6.1.44,10.0.2.16,40a4dc7cb837fdec,Android10,27lFw6EnfbYFsU3F8AEejYGQRRl1,2020-11-12 21:28:00.998000+00:00,CompanySettings,company,01dMOvAgsRTPSWXTDXIh,1141480.516 6,1.6.0.43,192.168.87.241,ec62706b2834bfcc,Android9,7XBtY5ZDxcPWYF7sGECDZxnH71b2,2020-11-12 21:41:33.126000+00:00,DiscoverCompanies,student,064kDvawK8cJoRl5if9d,1142292.644
错误原因拆解
你代码里的核心问题出在第二个循环:
for ip, screen in mhist.items(): k = screen # 这里的screen是一个列表(比如["StudentProfile"]) if k in chist.keys(): chist[k].append(time)
Python字典的键必须是可哈希类型(比如字符串、数字、元组),而列表是不可哈希的(因为它是可变对象),所以当你尝试把列表screen作为chist的键时,就会抛出unhashable type: list错误。
另外还有两个逻辑小问题:
- 你用
timeStamp截取的time并不是实际耗时,CSV里已经有现成的seconds字段,直接用它更准确 - 当前的
mhist只存了每个用户的屏幕列表,没有关联对应的耗时数据,无法完成统计需求
修正后的解决方案
我们需要调整数据结构,用嵌套字典来存储「用户IP → 屏幕名称 → 耗时列表」的关系,最后再统计每个用户每个屏幕的总耗时/平均耗时。
优化后的代码
import csv fn3 = 'screenviewclean.1.csv' # 用DictReader通过字段名访问数据,避免索引记错的问题 with open(fn3, "r", newline="", encoding="utf-8") as f2: reader = csv.DictReader(f2) # 构建嵌套字典:{IpAddress: {screenName: [耗时1, 耗时2, ...]}} user_screen_time = {} for row in reader: ip = row["IpAddress"] screen = row["screenName"] # 把seconds转成float类型,方便后续计算 try: seconds = float(row["seconds"]) except ValueError: # 处理无效的耗时数据,避免程序崩溃 print(f"跳过无效耗时数据:{row}") continue # 初始化用户的字典 if ip not in user_screen_time: user_screen_time[ip] = {} # 初始化屏幕的耗时列表 if screen not in user_screen_time[ip]: user_screen_time[ip][screen] = [] # 添加耗时数据 user_screen_time[ip][screen].append(seconds) # 统计每个用户每个屏幕的总耗时(也可以改成计算平均耗时) user_screen_total_time = {} for ip, screen_times in user_screen_time.items(): user_screen_total_time[ip] = { screen: sum(times) for screen, times in screen_times.items() } # 打印清晰的结果 print("每个用户各屏幕的总耗时:") for ip, screen_total in user_screen_total_time.items(): print(f"IP: {ip}") for screen, total in screen_total.items(): print(f" 屏幕 {screen}: {total:.2f} 秒")
代码说明
- 使用
csv.DictReader:通过字段名直接访问数据,比索引更易维护,不容易出错 - 嵌套字典结构:清晰存储每个用户对应每个屏幕的所有耗时记录
- 类型转换:把CSV中的字符串类型
seconds转成float,方便后续求和/计算平均值 - 异常处理:跳过无效的耗时数据,避免程序崩溃
- 最终统计:用字典推导式快速计算每个屏幕的总耗时,结果直观清晰
运行结果示例
对于你提供的CSV样本,运行后会输出:
每个用户各屏幕的总耗时: IP: 192.168.0.77 屏幕 StudentProfile: 1231534.55 秒 IP: 10.0.2.16 屏幕 CompanySettings: 1141480.52 秒 IP: 192.168.87.241 屏幕 DiscoverCompanies: 1142292.64 秒
内容的提问来源于stack exchange,提问作者Mohamed
相关产品推荐
相关产品推荐

