按15分钟时间区间统计WiFi连接用户位置的问题求助
WiFi连接日志15分钟区间分析问题
我正在处理一份WiFi连接日志文件,需求是按每15分钟的时间区间分析数据,将同一时间区间内用户的连接按连接位置分组,统计用户在该区间内连接次数最多的位置。
输入数据
{'User-Name': {0: 'Thiago', 1: 'Ana', 2: 'Ana', 3: 'Ana', 4: 'Jose', 5: 'Thiago', 6: 'Thiago', 7: 'Thiago', 8: 'Ana', 9: 'Thiago', 10: 'Thiago', 11: 'Thiago', 12: 'Thiago', 13: 'Ana', 14: 'Thiago', 15: 'Thiago', 16: 'Jose'}, 'NAS-Identifier': {0: 'Cantina', 1: 'Cantina', 2: 'Cantina', 3: 'Dex', 4: 'Dcc', 5: 'Cantina', 6: 'Cantina', 7: 'Cantina', 8: 'Dex', 9: 'Cantina', 10: 'Dex', 11: 'Dex', 12: 'Dcc', 13: 'Dex', 14: 'Dcc', 15: 'Dcc', 16: 'Dex'}, 'Event-Timestamp': {0: 'May 6 2021 08:01:00', 1: 'May 6 2021 08:01:00', 2: 'May 6 2021 08:03:00', 3: 'May 6 2021 08:04:00', 4: 'May 6 2021 08:05:00', 5: 'May 6 2021 08:06:00', 6: 'May 6 2021 08:07:00', 7: 'May 6 2021 08:08:00', 8: 'May 6 2021 08:09:00', 9: 'May 6 2021 08:10:00', 10: 'May 6 2021 08:11:00', 11: 'May 6 2021 08:12:00', 12: 'May 6 2021 08:13:00', 13: 'May 6 2021 08:14:00', 14: 'May 6 2021 08:18:00', 15: 'May 6 2021 08:23:00', 16: 'May 6 2021 08:25:00'}, 'Type': {0: 'update', 1: 'update', 2: 'stop', 3: 'start', 4: ' stop', 5: 'update', 6: 'start', 7: 'start', 8: 'update', 9: 'stop', 10: 'start', 11: 'stop', 12: 'start', 13: 'update', 14: 'update', 15: 'update', 16: ' start'}}
现有代码(判断记录所属时间区间)
df_registro = pd.read_csv("entrada_locais_aps.csv") df_registro['Event-Timestamp'] = df_registro['Event-Timestamp'].apply(lambda x: datetime.strptime(x,'%b %d %Y %H:%M:%S')) #converting Event-Timestamp column to datetime type horarios = { 'intervalo': { 1: (time(7, 0), time(7, 15)), # range 1 = 07:00h - 07:15h 2: (time(7, 15), time(7, 30)), # range 2 = 07:15h - 07:30h 3: (time(7, 30), time(7, 45)), # range 3 = 07:30h - 07:45h 4: (time(7, 45), time(8, 0)), # range 4 = 07:45h - 08:00h 5: (time(8, 0), time(8, 15)), # range 5 = 08:00h - 08:15h 6: (time(8, 15), time(8, 30)), # range 6 = 08:15h - 08:30h 7: (time(8, 30), time(8, 45)), # range 7 = 08:30h - 08:45h 8: (time(8, 45), time(9, 0)), # range 8 = 08:45h - 09:00h } } encontrado = False for nome, intervalo in horarios.items(): i=1 lista=[] while i < 9: #I have 8 intervals of 15 in 15 minutes y=0 inicio, fim = intervalo[i] while y < 15: #15 is number of register dataframe registro2 = df_registro['Event-Timestamp'].dt.time[y] if inicio <= registro2 < fim: # checks if time is between the start and end of the range print(df_registro['User-Name'][y], "-",df_registro['NAS-Identifier'][y]) print(f'The record is within the range {inicio:%H:%M} to {fim:%H:%M}') print("--------------") lista.append([df_registro['User-Name'][y],df_registro['NAS-Identifier'][y],df_registro['Event-Timestamp'].dt.time[y]]) encontrado = True y += 1 i += 1 if not encontrado: print('Register out of range')
代码执行结果
Thiago - Cantina The record is within the range 08:00 to 08:15 Ana - Cantina The record is within the range 08:00 to 08:15 Ana - Cantina The record is within the range 08:00 to 08:15 Ana - DEX The record is within the range 08:00 to 08:15 Jose - DCC The record is within the range 08:00 to 08:15 Thiago - Cantina The record is within the range 08:00 to 08:15 Thiago - Cantina The record is within the range 08:00 to 08:15 Thiago - Cantina The record is within the range 08:00 to 08:15 Ana - DEX The record is within the range 08:00 to 08:15 Thiago - Cantina The record is within the range 08:00 to 08:15 Thiago - DEX The record is within the range 08:00 to 08:15 Thiago - DEX The record is within the range 08:00 to 08:15 Thiago - DCC The record is within the range 08:00 to 08:15 Ana - DEX The record is within the range 08:00 to 08:15 Thiago - DCC The record is within the range 08:15 to 08:30
统计代码及错误结果
统计代码
df_registro3 = pd.DataFrame(lista,columns=['Name', 'Place', 'Hour']) df4 = df_registro3.groupby(['Name','Place']).count() print(df4)
错误统计结果
Name Place Ana Cantina 2 DEX 3 Jose DCC 1 Thiago Cantina 5 DCC 2 DEX 2
问题说明
上述统计结果存在错误:用户Thiago的DCC位置有一条记录属于08:15-08:30区间,另一条属于08:00-08:15区间,不应被合并计数为2次。问题出在列表创建环节,尝试为每个时间区间创建独立列表但未成功,需要解决该问题。
内容的提问来源于stack exchange,提问作者Thiago Ramos
相关产品推荐
相关产品推荐

