如何使用Python从日志文件中提取各TagId的RSSI、纬度、经度并生成CSV文件
Hey there! Let's figure out how to parse those log lines and generate the CSV you need. The key here is to properly extract and combine the JSON data from both types of log entries, since regex alone won't handle the structured data cleanly. Here's a step-by-step solution:
Step 1: Understand the Log Structure
Your logs have two distinct entry types for each tag:
- One entry with
anchordata (RSSI values per anchor ID) andbatteryLevel - Another entry with
geographicdata (lng/lat) andbatteryLevel
We need to pair these entries by tagId to build complete rows for the CSV.
Step 2: Solution Code
This code will parse each line, extract the JSON, cache partial data, and write the final CSV once we have all required info for a tag:
import json import csv from collections import defaultdict import re def extract_json_from_log(line): # Extract the JSON blob from the log line (matches everything wrapped in {}) json_match = re.search(r'\{.*\}', line) if json_match: return json.loads(json_match.group()) return None def main(): # Cache to hold partial tag data: key = tagId, value = dict of rssi, battery, lng, lat tag_cache = defaultdict(lambda: {"rssi": {}, "battery": None, "lng": None, "lat": None}) all_anchor_ids = set() # Track all unique anchor IDs for CSV headers final_records = [] with open("log.txt", "r") as f: for line in f: line = line.strip() if not line: continue data = extract_json_from_log(line) if not data: continue tag_id = data.get("tagId") if not tag_id: continue # Handle entry with anchor (RSSI data) if "anchor" in data: anchor_list = data["anchor"] rssi_data = {} for anchor in anchor_list: # Clean up anchor ID by removing the "P " prefix anchor_id = anchor["ID"].replace("P ", "") rssi_data[anchor_id] = anchor["RSSI"] all_anchor_ids.add(anchor_id) # Update cache with RSSI and battery level tag_cache[tag_id]["rssi"] = rssi_data tag_cache[tag_id]["battery"] = data["tagData"]["batteryLevel"] # Handle entry with geographic data elif "geographic" in data: geo = data["geographic"] tag_cache[tag_id]["lng"] = geo["lng"] tag_cache[tag_id]["lat"] = geo["lat"] # Update battery level if present in this entry if "tagSensorData" in data: for sensor in data["tagSensorData"]: if sensor["type"] == "batteryLevel": tag_cache[tag_id]["battery"] = float(sensor["value"]) # Check if we have complete data for this tag tag_info = tag_cache[tag_id] if tag_info["rssi"] and tag_info["lng"] is not None and tag_info["lat"] is not None: # Build the full record row record = {"tagId": tag_id} # Add RSSI values for each anchor (fill empty if missing) for anchor_id in sorted(all_anchor_ids): record[anchor_id] = tag_info["rssi"].get(anchor_id, "") # Add geographic and battery data record["lng"] = tag_info["lng"] record["lat"] = tag_info["lat"] record["batteryLevel"] = tag_info["battery"] final_records.append(record) # Clear cache for this tag (optional, works if logs are paired one-to-one) del tag_cache[tag_id] # Generate the final CSV # Define headers: sorted anchor IDs + required core fields headers = ["tagId"] + sorted(all_anchor_ids) + ["lng", "lat", "batteryLevel"] with open("tag_data.csv", "w", newline="") as csvfile: writer = csv.DictWriter(csvfile, fieldnames=headers) writer.writeheader() writer.writerows(final_records) print("CSV generated successfully!") if __name__ == "__main__": main()
Step 3: How This Works
- Extract JSON: The
extract_json_from_logfunction uses regex to yank out the full JSON payload from each log line, ignoring theINFO:prefix. - Cache Partial Data: We use a
defaultdictto store incomplete tag data until we have both RSSI and geographic information. - Clean Anchor IDs: We strip the "P " prefix from anchor IDs to match your desired CSV header format.
- Build Complete Records: Once a tag has all required data, we construct a full record with every anchor's RSSI, geographic coordinates, and battery level.
- Write CSV: Python's
csv.DictWriterhandles header creation and row writing automatically, ensuring the output matches your requested format.
Why Your Previous Attempts Had Issues
- Regex Only: Regex is great for simple value extraction, but it can't reliably map RSSI values to their corresponding anchor IDs or handle nested structured data like JSON.
- Incorrect JSON Parsing: Your earlier JSON code tried splitting lines on
;},which doesn't match the actual log structure—we need to extract the full JSON blob first instead of splitting arbitrary strings.
Testing with Your Example Log
If you run this code with your sample log lines, it will generate a CSV with the correct headers and combined data for the tag 8593e96004019a2e.
内容的提问来源于stack exchange,提问作者Hpatck

