从API自动下载CSV并处理空单元格、转换为同名JSON的Python实现问题
Hey there! Let's break down your two main issues and fix them step by step. You've got a solid foundation with your existing code, so we'll just tweak the parts that need adjustment to get things working smoothly with ArcGIS Pro and GeoEvent Server.
问题1:处理CSV中的空单元格(适配GeoEvent Server)
Empty cells (or empty strings) are likely tripping up GeoEvent Server, which usually expects consistent, non-empty field values. We have two practical solutions depending on your needs:
方案A:过滤掉含空值的行
If you only want to keep records where every field has a value, add a check before writing each row to the CSV. This ensures no empty cells make it into the final file:
# 在export_to_csv_job函数中,替换原有的行写入循环 for elem in to_proccess: # 生成当前行的所有值,用空字符串填充缺失字段 row_values = [elem.get(key, "") for key in most_keys.keys()] # 检查所有值是否都不为空(去除首尾空格后) if all(value.strip() != "" for value in row_values): csv_writer.writerow(row_values)
方案B:移除全为空的列
If some columns are completely empty across all records, GeoEvent might reject the CSV entirely. We can filter out these columns before writing:
# 在export_to_csv_job函数中,替换原有的表头和行写入逻辑 # 先收集所有行的数据 all_rows = [] for elem in to_proccess: all_rows.append([elem.get(key, "") for key in most_keys.keys()]) # 找出所有非全空的列索引 non_empty_col_indices = [] for col_idx in range(len(most_keys.keys())): # 检查该列是否有至少一个非空值 if any(row[col_idx].strip() != "" for row in all_rows): non_empty_col_indices.append(col_idx) # 筛选表头和行数据 filtered_headers = [list(most_keys.keys())[idx] for idx in non_empty_col_indices] filtered_rows = [[row[idx] for idx in non_empty_col_indices] for row in all_rows] # 写入CSV csv_writer.writerow(filtered_headers) for row in filtered_rows: csv_writer.writerow(row)
Pick the solution that fits your GeoEvent Server requirements—either removing bad rows or trimming useless columns will resolve the empty cell issue.
问题2:CSV转JSON时生成同名文件
Your current code overwrites a single data.json file every time, which isn't what you want. Let's modify it to generate a JSON file that matches the name of its source CSV:
修复后的CSV转JSON逻辑
import csv import json import time import os directory = "../csvFiles" def csv_to_json(csvFilePath, jsonFilePath): jsonArray = [] with open(csvFilePath, encoding="utf-8") as csvf: csvReader = csv.DictReader(csvf) # 可选:在这里也过滤掉含空值的行,保持数据一致性 for row in csvReader: if all(value.strip() != "" for value in row.values()): jsonArray.append(row) with open(jsonFilePath, "w", encoding="utf-8") as jsonf: jsonString = json.dumps(jsonArray, indent=4) jsonf.write(jsonString) # 遍历文件夹中的所有CSV,生成同名JSON for file in os.listdir(directory): if file.endswith(".csv"): # 拼接完整文件路径,避免找不到文件的问题 csv_full_path = os.path.join(directory, file) # 生成对应的JSON文件名(替换.csv为.json) json_file_name = os.path.splitext(file)[0] + ".json" json_full_path = os.path.join(directory, json_file_name) start = time.perf_counter() csv_to_json(csv_full_path, json_full_path) finish = time.perf_counter() print(f"Converted {file} to {json_file_name} successfully in {finish - start:0.4f} seconds")
关键改进点:
- Use
os.path.jointo create full file paths (prevents "file not found" errors if your script runs from a different directory) - Use
os.path.splitext(file)[0]to get the CSV filename without its extension, then append.jsonto make the matching JSON filename
额外优化:生成CSV后立即转JSON
Instead of running the conversion separately, you can trigger it right after generating the CSV in your export_to_csv_job function. This ensures your JSON is always up-to-date:
# 在export_to_csv_job函数中,data_file.close()之后添加: csv_file_path = f"csvFiles/data_{old_time_from.strftime('%m_%d_%Y_%H_%M_%S')}_{datetime.now().strftime('%m_%d_%Y_%H_%M_%S')}.csv" json_file_path = os.path.splitext(csv_file_path)[0] + ".json" csv_to_json(csv_file_path, json_file_path)
Just make sure the csv_to_json function is accessible in this script (either import it or define it directly).
备注:内容来源于stack exchange,提问作者blameUnicorn96

