Python字典如何将多ID映射到同一城市?爬虫代码优化求助
嘿,我来帮你搞定这个问题!
核心问题:字典键的类型限制
首先得说清楚为什么你原来的写法行不通:Python里字典的键必须是可哈希的不可变类型(比如整数、字符串、元组),而你尝试用的{522, 4745, 2718}是集合——属于可变类型,没法作为字典键,所以肯定达不到预期效果。
正确的思路是反转字典结构:用城市名作为键,对应的Region ID列表/集合作为值,这样同一城市的所有ID就自然归到一起了。
具体实现代码
方法1:用collections.defaultdict(最简便)
借助Python内置的defaultdict可以快速实现分组:
from collections import defaultdict # 你的原始Region ID映射 RegionIDArray = {522: "London", 4745: "London", 2718: "London", 3487: "Tokyo"} # 构建城市到Region ID的映射 city_to_regions = defaultdict(list) for region_id, city in RegionIDArray.items(): city_to_regions[city].append(region_id) # 输出结果:{'London': [522, 4745, 2718], 'Tokyo': [3487]} print(city_to_regions)
方法2:用普通字典(无需导入模块)
如果不想依赖额外模块,用基础字典也能实现:
RegionIDArray = {522: "London", 4745: "London", 2718: "London", 3487: "Tokyo"} city_to_regions = {} for region_id, city in RegionIDArray.items(): # 如果城市不在字典里,先初始化一个空列表 if city not in city_to_regions: city_to_regions[city] = [] city_to_regions[city].append(region_id)
爬虫代码优化示例
用这个反转后的字典结构,你的爬虫逻辑会更清晰,还能通过并发请求提高抓取效率,举个实用的示例:
import requests from threading import Thread import json import random import time # 定义抓取单个城市活动的函数 def fetch_city_events(city, region_ids): print(f"开始抓取【{city}】的活动...") all_events = [] # 遍历该城市的所有Region ID for region_id in region_ids: # 替换成你实际的目标URL target_url = f"https://your-target-site.com/events?region={region_id}" try: # 设置请求头,模拟浏览器访问(反爬必备) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 添加随机延迟,避免触发反爬 time.sleep(random.uniform(1, 3)) response = requests.get(target_url, headers=headers, timeout=10) response.raise_for_status() # 主动抛出HTTP错误 # 这里根据目标网站的结构解析数据,示例假设返回JSON events = response.json().get("data", {}).get("events", []) all_events.extend(events) except Exception as e: print(f"抓取Region ID {region_id}时出错:{str(e)}") continue # 去重(避免不同Region ID返回重复活动) unique_events = list({event["id"]: event for event in all_events}.values()) # 保存结果到JSON文件 with open(f"{city}_events.json", "w", encoding="utf-8") as f: json.dump(unique_events, f, ensure_ascii=False, indent=2) print(f"【{city}】抓取完成,共{len(unique_events)}条有效活动") # 启动多线程同时处理不同城市 threads = [] for city, region_ids in city_to_regions.items(): thread = Thread(target=fetch_city_events, args=(city, region_ids)) threads.append(thread) thread.start() # 等待所有线程结束 for thread in threads: thread.join() print("所有城市的活动抓取任务全部完成!")
额外实用建议
- 反爬应对:除了设置请求头和延迟,必要时可以使用代理IP池,避免被目标网站封禁。
- 数据校验:解析数据时增加字段校验,确保抓取到的活动数据格式符合预期。
- 可扩展性:后续新增Region ID时,直接加到原始的
RegionIDArray里,反转后的字典会自动完成分组,无需修改其他逻辑。
内容的提问来源于stack exchange,提问作者Serious Ruffy
相关产品推荐
相关产品推荐

