You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python字典如何将多ID映射到同一城市?爬虫代码优化求助

嘿,我来帮你搞定这个问题!

核心问题:字典键的类型限制

首先得说清楚为什么你原来的写法行不通:Python里字典的键必须是可哈希的不可变类型(比如整数、字符串、元组),而你尝试用的{522, 4745, 2718}是集合——属于可变类型,没法作为字典键,所以肯定达不到预期效果。

正确的思路是反转字典结构:用城市名作为键,对应的Region ID列表/集合作为值,这样同一城市的所有ID就自然归到一起了。

具体实现代码

方法1:用collections.defaultdict(最简便)

借助Python内置的defaultdict可以快速实现分组:

from collections import defaultdict

# 你的原始Region ID映射
RegionIDArray = {522: "London", 4745: "London", 2718: "London", 3487: "Tokyo"}

# 构建城市到Region ID的映射
city_to_regions = defaultdict(list)
for region_id, city in RegionIDArray.items():
    city_to_regions[city].append(region_id)

# 输出结果:{'London': [522, 4745, 2718], 'Tokyo': [3487]}
print(city_to_regions)

方法2:用普通字典(无需导入模块)

如果不想依赖额外模块,用基础字典也能实现:

RegionIDArray = {522: "London", 4745: "London", 2718: "London", 3487: "Tokyo"}

city_to_regions = {}
for region_id, city in RegionIDArray.items():
    # 如果城市不在字典里,先初始化一个空列表
    if city not in city_to_regions:
        city_to_regions[city] = []
    city_to_regions[city].append(region_id)

爬虫代码优化示例

用这个反转后的字典结构,你的爬虫逻辑会更清晰,还能通过并发请求提高抓取效率,举个实用的示例:

import requests
from threading import Thread
import json
import random
import time

# 定义抓取单个城市活动的函数
def fetch_city_events(city, region_ids):
    print(f"开始抓取【{city}】的活动...")
    all_events = []
    # 遍历该城市的所有Region ID
    for region_id in region_ids:
        # 替换成你实际的目标URL
        target_url = f"https://your-target-site.com/events?region={region_id}"
        try:
            # 设置请求头,模拟浏览器访问(反爬必备)
            headers = {
                "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
            }
            # 添加随机延迟,避免触发反爬
            time.sleep(random.uniform(1, 3))
            response = requests.get(target_url, headers=headers, timeout=10)
            response.raise_for_status()  # 主动抛出HTTP错误
            
            # 这里根据目标网站的结构解析数据,示例假设返回JSON
            events = response.json().get("data", {}).get("events", [])
            all_events.extend(events)
            
        except Exception as e:
            print(f"抓取Region ID {region_id}时出错:{str(e)}")
            continue
    
    # 去重(避免不同Region ID返回重复活动)
    unique_events = list({event["id"]: event for event in all_events}.values())
    
    # 保存结果到JSON文件
    with open(f"{city}_events.json", "w", encoding="utf-8") as f:
        json.dump(unique_events, f, ensure_ascii=False, indent=2)
    
    print(f"【{city}】抓取完成,共{len(unique_events)}条有效活动")

# 启动多线程同时处理不同城市
threads = []
for city, region_ids in city_to_regions.items():
    thread = Thread(target=fetch_city_events, args=(city, region_ids))
    threads.append(thread)
    thread.start()

# 等待所有线程结束
for thread in threads:
    thread.join()

print("所有城市的活动抓取任务全部完成!")

额外实用建议

  • 反爬应对:除了设置请求头和延迟,必要时可以使用代理IP池,避免被目标网站封禁。
  • 数据校验:解析数据时增加字段校验,确保抓取到的活动数据格式符合预期。
  • 可扩展性:后续新增Region ID时,直接加到原始的RegionIDArray里,反转后的字典会自动完成分组,无需修改其他逻辑。

内容的提问来源于stack exchange,提问作者Serious Ruffy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:56:37