如何将每个短代码搜索所得链接存储至列表或字典?
嘿,我来帮你搞定这个需求!要把每个短代码对应的搜索链接存起来,用字典或者列表都很顺手,我给你两种实用的实现方式参考:
方案1:用字典存储(首推,查找更方便)
这种方式把短代码作为键,对应的搜索链接列表作为值,后续要找某个短代码的结果直接通过键就能拿到,非常高效。
你只需要在代码开头初始化一个空字典,然后在遍历搜索结果的时候,把链接逐个添加到对应短代码的列表里就行。修改后的完整代码如下:
import pandas as pd import numpy as np import time from tqdm import tqdm import requests from googlesearch import search from requests.exceptions import HTTPError # 初始化存储结果的字典:键是短代码,值是对应的链接列表 shortcode_links_dict = {} try: from googlesearch import search except ImportError: print("No module named 'google' found") exit() # 读取Excel里的短代码列表 with open('Unknown.xlsx', "rb") as f: df = pd.read_excel(f) shortcode_list = df['Short Code'].tolist() def stopwatch(sec): while sec: minn, sec = divmod(sec, 60) timeformat = '{:02d}:{:02d}'.format(minn, sec) print(timeformat, end='\r') time.sleep(1) sec -= 1 # 记得定义请求头,避免被Google拦截(示例) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 遍历每个短代码进行搜索 for i in tqdm(range(len(shortcode_list))): shortcode = shortcode_list[i] # 初始化当前短代码的链接列表 shortcode_links_dict[shortcode] = [] try: # 随机延迟和暂停,避免被封 delay = np.random.choice(np.arange(1, 60, 1).tolist()) pause = np.random.choice(np.arange(2, 8, 1).tolist()) stopwatch(delay) query = f"text * to \"{shortcode}\"" url = f'https://www.google.com?q={query}' res = requests.get(url, headers=headers) print(f"\nThe query will be {query} {res}") # 搜索并存储链接 for k in search(query, tld="co.in", num=10, stop=10, pause=pause, country='US', user_agent=googlesearch.get_random_user_agent(), verify_ssl=True): print(k) shortcode_links_dict[shortcode].append(k) except HTTPError as exception: if exception.code == 429: print(exception) print("Waiting for 8 minutes and Continue") stopwatch(480) # 重新尝试当前短代码的搜索 continue else: # 其他HTTP错误,标记当前短代码无结果 shortcode_links_dict[shortcode] = ["搜索失败:" + str(exception)] except Exception as e: # 其他意外错误 shortcode_links_dict[shortcode] = ["搜索出错:" + str(e)] # 打印存储的结果(可选) print("\n所有短代码的搜索结果:") for sc, links in shortcode_links_dict.items(): print(f"\n短代码 {sc} 的链接:") for link in links: print(link)
方案2:用列表存储(适合按顺序遍历)
如果你需要保留短代码的搜索顺序,用列表存储也可以,每个元素是一个包含短代码和对应链接列表的字典(或者元组):
修改核心部分的代码:
# 初始化存储结果的列表 shortcode_links_list = [] # 遍历每个短代码时的存储逻辑 for i in tqdm(range(len(shortcode_list))): shortcode = shortcode_list[i] current_result = {"shortcode": shortcode, "links": []} try: # 原有的延迟、搜索逻辑和上面一致 delay = np.random.choice(np.arange(1, 60, 1).tolist()) pause = np.random.choice(np.arange(2, 8, 1).tolist()) stopwatch(delay) query = f"text * to \"{shortcode}\"" url = f'https://www.google.com?q={query}' res = requests.get(url, headers=headers) print(f"\nThe query will be {query} {res}") for k in search(query, tld="co.in", num=10, stop=10, pause=pause, country='US', user_agent=googlesearch.get_random_user_agent(), verify_ssl=True): print(k) current_result["links"].append(k) # 把当前结果加入列表 shortcode_links_list.append(current_result) except Exception as e: current_result["links"] = ["搜索出错:" + str(e)] shortcode_links_list.append(current_result) # 打印结果示例 for item in shortcode_links_list: print(f"\n短代码 {item['shortcode']} 的链接:") for link in item['links']: print(link)
小提醒
- 一定要定义
headers变量模拟浏览器请求,不然Google很容易拦截你的请求。 - 如果某个短代码搜索失败,两种方案都会把错误信息存进去,方便后续排查问题。
- 要是想把结果导出成文件,比如用
pd.DataFrame.from_dict(shortcode_links_dict, orient='index').to_excel('搜索结果.xlsx')就能把字典导出成Excel。
内容的提问来源于stack exchange,提问作者Mohamed Elgendy
相关产品推荐
相关产品推荐

