Python批量谷歌搜索脚本报429错误,改用谷歌API可行吗?
问题:批量谷歌搜索触发429错误,改用谷歌API能否解决?
我编写了如下Python脚本,用于批量对多个关键词执行谷歌搜索,筛选URL中包含“burj”的结果并写入CSV文件:
import csv import time from googlesearch import search def fetch_search_results(keywords): search_results = [] for keyword in keywords: results = [] for url in search(keyword, num_results=3): if 'burj' in url: results.append(url) search_results.append(','.join(results)) delay() # Add a delay between requests return search_results # Function to introduce a delay between requests def delay(): time.sleep(2) # Adjust the delay duration as needed # Read keywords from a CSV file def read_keywords_from_csv(file_path): keywords = [] with open(file_path, 'r') as file: reader = csv.reader(file) for row in reader: keyword = row[0] # Assuming the keyword is in the first column keywords.append(keyword) return keywords # Write search results to a CSV file def write_results_to_csv(results, file_path): with open(file_path, 'w', newline='') as file: writer = csv.writer(file) writer.writerow(['Keyword', 'Search Results']) for keyword, result in zip(keywords, results): writer.writerow([keyword, result]) # Example usage csv_file_path = "Surat_unmatched.csv" # Update with your CSV file path keywords = read_keywords_from_csv(csv_file_path) results = fetch_search_results(keywords) output_csv_file_path = "search_results.csv" # Update with desired output file path write_results_to_csv(results, output_csv_file_path)
但脚本运行时抛出如下错误:
429 Client Error: Too Many Requests for url: https://www.google.com/sorry/index?continue=https://www.google.com/search%3Fq%3DAvadh%252BUtopia%252BSurat%26num%3D3%26hl%3Den%26start%3D0&hl=en
谷歌判定该请求存在可疑行为,我原本期望得到包含Burj相关网页链接的CSV输出,请问改用谷歌API能否解决这个问题?
解答
改用谷歌官方的Google Custom Search JSON API完全可以解决这个429错误问题,核心原因如下:
- 你当前使用的
googlesearch库属于非官方爬虫工具,是模拟浏览器发送无授权请求,谷歌的反爬机制会识别这类行为并限制请求频率,最终触发429错误。 - 官方API是谷歌提供的合法请求渠道,需要通过API密钥授权访问,谷歌会为你分配固定的请求配额(免费版每月有100次免费请求,付费版可根据需求扩容),只要在配额内发起请求,就不会触发反爬限制,请求稳定性也更高。
改造思路及示例代码
- 先在谷歌云平台创建Custom Search Engine,获取你的API密钥和搜索引擎ID;
- 替换原脚本中的
googlesearch调用为官方API请求; - 修复原脚本中
write_results_to_csv依赖全局变量的bug,改为传入参数; - 保留原有的读取CSV、筛选URL的逻辑。
改造后的示例代码:
import csv import requests # 替换为你自己的API密钥和搜索引擎ID API_KEY = "YOUR_GOOGLE_API_KEY" SEARCH_ENGINE_ID = "YOUR_SEARCH_ENGINE_ID" def fetch_search_results(keywords): search_results = [] base_api_url = "https://www.googleapis.com/customsearch/v1" for keyword in keywords: # 构造API请求参数 request_params = { "q": keyword, "key": API_KEY, "cx": SEARCH_ENGINE_ID, "num": 3 # 对应原脚本的num_results参数 } # 发送API请求 response = requests.get(base_api_url, params=request_params) response_data = response.json() # 提取并筛选含"burj"的URL matched_urls = [] if "items" in response_data: for item in response_data["items"]: url = item["link"] if "burj" in url.lower(): # 忽略大小写匹配更全面 matched_urls.append(url) search_results.append(','.join(matched_urls)) return search_results # 读取CSV中的关键词(复用原逻辑,添加空行判断) def read_keywords_from_csv(file_path): keywords = [] with open(file_path, 'r', encoding='utf-8') as file: reader = csv.reader(file) for row in reader: if row: # 跳过空行 keywords.append(row[0]) return keywords # 写入结果到CSV(修复全局变量问题,添加keywords参数) def write_results_to_csv(keywords, results, file_path): with open(file_path, 'w', newline='', encoding='utf-8') as file: writer = csv.writer(file) writer.writerow(['Keyword', 'Search Results']) for keyword, result in zip(keywords, results): writer.writerow([keyword, result]) # 示例调用 if __name__ == "__main__": input_csv_path = "Surat_unmatched.csv" output_csv_path = "search_results.csv" keywords = read_keywords_from_csv(input_csv_path) results = fetch_search_results(keywords) write_results_to_csv(keywords, results, output_csv_path)
额外提示
- 免费版API配额有限,如果关键词数量较多,建议升级到付费版;
- API请求无需额外添加延迟,谷歌会根据配额处理请求,不会触发429错误;
- 原脚本中
fetch_search_results函数里的缩进错误(results.append(url)前缺少缩进)也需要修复,否则会抛出语法错误。
内容的提问来源于stack exchange,提问作者laviano
相关产品推荐
相关产品推荐

