如何通过API完整获取赫尔辛基市场数据?代码问题排查求助
无法从Nasdaq OMX Nordic API获取全部数据的问题
我尝试抓取Nasdaq OMX Nordic网站赫尔辛基主市场的全部公司新闻数据,但结果中缺失了部分内容(比如2022-01-03 8:00的Vikings Exchange相关数据)。以下是我的代码:
import requests import json import time import csv import pandas start=250 with open('C:/Users/apskaita3/Desktop/number2.txt', "r") as f: start= f.readlines() start=int(start[0]) start=start + 70 results = {"item": {}} # Todo load json for i in range(0,9800): #<----- Just change range here to increase number of requests URL = f"https://api.news.eu.nasdaq.com/news/query.action?type=handleResponse&showAttachments=true&showCnsSpecific=true&showCompany=true&countResults=false&freeText=&company=&market=Main%20Market%2C+Helsinki&cnscategory=&fromDate=&toDate=&globalGroup=exchangeNotice&globalName=NordicMainMarkets&displayLanguage=en&language=en&timeZone=CET&dateMask=yyyy-MM-dd+HH%3Amm%3Ass&limit=19&start={i}&dir=ASC" r = requests.get(url = URL) #time.sleep(1) res = r.text.replace("handleResponse(", "") res_json = json.loads(res) data = res_json print("Doing: " + str(i + 1) + "th") downloaded_entries = data["results"]["item"] new_entries = [d for d in downloaded_entries if d["headline"] not in results["item"]] start=str(start) for entry in new_entries: if entry["market"] == 'Main Market, Helsinki' and entry["published"]>="2021-10-20 06:30:00": headline = entry["headline"].strip() published = entry["published"] market=entry["market"] market="Main Market, Helsinki" results["item"][headline] = {"company": entry["company"], "messageUrl": entry["messageUrl"], "published": entry["published"], "headline": headline} print(entry['market']) print(f"Market: {market}/nDate: {published}/n") with open("C:/Users/apskaita3/Finansų analizės ir valdymo sprendimai, UAB/Rokas Toomsalu - Power BI analitika/Integracijos/1_Public comapnies analytics/Databasesets/Others/market_news_helsinki.json", "w") as outfile: json_object = json.dumps({"item": list(results["item"].values())}, indent = 4) outfile.write(json_object) with open("C:/Users/apskaita3/Desktop/number2.txt", "w") as outfile1: outfile1.write(start) # type: ignore
问题原因分析
- 分页逻辑错误:API的
start参数是分页起始位置,配合limit=19,每页返回19条数据。你用range(0,9800)循环,每次start={i},相当于每次只偏移1条,会重复请求大量重复数据,同时跳过绝大多数页面,直接导致漏抓数据。 - 去重逻辑缺陷:用
headline作为去重key,存在相同标题但不同内容的新闻,会被误判为重复而跳过,丢失有效数据。 - 无请求频率控制:注释了
time.sleep(1),高频请求可能被API限流,导致部分请求返回不完整或失败。 - 未处理边界情况:当
start超过总数据量时,API返回空条目,代码仍会继续循环,做无效请求且无法终止。
修正后的代码
import requests import json import time # 读取上次起始位置,不存在则从0开始 start = 0 try: with open('C:/Users/apskaita3/Desktop/number2.txt', "r") as f: start = int(f.readlines()[0]) except FileNotFoundError: pass results = {} # 用messageUrl作为唯一标识去重 limit = 19 # 每页固定返回19条 max_requests = 500 # 防止无限循环,可按需调整 target_market = 'Main Market, Helsinki' min_date = "2021-10-20 06:30:00" total_requested = 0 while total_requested < max_requests: # 构造正确的分页URL URL = f"https://api.news.eu.nasdaq.com/news/query.action?type=handleResponse&showAttachments=true&showCnsSpecific=true&showCompany=true&countResults=false&freeText=&company=&market=Main%20Market%2C+Helsinki&cnscategory=&fromDate=&toDate=&globalGroup=exchangeNotice&globalName=NordicMainMarkets&displayLanguage=en&language=en&timeZone=CET&dateMask=yyyy-MM-dd+HH%3Amm%3Ass&limit={limit}&start={start}&dir=ASC" try: r = requests.get(url=URL) r.raise_for_status() # 检查请求是否成功 except requests.exceptions.RequestException as e: print(f"请求失败: {e}") time.sleep(5) continue # 处理API的包装格式,注意去掉结尾的括号 res = r.text.replace("handleResponse(", "").rstrip(")") try: res_json = json.loads(res) except json.JSONDecodeError as e: print(f"解析JSON失败: {e}") time.sleep(5) continue downloaded_entries = res_json["results"]["item"] if not downloaded_entries: print("无更多数据,终止循环") break new_count = 0 for entry in downloaded_entries: if entry["market"] == target_market and entry["published"] >= min_date: unique_key = entry["messageUrl"] if unique_key not in results: results[unique_key] = { "company": entry["company"], "messageUrl": entry["messageUrl"], "published": entry["published"], "headline": entry["headline"].strip(), "market": entry["market"] } new_count += 1 print(f"新增数据: {entry['published']} - {entry['headline'][:50]}...") print(f"完成第{total_requested+1}次请求,新增{new_count}条数据") start += limit total_requested += 1 time.sleep(1) # 控制请求频率,避免限流 # 保存结果 output_path = "C:/Users/apskaita3/Finansų analizės ir valdymo sprendimai, UAB/Rokas Toomsalu - Power BI analitika/Integracijos/1_Public comapnies analytics/Databasesets/Others/market_news_helsinki.json" with open(output_path, "w", encoding="utf-8") as outfile: json.dump({"item": list(results.values())}, outfile, indent=4, ensure_ascii=False) # 保存下次起始位置 with open("C:/Users/apskaita3/Desktop/number2.txt", "w") as outfile1: outfile1.write(str(start)) print(f"抓取完成,共获取{len(results)}条有效数据")
关键优化点
- 修正分页逻辑:用
start += limit确保每次请求下一页完整数据,无跳过无重复 - 优化去重逻辑:用
messageUrl作为唯一标识,避免因标题重复丢失数据 - 增加错误处理:处理请求失败、JSON解析失败的情况,避免程序崩溃
- 控制请求频率:恢复
time.sleep(1),降低被API限流的风险 - 自动终止循环:API返回空条目时自动停止,减少无效请求
内容的提问来源于stack exchange,提问作者jimm denison
相关产品推荐
相关产品推荐

