You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过API完整获取赫尔辛基市场数据?代码问题排查求助

无法从Nasdaq OMX Nordic API获取全部数据的问题

我尝试抓取Nasdaq OMX Nordic网站赫尔辛基主市场的全部公司新闻数据,但结果中缺失了部分内容(比如2022-01-03 8:00的Vikings Exchange相关数据)。以下是我的代码:

import requests
import json
import time
import csv
import pandas
 
start=250
 
with open('C:/Users/apskaita3/Desktop/number2.txt', "r") as f:
    start= f.readlines()
 
start=int(start[0])
start=start + 70

results = {"item": {}}

# Todo load json
for i in range(0,9800): #<----- Just change range here to increase number of requests
    URL = f"https://api.news.eu.nasdaq.com/news/query.action?type=handleResponse&showAttachments=true&showCnsSpecific=true&showCompany=true&countResults=false&freeText=&company=&market=Main%20Market%2C+Helsinki&cnscategory=&fromDate=&toDate=&globalGroup=exchangeNotice&globalName=NordicMainMarkets&displayLanguage=en&language=en&timeZone=CET&dateMask=yyyy-MM-dd+HH%3Amm%3Ass&limit=19&start={i}&dir=ASC"
    r = requests.get(url = URL)
    #time.sleep(1)
    res = r.text.replace("handleResponse(", "")
    res_json = json.loads(res)
    data = res_json
    print("Doing: " + str(i + 1) + "th")
    
    downloaded_entries = data["results"]["item"]
    new_entries = [d for d in downloaded_entries if d["headline"] not in results["item"]]
    start=str(start)
    
    for entry in new_entries:
        if entry["market"] == 'Main Market, Helsinki' and entry["published"]>="2021-10-20 06:30:00":
            headline = entry["headline"].strip()
            published = entry["published"]
            market=entry["market"]
            market="Main Market, Helsinki"
            results["item"][headline] = {"company": entry["company"], "messageUrl": entry["messageUrl"], "published": entry["published"], "headline": headline}
            print(entry['market'])
            print(f"Market: {market}/nDate: {published}/n")
            
with open("C:/Users/apskaita3/Finansų analizės ir valdymo sprendimai, UAB/Rokas Toomsalu - Power BI analitika/Integracijos/1_Public comapnies analytics/Databasesets/Others/market_news_helsinki.json", "w") as outfile:    
    json_object = json.dumps({"item": list(results["item"].values())}, indent = 4)
    outfile.write(json_object)
with open("C:/Users/apskaita3/Desktop/number2.txt", "w") as outfile1:    
    outfile1.write(start)  # type: ignore

问题原因分析

  1. 分页逻辑错误:API的start参数是分页起始位置,配合limit=19,每页返回19条数据。你用range(0,9800)循环,每次start={i},相当于每次只偏移1条,会重复请求大量重复数据,同时跳过绝大多数页面,直接导致漏抓数据。
  2. 去重逻辑缺陷:用headline作为去重key,存在相同标题但不同内容的新闻,会被误判为重复而跳过,丢失有效数据。
  3. 无请求频率控制:注释了time.sleep(1),高频请求可能被API限流,导致部分请求返回不完整或失败。
  4. 未处理边界情况:当start超过总数据量时,API返回空条目,代码仍会继续循环,做无效请求且无法终止。

修正后的代码

import requests
import json
import time

# 读取上次起始位置,不存在则从0开始
start = 0
try:
    with open('C:/Users/apskaita3/Desktop/number2.txt', "r") as f:
        start = int(f.readlines()[0])
except FileNotFoundError:
    pass

results = {}  # 用messageUrl作为唯一标识去重
limit = 19  # 每页固定返回19条
max_requests = 500  # 防止无限循环,可按需调整
target_market = 'Main Market, Helsinki'
min_date = "2021-10-20 06:30:00"

total_requested = 0
while total_requested < max_requests:
    # 构造正确的分页URL
    URL = f"https://api.news.eu.nasdaq.com/news/query.action?type=handleResponse&showAttachments=true&showCnsSpecific=true&showCompany=true&countResults=false&freeText=&company=&market=Main%20Market%2C+Helsinki&cnscategory=&fromDate=&toDate=&globalGroup=exchangeNotice&globalName=NordicMainMarkets&displayLanguage=en&language=en&timeZone=CET&dateMask=yyyy-MM-dd+HH%3Amm%3Ass&limit={limit}&start={start}&dir=ASC"
    
    try:
        r = requests.get(url=URL)
        r.raise_for_status()  # 检查请求是否成功
    except requests.exceptions.RequestException as e:
        print(f"请求失败: {e}")
        time.sleep(5)
        continue
    
    # 处理API的包装格式,注意去掉结尾的括号
    res = r.text.replace("handleResponse(", "").rstrip(")")
    try:
        res_json = json.loads(res)
    except json.JSONDecodeError as e:
        print(f"解析JSON失败: {e}")
        time.sleep(5)
        continue
    
    downloaded_entries = res_json["results"]["item"]
    if not downloaded_entries:
        print("无更多数据,终止循环")
        break
    
    new_count = 0
    for entry in downloaded_entries:
        if entry["market"] == target_market and entry["published"] >= min_date:
            unique_key = entry["messageUrl"]
            if unique_key not in results:
                results[unique_key] = {
                    "company": entry["company"],
                    "messageUrl": entry["messageUrl"],
                    "published": entry["published"],
                    "headline": entry["headline"].strip(),
                    "market": entry["market"]
                }
                new_count += 1
                print(f"新增数据: {entry['published']} - {entry['headline'][:50]}...")
    
    print(f"完成第{total_requested+1}次请求,新增{new_count}条数据")
    start += limit
    total_requested += 1
    time.sleep(1)  # 控制请求频率,避免限流

# 保存结果
output_path = "C:/Users/apskaita3/Finansų analizės ir valdymo sprendimai, UAB/Rokas Toomsalu - Power BI analitika/Integracijos/1_Public comapnies analytics/Databasesets/Others/market_news_helsinki.json"
with open(output_path, "w", encoding="utf-8") as outfile:
    json.dump({"item": list(results.values())}, outfile, indent=4, ensure_ascii=False)

# 保存下次起始位置
with open("C:/Users/apskaita3/Desktop/number2.txt", "w") as outfile1:
    outfile1.write(str(start))

print(f"抓取完成,共获取{len(results)}条有效数据")

关键优化点

  • 修正分页逻辑:用start += limit确保每次请求下一页完整数据,无跳过无重复
  • 优化去重逻辑:用messageUrl作为唯一标识,避免因标题重复丢失数据
  • 增加错误处理:处理请求失败、JSON解析失败的情况,避免程序崩溃
  • 控制请求频率:恢复time.sleep(1),降低被API限流的风险
  • 自动终止循环:API返回空条目时自动停止,减少无效请求

内容的提问来源于stack exchange,提问作者jimm denison

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 06:20:27