You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用twarc2按日限制采集Twitter数据的实现方式咨询

使用twarc2按日限制采集Twitter数据的实现方式

可以一次性完成采集,无需分10次手动执行代码,有两种常用实现方式:

方式一:批量查询文件执行

  1. 创建一个文本文件(比如daily_queries.txt),每行写入对应日期的查询语句,包含时间范围和数量限制:

    your_search_query since:2023-07-01 until:2023-07-02 limit:100
    your_search_query since:2023-07-02 until:2023-07-03 limit:100
    ...
    your_search_query since:2023-07-10 until:2023-07-11 limit:100
    

    注意:until的日期要比目标结束日期晚一天(Twitter的until参数为排他性);your_search_query替换为实际搜索关键词,可按需添加过滤条件(如-is:retweet排除转发)。

  2. 执行单次twarc2命令批量处理所有查询:

    twarc2 search --queries daily_queries.txt output.jsonl
    

    所有每日100条数据会统一写入output.jsonl,也可在查询文件每行末尾追加> tweets_{date}.jsonl,让每日数据单独输出到对应文件。

方式二:单次命令配合后处理

  1. 先采集整个时间范围内的总数据,指定总限制为1000条(10天×100条):

    twarc2 search "your_search_query since:2023-07-01 until:2023-07-11" --limit 1000 full_output.jsonl
    
  2. 用脚本读取全量数据,按推文创建日期分组并保留每日最多100条,再拆分写入每日文件。示例Python代码片段:

    import json
    from collections import defaultdict
    
    daily_tweets = defaultdict(list)
    
    with open("full_output.jsonl", "r") as f:
        for line in f:
            tweet = json.loads(line)
            date = tweet["created_at"].split("T")[0]
            if len(daily_tweets[date]) < 100:
                daily_tweets[date].append(tweet)
    
    for date, tweets in daily_tweets.items():
        with open(f"tweets_{date}.jsonl", "w") as f:
            for tweet in tweets:
                json.dump(tweet, f)
                f.write("\n")
    

两种方式都能实现一次性执行完成采集,方式一更直接控制每日数量,方式二适合需要先获取全量数据再筛选的场景。

内容的提问来源于stack exchange,提问作者code_lover

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 23:33:29