You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理Twitter API JSON调用tldextract提取域名报错如何解决

问题根因
  • 第一类报错The JSON object must be str, bytes or bytearray, not list:多是部分待解析的行是空行、内容不合法,或是误将已经解析完成的列表类型传给了json.loads方法导致
  • 第二类TypeError: expected string or bytes-like object:你遍历了url对象的所有键值对,其中indices对应的值是数字列表类型,不是URL字符串,直接传入tldextract.extract就会触发类型错误,不需要遍历所有键,只需要取expanded_url或者url字段的值提取域名即可
修正后的实现代码
import json
import tldextract

# 存储所有提取到的域名,若只需单条推文的域名可将该变量放在行循环内部
all_domains = []

with open(f"{filepath}decompressed_twitter_lot1file1.txt", 'r', encoding='utf-8') as fh:
    for line in fh:
        # 跳过空行避免解析错误
        line = line.strip()
        if not line:
            continue
        # 捕获解析异常,跳过不符合格式的行
        try:
            tweet_obj = json.loads(line)
        except json.JSONDecodeError:
            continue
        urls_in_tweet = tweet_obj.get('entities', {}).get('urls', [])
        domains_in_tweet = []
        for url_item in urls_in_tweet:
            # 优先取完整的扩展链接,无扩展链接则取短链
            target_url = url_item.get('expanded_url') or url_item.get('url')
            if not target_url:
                continue
            # 捕获异常避免特殊URL导致程序中断
            try:
                domain = tldextract.extract(target_url).registered_domain
                if domain:
                    domains_in_tweet.append(domain)
                    all_domains.append(domain)
            except:
                continue
        # 可打印单条推文提取到的域名
        # print(domains_in_tweet)

# 打印所有提取到的域名
print(all_domains)
优化说明
  • 替换了保留字object作为变量名,避免命名冲突
  • 增加空行判断、JSON解析异常捕获,解决第一类JSON解析报错
  • 不再遍历URL对象的所有键值对,直接取需要的URL字段,避免传入非字符串类型触发类型错误
  • 增加域名提取异常捕获,兼容特殊格式的URL
  • 区分单条推文域名列表和全局域名列表,可根据需求取用

内容的提问来源于stack exchange,提问作者Trishala Suryavanshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 22:36:00