You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何抓取印度经济时报2020年头条的Python代码无结果?

印度经济时报头条抓取CSV无数据的排查与解决方法

可能的原因及对应解决步骤

1. 反爬机制拦截请求

  • 直接用requests发起的请求可能被识别为爬虫,返回空页面或非目标内容。给请求添加浏览器UA头:
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
    response = requests.get(url, headers=headers)
    
  • 若网站需要Cookie验证,可从浏览器复制对应Cookie加入请求头,或用requests.Session()维持会话。

2. 日期URL构造错误

  • 手动访问你生成的每日URL,确认是否能正常打开并显示当日头条。比如印度经济时报历史存档的URL格式可能为https://economictimes.indiatimes.com/archivelist/year-2020,month-1,starttime-1200489600.cms,需确保日期转换后的参数完全匹配。
  • 检查日期循环逻辑:2020年是闰年,2月有29天,避免生成无效日期的URL导致请求失败。

3. BeautifulSoup解析逻辑错误

  • 打印response.text查看是否抓取到完整页面内容,若内容为空,优先排查反爬问题;若内容正常,调整标签选择器:
    比如原选择器可能无法定位头条,可通过浏览器开发者工具查看头条元素的标签和class,替换为soup.find_all('h2', class_='story__heading')这类精准选择器。
  • 确认解析器使用正确,推荐用html.parser或lxml,避免因解析器不兼容导致元素查找失败。

4. CSV写入逻辑问题

  • 检查是否在数据为空时就打开了CSV文件,或未调用writerow()/writerows()写入数据。
  • 确保文件打开方式正确,避免编码或换行符问题:
    with open('headlines.csv', 'w', newline='', encoding='utf-8') as f:
        writer = csv.writer(f)
        writer.writerow(['日期', '头条'])
    
  • 添加异常捕获,排查请求或解析过程中的隐性错误,避免程序无报错但无数据输出。

修正后的示例代码片段

import requests
from bs4 import BeautifulSoup
import csv
from datetime import datetime, timedelta

start_date = datetime(2020, 1, 1)
end_date = datetime(2020, 10, 31)
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}

with open('et_headlines.csv', 'w', newline='', encoding='utf-8') as csvfile:
    writer = csv.writer(csvfile)
    writer.writerow(['日期', '头条标题'])

    current_date = start_date
    while current_date <= end_date:
        # 替换为印度经济时报历史头条的真实URL格式
        date_str = current_date.strftime('%Y/%m/%d')
        url = f'https://economictimes.indiatimes.com/date/{date_str}'
        
        try:
            response = requests.get(url, headers=headers)
            response.raise_for_status()
            soup = BeautifulSoup(response.text, 'html.parser')
            
            # 替换为页面实际的头条元素选择器
            headlines = soup.find_all('h2', class_='story__heading')
            if headlines:
                for headline in headlines:
                    writer.writerow([current_date.strftime('%Y-%m-%d'), headline.get_text(strip=True)])
            else:
                print(f"{current_date.strftime('%Y-%m-%d')} 未抓取到头条")
        except Exception as e:
            print(f"{current_date.strftime('%Y-%m-%d')} 抓取失败: {str(e)}")
        
        current_date += timedelta(days=1)

内容的提问来源于stack exchange,提问作者Doubts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 12:40:20