You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取指定网页文章内容?代码返回No Data解决

解决网页文章内容提取失败的方案

针对你提取印度经济时报文章内容时输出“No Data”的问题,可按以下步骤排查修复:

1. 解决请求被反爬拦截问题

很多网站会通过检测User-Agent识别非浏览器请求,直接调用requests.get()会被拦截,返回空白或反爬页面,导致无法解析内容。需添加模拟浏览器的请求头:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 模拟Chrome浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

url = 'https://economictimes.indiatimes.com/industry/cons-products/food/heinz-braces-up-for-aggressive-marketing/articleshow/5417995.cms'
response = requests.get(url, headers=headers)
# 检查请求是否成功,失败则直接抛出异常
response.raise_for_status()

2. 修正文章内容的定位选择器

网页结构可能已变更,原代码的选择器无法匹配目标内容。通过浏览器开发者工具查看该页面源码,可知文章正文位于class="artText"的div标签下,包含多个p标签。更新解析逻辑:

soup = BeautifulSoup(response.text, 'html.parser')
# 定位文章内容容器
article_container = soup.find('div', class_='artText')

if article_container:
    # 提取所有有效段落并拼接成完整文本
    paragraphs = article_container.find_all('p')
    full_article = '\n\n'.join([p.get_text(strip=True) for p in paragraphs if p.get_text(strip=True)])
    # 用pandas存储或输出结果
    df = pd.DataFrame({'文章内容': [full_article]})
    print(df)
else:
    print("No Data")

3. 实用调试技巧

  • 打印response.text查看返回内容是否为正常网页,若显示反爬提示或空白,需补充更多请求头字段(如Referer、Accept-Language)。
  • 用浏览器开发者工具的“检查元素”功能,右键文章内容选择“复制CSS选择器”,确保定位节点的准确性。

内容的提问来源于stack exchange,提问作者airmon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 15:52:35