You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的requests与BeautifulSoup爬取RSS新闻数据提取异常求助

解决印度时报RSS爬取的描述和图片提取问题

我明白你遇到的麻烦了——这个RSS的<description>标签里嵌套了HTML结构,直接取text肯定会带冗余代码,img标签也没法直接从外层soup拿到。咱们可以通过二次解析description的内容来解决这个问题,具体调整如下:

修改后的完整代码

import requests
from bs4 import BeautifulSoup

def timesofindiaNews():
    URL = 'https://timesofindia.indiatimes.com/rssfeeds_us/72258322.cms'
    page = requests.get(URL)
    # 用xml解析整个RSS feed
    soup = BeautifulSoup(page.content, features='xml')
    news_elems = soup.find_all('item')
    news = []
    
    for news_elem in news_elems:
        title = news_elem.title.text
        news_link = news_elem.link.text
        
        # 关键:把description里的HTML内容单独拿出来二次解析
        desc_content = news_elem.description.text
        desc_soup = BeautifulSoup(desc_content, 'html.parser')
        
        # 提取纯文本描述(自动去掉所有HTML标签)
        news_description = desc_soup.get_text(strip=True)
        
        # 提取图片链接:检查是否有img标签,有的话取src属性,否则设为None
        image_tag = desc_soup.find('img')
        image = image_tag['src'] if image_tag else None
        
        # 把整理好的信息加入列表
        news.append({
            'title': title,
            'description': news_description,
            'image': image,
            'link': news_link
        })
    
    return news

# 测试调用
if __name__ == '__main__':
    news_list = timesofindiaNews()
    for item in news_list[:3]:  # 打印前3条看效果
        print(f"标题: {item['title']}")
        print(f"描述: {item['description']}")
        print(f"图片链接: {item['image']}")
        print(f"新闻链接: {item['link']}\n")

关键调整说明

  • 二次解析description:因为<description>里的内容是HTML格式,所以我们用html.parser再解析一次,这样就能像操作普通HTML页面一样提取内部元素。
  • 纯文本提取:用get_text(strip=True)可以自动去除所有HTML标签,同时清理多余的空格和换行。
  • 图片链接提取:先检查是否存在<img>标签,避免找不到标签时报错,存在的话直接取src属性值。

这样修改后,你就能得到干净的纯文本描述和正确的图片链接了,不会再出现HTML代码和null的image字段~

内容的提问来源于stack exchange,提问作者Mehul Dhariyaparmar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 08:58:13