使用Python的requests与BeautifulSoup爬取RSS新闻数据提取异常求助
解决印度时报RSS爬取的描述和图片提取问题
我明白你遇到的麻烦了——这个RSS的<description>标签里嵌套了HTML结构,直接取text肯定会带冗余代码,img标签也没法直接从外层soup拿到。咱们可以通过二次解析description的内容来解决这个问题,具体调整如下:
修改后的完整代码
import requests from bs4 import BeautifulSoup def timesofindiaNews(): URL = 'https://timesofindia.indiatimes.com/rssfeeds_us/72258322.cms' page = requests.get(URL) # 用xml解析整个RSS feed soup = BeautifulSoup(page.content, features='xml') news_elems = soup.find_all('item') news = [] for news_elem in news_elems: title = news_elem.title.text news_link = news_elem.link.text # 关键:把description里的HTML内容单独拿出来二次解析 desc_content = news_elem.description.text desc_soup = BeautifulSoup(desc_content, 'html.parser') # 提取纯文本描述(自动去掉所有HTML标签) news_description = desc_soup.get_text(strip=True) # 提取图片链接:检查是否有img标签,有的话取src属性,否则设为None image_tag = desc_soup.find('img') image = image_tag['src'] if image_tag else None # 把整理好的信息加入列表 news.append({ 'title': title, 'description': news_description, 'image': image, 'link': news_link }) return news # 测试调用 if __name__ == '__main__': news_list = timesofindiaNews() for item in news_list[:3]: # 打印前3条看效果 print(f"标题: {item['title']}") print(f"描述: {item['description']}") print(f"图片链接: {item['image']}") print(f"新闻链接: {item['link']}\n")
关键调整说明
- 二次解析description:因为
<description>里的内容是HTML格式,所以我们用html.parser再解析一次,这样就能像操作普通HTML页面一样提取内部元素。 - 纯文本提取:用
get_text(strip=True)可以自动去除所有HTML标签,同时清理多余的空格和换行。 - 图片链接提取:先检查是否存在
<img>标签,避免找不到标签时报错,存在的话直接取src属性值。
这样修改后,你就能得到干净的纯文本描述和正确的图片链接了,不会再出现HTML代码和null的image字段~
内容的提问来源于stack exchange,提问作者Mehul Dhariyaparmar
相关产品推荐
相关产品推荐

