You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从目标网站中提取有效RSS链接

提取网站RSS链接的优化方案

一、整合标准RSS元素查找

网页里的RSS订阅链接通常会在<link>标签中以type="application/rss+xml"或type="application/atom+xml"声明,这是更规范的提取方式,直接把这段逻辑加到你的现有代码里即可:

# 提取页面头部声明的标准RSS/Atom链接
rss_meta_links = soup.find_all('link', type=['application/rss+xml', 'application/atom+xml'])
rss_meta_links = [link.get('href') for link in rss_meta_links if link.get('href')]

将这部分和你原有的页面链接筛选逻辑结合,就能同时覆盖两种来源的RSS链接:页面头部的订阅声明、页面中的可见链接。

二、验证RSS链接有效性

提取到候选链接后,需要验证它是否能正常访问且返回合法的RSS内容,避免无效链接:

def is_valid_rss(base_url, url):
    try:
        # 处理相对链接,拼接成完整URL
        if not url.startswith(('http://', 'https://')):
            url = requests.compat.urljoin(base_url, url)
        response = requests.get(url, timeout=10)
        content_type = response.headers.get('Content-Type', '')
        # 检查状态码和内容类型
        if response.status_code == 200 and ('application/rss+xml' in content_type or 'application/atom+xml' in content_type):
            return True, url
        return False, url
    except Exception as e:
        return False, url

调用这个函数就能过滤掉不可访问或非RSS格式的链接。

三、首页爬取是否足够?

  • 如果只需要核心频道的RSS链接:爬首页基本足够,正规媒体通常会把主要频道的订阅链接放在首页头部meta或页脚区域。
  • 如果需要全频道的RSS链接:仅爬首页不够,部分细分频道的RSS可能只在对应频道页存在,这时需要先从首页提取频道入口链接,再对频道页重复RSS提取逻辑。

四、完整优化代码

结合上述逻辑,优化后的完整代码如下:

import requests
from bs4 import BeautifulSoup

website_links = ["https://www.diepresse.com/", 
"https://www.sueddeutsche.de/", 
"https://www.berliner-zeitung.de/", 
"https://www.aargauerzeitung.ch/", 
"https://www.luzernerzeitung.ch/", 
"https://www.nzz.ch/",
"https://www.spiegel.de/", 
"https://www.blick.ch/",
"https://www.berliner-zeitung.de/", 
"https://www.ostsee-zeitung.de/", 
"https://www.kleinezeitung.at/", 
"https://www.blick.ch/", 
"https://www.ksta.de/", 
"https://www.tagblatt.ch/", 
"https://www.srf.ch/", 
"https://www.derstandard.at/"]

def is_valid_rss(base_url, url):
    try:
        if not url.startswith(('http://', 'https://')):
            url = requests.compat.urljoin(base_url, url)
        response = requests.get(url, timeout=10)
        content_type = response.headers.get('Content-Type', '')
        if response.status_code == 200 and ('application/rss+xml' in content_type or 'application/atom+xml' in content_type):
            return True, url
        return False, url
    except Exception as e:
        return False, url

for base_url in website_links:
    print(f"=== 处理网站: {base_url} ===")
    try:
        # 伪装请求头避免被拦截
        headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'}
        response = requests.get(base_url, headers=headers, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.content, 'html.parser')
        
        # 提取头部meta声明的RSS链接
        rss_meta_links = soup.find_all('link', type=['application/rss+xml', 'application/atom+xml'])
        rss_meta_links = [link.get('href') for link in rss_meta_links if link.get('href')]
        
        # 提取页面中含RSS特征的链接
        all_page_links = [link.get('href') for link in soup.select("a[href]") if link.get('href')]
        rss_page_links = [l for l in all_page_links if any(keyword in l.lower() for keyword in ['rss', '.rss', 'rss.xml', '/rss/'])]
        
        # 合并并去重候选链接
        all_candidate_rss = list(set(rss_meta_links + rss_page_links))
        
        # 验证有效性并输出
        valid_rss = []
        for candidate in all_candidate_rss:
            is_valid, full_url = is_valid_rss(base_url, candidate)
            if is_valid:
                valid_rss.append(full_url)
        
        print(f"找到有效RSS链接: {len(valid_rss)}个")
        for rss in valid_rss:
            print(rss)
        print()
        
    except Exception as e:
        print(f"访问网站失败: {str(e)}")
        print()

五、额外优化建议

  • 去重处理:用set对候选链接去重,避免同一RSS链接被多次提取
  • UA伪装:添加请求头模拟浏览器访问,降低被反爬拦截的概率
  • 异步请求:如果网站数量较多,改用aiohttp做异步请求,能大幅提升爬取效率

内容的提问来源于stack exchange,提问作者taga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 01:05:27