You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Web Apps部署Flask RSS阅读器无法拉取Substack源排查

问题背景

我开发了一个简易Flask应用,用于从若干RSS源获取文章,核心代码如下:

import requests
import xml.etree.ElementTree as ET
from dateutil import parser
import re
from feeds_config import FEEDS

def extract_image_from_content(content):
    """Extract the first image URL from the content using regex."""
    match = re.search(r'<img[^>]+src="([^">]+)"', content)
    return match.group(1) if match else None

def fetch_articles():
    """
    Fetch articles from the feeds listed in FEEDS configuration.
    For each feed, parse the RSS feed and extract relevant article details.
    """
    articles = []
    
    for feed in FEEDS:
        response = requests.get(feed['url'])  # Fetch the RSS feed
        
        if response.status_code == 200:
            root = ET.fromstring(response.content)  # Parse XML content
            ns = feed.get('image_ns', {})  # Get namespaces for images
            content_ns = feed.get('content_ns', {})  # Get namespaces for content
            
            # List to temporarily store feed articles
            feed_articles = []
            
            # Iterate through each item (article) in the feed
            for item in root.findall(".//item"):
                title = item.find("title").text  # Extract article title
                link = item.find("link").text  # Extract article link
                pub_date = item.find("pubDate").text  # Extract publication date
                timestamp = parser.parse(pub_date)  # Parse date to a datetime object
                
                # Extract image URL from <enclosure> or other image tags if available
                image = item.find(feed.get('image_xpath', '.'), namespaces=ns)
                image_url = image.get("url") if image is not None else None
                
                # If no image found, attempt to extract it from content
                if not image_url and feed.get('content_xpath'):
                    content = item.find(feed['content_xpath'], namespaces=content_ns)
                    content_text = content.text if content is not None else ""
                    image_url = extract_image_from_content(content_text)
                
                # Append article details to feed_articles list
                feed_articles.append({
                    "title": title,
                    "link": link,
                    "timestamp": timestamp,
                    "source": feed['source'],
                    "image": image_url,
                    "source_url": feed['source_url']
                })
            
            # Remove duplicate if the first two items have the same title
            if len(feed_articles) > 1 and feed_articles[0]['title'] == feed_articles[1]['title']:
                feed_articles.pop(0)
            
            # Add the remaining articles to the main articles list
            articles.extend(feed_articles)
    
    return articles

使用的FEEDS配置:

FEEDS = [
    {
        'url': 'https://hedgehogreview.com/web-features/feed',
        'source': 'Hedgehog Review',
        'source_url': 'https://hedgehogreview.com/',
        'image_xpath': './enclosure',
        'image_ns': {},
        'content_xpath': './content:encoded',
        'content_ns': {'content': 'http://purl.org/rss/1.0/modules/content/'}
    },
    {
        'url': 'https://mcrawford.substack.com/feed',
        'source': 'M.B. Crawford Substack',
        'source_url': 'https://mcrawford.substack.com',
        'image_xpath': './enclosure',
        'image_ns': {}
    },
    {
        'url': 'https://mattdinan.substack.com/feed',
        'source': 'Matt Dinan Substack',
        'source_url': 'https://mattdinan.substack.com',
        'image_xpath': './enclosure',
        'image_ns': {}
    }
]

本地运行时所有源均可正常加载,但部署到Azure免费版App Service后,仅Hedgehog Review源能加载,Substack源无法拉取。已确认出入站流量允许、依赖已正确部署,请问原因是什么?


可能的原因及解决办法

1. Azure免费层共享IP被Substack反爬拦截

Azure免费版App Service的出站IP是多租户共享的,大量用户用这些IP进行爬取操作,很容易被Substack的反爬虫系统标记为恶意请求,直接拒绝访问。而Hedgehog Review的反爬策略相对宽松,所以能正常返回内容。

解决办法:

  • 给请求添加模拟浏览器的请求头,让请求看起来更像正常用户访问:
    修改requests.get的代码,增加headers参数:
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8'
    }
    response = requests.get(feed['url'], headers=headers)
    

2. 请求头过于简单,被Substack识别为爬虫

本地环境中,requests可能继承了系统或终端的一些默认头,而Azure App Service环境下的请求头非常简洁,缺少User-Agent等关键标识,被Substack判定为非人类请求,从而拦截。

解决办法:

  • 除了添加User-Agent,还可以补充其他常见请求头,比如Accept-Language、Referer等,进一步模拟正常浏览器请求。
  • 在代码中添加日志,输出请求的状态码和响应内容,方便排查:
    response = requests.get(feed['url'], headers=headers)
    # 记录请求状态和部分响应内容,可在Azure App Service日志中查看
    print(f"Source: {feed['source']}, Status Code: {response.status_code}")
    if response.status_code != 200:
        print(f"Response snippet: {response.text[:300]}")
    
    通过日志可以确认是否返回403(拒绝访问)、429(请求频率过高)等状态码,明确拦截类型。

3. 请求频率过高触发限制

如果应用部署后频繁请求Substack的Feed,可能触发Substack的请求频率限制,导致被临时拦截。

解决办法:

  • 在代码中添加请求间隔,降低请求频率:
    import time
    # ...
    response = requests.get(feed['url'], headers=headers)
    time.sleep(1)  # 每次请求后暂停1秒
    # ...
    

内容的提问来源于stack exchange,提问作者John Mehler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 16:05:54