Azure Web Apps部署Flask RSS阅读器无法拉取Substack源排查
问题背景
我开发了一个简易Flask应用,用于从若干RSS源获取文章,核心代码如下:
import requests import xml.etree.ElementTree as ET from dateutil import parser import re from feeds_config import FEEDS def extract_image_from_content(content): """Extract the first image URL from the content using regex.""" match = re.search(r'<img[^>]+src="([^">]+)"', content) return match.group(1) if match else None def fetch_articles(): """ Fetch articles from the feeds listed in FEEDS configuration. For each feed, parse the RSS feed and extract relevant article details. """ articles = [] for feed in FEEDS: response = requests.get(feed['url']) # Fetch the RSS feed if response.status_code == 200: root = ET.fromstring(response.content) # Parse XML content ns = feed.get('image_ns', {}) # Get namespaces for images content_ns = feed.get('content_ns', {}) # Get namespaces for content # List to temporarily store feed articles feed_articles = [] # Iterate through each item (article) in the feed for item in root.findall(".//item"): title = item.find("title").text # Extract article title link = item.find("link").text # Extract article link pub_date = item.find("pubDate").text # Extract publication date timestamp = parser.parse(pub_date) # Parse date to a datetime object # Extract image URL from <enclosure> or other image tags if available image = item.find(feed.get('image_xpath', '.'), namespaces=ns) image_url = image.get("url") if image is not None else None # If no image found, attempt to extract it from content if not image_url and feed.get('content_xpath'): content = item.find(feed['content_xpath'], namespaces=content_ns) content_text = content.text if content is not None else "" image_url = extract_image_from_content(content_text) # Append article details to feed_articles list feed_articles.append({ "title": title, "link": link, "timestamp": timestamp, "source": feed['source'], "image": image_url, "source_url": feed['source_url'] }) # Remove duplicate if the first two items have the same title if len(feed_articles) > 1 and feed_articles[0]['title'] == feed_articles[1]['title']: feed_articles.pop(0) # Add the remaining articles to the main articles list articles.extend(feed_articles) return articles
使用的FEEDS配置:
FEEDS = [ { 'url': 'https://hedgehogreview.com/web-features/feed', 'source': 'Hedgehog Review', 'source_url': 'https://hedgehogreview.com/', 'image_xpath': './enclosure', 'image_ns': {}, 'content_xpath': './content:encoded', 'content_ns': {'content': 'http://purl.org/rss/1.0/modules/content/'} }, { 'url': 'https://mcrawford.substack.com/feed', 'source': 'M.B. Crawford Substack', 'source_url': 'https://mcrawford.substack.com', 'image_xpath': './enclosure', 'image_ns': {} }, { 'url': 'https://mattdinan.substack.com/feed', 'source': 'Matt Dinan Substack', 'source_url': 'https://mattdinan.substack.com', 'image_xpath': './enclosure', 'image_ns': {} } ]
本地运行时所有源均可正常加载,但部署到Azure免费版App Service后,仅Hedgehog Review源能加载,Substack源无法拉取。已确认出入站流量允许、依赖已正确部署,请问原因是什么?
可能的原因及解决办法
1. Azure免费层共享IP被Substack反爬拦截
Azure免费版App Service的出站IP是多租户共享的,大量用户用这些IP进行爬取操作,很容易被Substack的反爬虫系统标记为恶意请求,直接拒绝访问。而Hedgehog Review的反爬策略相对宽松,所以能正常返回内容。
解决办法:
- 给请求添加模拟浏览器的请求头,让请求看起来更像正常用户访问:
修改requests.get的代码,增加headers参数:headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8' } response = requests.get(feed['url'], headers=headers)
2. 请求头过于简单,被Substack识别为爬虫
本地环境中,requests可能继承了系统或终端的一些默认头,而Azure App Service环境下的请求头非常简洁,缺少User-Agent等关键标识,被Substack判定为非人类请求,从而拦截。
解决办法:
- 除了添加
User-Agent,还可以补充其他常见请求头,比如Accept-Language、Referer等,进一步模拟正常浏览器请求。 - 在代码中添加日志,输出请求的状态码和响应内容,方便排查:
通过日志可以确认是否返回403(拒绝访问)、429(请求频率过高)等状态码,明确拦截类型。response = requests.get(feed['url'], headers=headers) # 记录请求状态和部分响应内容,可在Azure App Service日志中查看 print(f"Source: {feed['source']}, Status Code: {response.status_code}") if response.status_code != 200: print(f"Response snippet: {response.text[:300]}")
3. 请求频率过高触发限制
如果应用部署后频繁请求Substack的Feed,可能触发Substack的请求频率限制,导致被临时拦截。
解决办法:
- 在代码中添加请求间隔,降低请求频率:
import time # ... response = requests.get(feed['url'], headers=headers) time.sleep(1) # 每次请求后暂停1秒 # ...
内容的提问来源于stack exchange,提问作者John Mehler
相关产品推荐
相关产品推荐

