Python编写MyFitnessPal爬虫遇认证问题,求解决登录故障
MyFitnessPal爬虫登录问题排查与解决
问题描述
基于GitHub仓库holtchesley/mfp_food_and_excercise开发Python爬虫,用于提取MyFitnessPal(MFP)的数据,但在程序完成MFP网站认证登录时遇到障碍:爬虫尝试登录后经多次重定向返回登录页面,最终在抓取完成前自动关闭。
已完成验证操作
- 确认爬虫使用的登录凭证与手动登录完全一致
- 调试确认日期变量及生成的URL正确,将URL粘贴至浏览器可正常跳转至目标页面
问题现象日志
1. 2023-05-18 16:32:11 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.myfitnesspal.com/account/login?email=&username=prayagpurohit1%40gmail.com&password=**> (referer: https://www.myfitnesspal.com/account/login) 2. 2023-05-18 16:32:11 [scrapy.downloadermiddlewares.redirect] DEBUG: Redirecting (301) to <GET https://www.myfitnesspal.com/reports/printable_diary?from=2023-03-24&to=2023-03-24> from <GET http://www.myfitnesspal.com/reports/printable_diary?from=2023-03-24&to=2023-03-24> *手动在浏览器中打开第二条日志的URL(登录后)可进入正确页面 3. 2023-05-18 16:32:14 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.myfitnesspal.com/account/login?callbackUrl=https%3A%2F%2Fwww.myfitnesspal.com%2Freports%2Fprintable-diary%3Ffrom%3D2023-03-24%26to%3D2023-03-24> (referer: None)
待排查代码
import os import scrapy from datetime import datetime, timedelta from dateutil import parser as dtparser class MfpSpider(scrapy.Spider): name = "mfp" allowed_domains = ["myfitnesspal.com"] start_urls = ( 'https://www.myfitnesspal.com/account/login', ) def __init__(self, username='', password='', start_date=None, target_date='', target_file='', outdir='output', *args, **kwargs): super(MfpSpider, self).__init__(*args, **kwargs) self.username = username self.password = password self.target_date = target_date self.target_file = target_file self.outdir = outdir self.start_date = dtparser.parse(target_date) self.end_date = dtparser.parse(target_date) if start_date is not None: self.start_date = dtparser.parse(start_date) self.dts = self.get_dts() def get_dts(self): ndays = (self.end_date - self.start_date).days return [self.start_date + timedelta(i) for i in range(ndays + 1)] def parse(self, response): return scrapy.FormRequest.from_response(response, formdata={ 'username': self.username, 'password': self.password}, callback=self.logged_in) # Login def logged_in(self, response): print('LOGGED IN. PROCESSING...') # print(response.body) for dt in self.dts: dtstr = dt.strftime('%Y-%m-%d') yield scrapy.Request(url="http://www.myfitnesspal.com/reports/printable_diary?from=" + dtstr + "&to=" + dtstr, callback=self.dump_tables) # Go to this URL after logging in which gives a diary for each date def write_table(self, rows, filename): with open(filename, "w") as f: f.write('\n'.join(rows)) def get_outfile(self, name): return os.path.join(self.outdir, name + '.csv') def read_tables(self, response, name, delimiter='\t'): frmt = lambda x: delimiter.join(x).encode('UTF-8') # + u"\n".encode('UTF-8') outputs = [] for table in response.xpath("//table[@id='" + name + "']"): for row in table.xpath(".//tr"): contents = row.xpath("./td/text()").extract() outputs.append(frmt(contents)) return outputs def dump_tables(self, response): # get data # print(response.body) dts = response.xpath('//h2/text()').extract() foods = self.read_tables(response, 'food') exercise = self.read_tables(response, 'exercise') # n.b. the typo title = dtparser.parse(dts[0]).strftime('%Y-%m-%d') if dts else '' # write data self.write_table(foods, self.get_outfile(title + "-food")) self.write_table(exercise, self.get_outfile(title + "-exercise"))
问题分析
- 表单字段缺失:
scrapy.FormRequest.from_response可能未捕获MFP登录表单所需的隐藏验证字段(如CSRF Token),导致登录请求不被服务器认可。 - HTTP/HTTPS重定向丢失会话:
logged_in方法中请求的是http://开头的URL,触发301重定向到https://时,会话Cookie可能未正确跟随,导致后续请求未携带登录状态。 - 请求头未模拟浏览器:缺少浏览器标识(User-Agent)、Accept等关键请求头,被MFP识别为非合法请求。
- 登录状态未校验:
logged_in方法直接跳转目标页面,未先验证是否真的登录成功,导致无效请求触发重定向回登录页。
解决建议
1. 补全登录表单字段
先打印登录页面的所有表单字段,确认是否存在CSRF Token等隐藏字段,添加到formdata中:
def parse(self, response): # 打印所有表单字段,排查缺失项 form_fields = response.xpath('//form//input/@name').extract() self.logger.info(f"登录表单字段:{form_fields}") # 提取CSRF Token(示例,需根据实际字段名调整) csrf_token = response.xpath('//input[@name="authenticity_token"]/@value').extract_first() return scrapy.FormRequest.from_response( response, formdata={ 'username': self.username, 'password': self.password, 'authenticity_token': csrf_token # 添加找到的隐藏字段 }, callback=self.logged_in, dont_filter=True )
2. 统一使用HTTPS请求
修改logged_in方法中的URL为HTTPS协议,避免重定向导致会话丢失:
def logged_in(self, response): # 先验证登录状态 if response.xpath('//*[contains(text(), "Welcome")]'): print('LOGGED IN. PROCESSING...') for dt in self.dts: dtstr = dt.strftime('%Y-%m-%d') # 使用HTTPS URL url = f"https://www.myfitnesspal.com/reports/printable_diary?from={dtstr}&to={dtstr}" yield scrapy.Request( url=url, callback=self.dump_tables, headers={'Referer': response.url} # 添加Referer头 ) else: self.logger.error("登录失败,未跳转到登录后页面")
3. 添加浏览器请求头模拟
在Spider中配置默认请求头,模拟真实浏览器请求:
class MfpSpider(scrapy.Spider): name = "mfp" allowed_domains = ["myfitnesspal.com"] start_urls = ( 'https://www.myfitnesspal.com/account/login', ) # 配置浏览器请求头 custom_settings = { 'DEFAULT_REQUEST_HEADERS': { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/113.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Connection': 'keep-alive' } }
4. 确保Cookie会话保持
检查Scrapy配置文件settings.py,确认Cookie中间件已启用:
COOKIES_ENABLED = True DOWNLOADER_MIDDLEWARES = { 'scrapy.downloadermiddlewares.cookies.CookiesMiddleware': 700, }
内容的提问来源于stack exchange,提问作者prayag purohit
相关产品推荐
相关产品推荐

