You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python编写MyFitnessPal爬虫遇认证问题,求解决登录故障

MyFitnessPal爬虫登录问题排查与解决

问题描述

基于GitHub仓库holtchesley/mfp_food_and_excercise开发Python爬虫,用于提取MyFitnessPal(MFP)的数据,但在程序完成MFP网站认证登录时遇到障碍:爬虫尝试登录后经多次重定向返回登录页面,最终在抓取完成前自动关闭。

已完成验证操作

  • 确认爬虫使用的登录凭证与手动登录完全一致
  • 调试确认日期变量及生成的URL正确,将URL粘贴至浏览器可正常跳转至目标页面

问题现象日志

1. 2023-05-18 16:32:11 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.myfitnesspal.com/account/login?email=&username=prayagpurohit1%40gmail.com&password=**> (referer: https://www.myfitnesspal.com/account/login)
2. 2023-05-18 16:32:11 [scrapy.downloadermiddlewares.redirect] DEBUG: Redirecting (301) to <GET https://www.myfitnesspal.com/reports/printable_diary?from=2023-03-24&to=2023-03-24> from <GET http://www.myfitnesspal.com/reports/printable_diary?from=2023-03-24&to=2023-03-24>
*手动在浏览器中打开第二条日志的URL(登录后)可进入正确页面
3. 2023-05-18 16:32:14 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.myfitnesspal.com/account/login?callbackUrl=https%3A%2F%2Fwww.myfitnesspal.com%2Freports%2Fprintable-diary%3Ffrom%3D2023-03-24%26to%3D2023-03-24> (referer: None)

待排查代码

import os
import scrapy
from datetime import datetime, timedelta
from dateutil import parser as dtparser


class MfpSpider(scrapy.Spider):
    name = "mfp"
    allowed_domains = ["myfitnesspal.com"]
    start_urls = (
        'https://www.myfitnesspal.com/account/login',
    )

    def __init__(self, username='', password='', start_date=None, target_date='', 
                target_file='', outdir='output', *args, **kwargs):
        super(MfpSpider, self).__init__(*args, **kwargs)
        self.username = username
        self.password = password
        self.target_date = target_date
        self.target_file = target_file
        self.outdir = outdir

        self.start_date = dtparser.parse(target_date)
        self.end_date = dtparser.parse(target_date)
        if start_date is not None:
            self.start_date = dtparser.parse(start_date)
        self.dts = self.get_dts()

    def get_dts(self):
        ndays = (self.end_date - self.start_date).days
        return [self.start_date + timedelta(i) for i in range(ndays + 1)]

    def parse(self, response):
        return scrapy.FormRequest.from_response(response,
                                                formdata={
                                                    'username': self.username,
                                                    'password': self.password},
                                                callback=self.logged_in)  # Login

    def logged_in(self, response):
        print('LOGGED IN. PROCESSING...')
        # print(response.body)
        for dt in self.dts:
            dtstr = dt.strftime('%Y-%m-%d')
            yield scrapy.Request(url="http://www.myfitnesspal.com/reports/printable_diary?from=" + dtstr + "&to=" + dtstr,
                                callback=self.dump_tables)  # Go to this URL after logging in which gives a diary for each date

    def write_table(self, rows, filename):
        with open(filename, "w") as f:
            f.write('\n'.join(rows))

    def get_outfile(self, name):
        return os.path.join(self.outdir, name + '.csv')

    def read_tables(self, response, name, delimiter='\t'):
        frmt = lambda x: delimiter.join(x).encode('UTF-8')  # + u"\n".encode('UTF-8')
        outputs = []
        for table in response.xpath("//table[@id='" + name + "']"):
            for row in table.xpath(".//tr"):
                contents = row.xpath("./td/text()").extract()
                outputs.append(frmt(contents))
        return outputs

    def dump_tables(self, response):
        # get data
        # print(response.body)
        dts = response.xpath('//h2/text()').extract()
        foods = self.read_tables(response, 'food')
        exercise = self.read_tables(response, 'exercise')  # n.b. the typo
        title = dtparser.parse(dts[0]).strftime('%Y-%m-%d') if dts else ''

        # write data
        self.write_table(foods, self.get_outfile(title + "-food"))
        self.write_table(exercise, self.get_outfile(title + "-exercise"))

问题分析

  1. 表单字段缺失:scrapy.FormRequest.from_response可能未捕获MFP登录表单所需的隐藏验证字段(如CSRF Token),导致登录请求不被服务器认可。
  2. HTTP/HTTPS重定向丢失会话:logged_in方法中请求的是http://开头的URL,触发301重定向到https://时,会话Cookie可能未正确跟随,导致后续请求未携带登录状态。
  3. 请求头未模拟浏览器:缺少浏览器标识(User-Agent)、Accept等关键请求头,被MFP识别为非合法请求。
  4. 登录状态未校验:logged_in方法直接跳转目标页面,未先验证是否真的登录成功,导致无效请求触发重定向回登录页。

解决建议

1. 补全登录表单字段

先打印登录页面的所有表单字段,确认是否存在CSRF Token等隐藏字段,添加到formdata中:

def parse(self, response):
    # 打印所有表单字段,排查缺失项
    form_fields = response.xpath('//form//input/@name').extract()
    self.logger.info(f"登录表单字段:{form_fields}")
    
    # 提取CSRF Token(示例,需根据实际字段名调整)
    csrf_token = response.xpath('//input[@name="authenticity_token"]/@value').extract_first()
    
    return scrapy.FormRequest.from_response(
        response,
        formdata={
            'username': self.username,
            'password': self.password,
            'authenticity_token': csrf_token  # 添加找到的隐藏字段
        },
        callback=self.logged_in,
        dont_filter=True
    )

2. 统一使用HTTPS请求

修改logged_in方法中的URL为HTTPS协议,避免重定向导致会话丢失:

def logged_in(self, response):
    # 先验证登录状态
    if response.xpath('//*[contains(text(), "Welcome")]'):
        print('LOGGED IN. PROCESSING...')
        for dt in self.dts:
            dtstr = dt.strftime('%Y-%m-%d')
            # 使用HTTPS URL
            url = f"https://www.myfitnesspal.com/reports/printable_diary?from={dtstr}&to={dtstr}"
            yield scrapy.Request(
                url=url,
                callback=self.dump_tables,
                headers={'Referer': response.url}  # 添加Referer头
            )
    else:
        self.logger.error("登录失败,未跳转到登录后页面")

3. 添加浏览器请求头模拟

在Spider中配置默认请求头,模拟真实浏览器请求:

class MfpSpider(scrapy.Spider):
    name = "mfp"
    allowed_domains = ["myfitnesspal.com"]
    start_urls = (
        'https://www.myfitnesspal.com/account/login',
    )
    
    # 配置浏览器请求头
    custom_settings = {
        'DEFAULT_REQUEST_HEADERS': {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/113.0.0.0 Safari/537.36',
            'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
            'Accept-Language': 'en-US,en;q=0.5',
            'Connection': 'keep-alive'
        }
    }

4. 确保Cookie会话保持

检查Scrapy配置文件settings.py,确认Cookie中间件已启用:

COOKIES_ENABLED = True
DOWNLOADER_MIDDLEWARES = {
    'scrapy.downloadermiddlewares.cookies.CookiesMiddleware': 700,
}

内容的提问来源于stack exchange,提问作者prayag purohit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 18:45:02