You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy脚本中配置凭证,爬取需认证的私有网站文章

Scrapy爬取需认证私有网站的凭证添加方法

常见的私有网站认证分两种场景,对应不同的凭证添加方式,直接给你改好的代码示例:


一、表单账号密码登录(网页端输入账号密码提交的场景)

这种是最常见的登录模式,需要先发送登录请求获取登录状态,再爬取内容。你需要先用浏览器开发者工具抓包,确认登录接口地址、表单字段名(比如账号字段是username还是user)。

修改后的代码:

import scrapy
from scrapy.crawler import CrawlerProcess
from scrapy.http import FormRequest

class EnergyGeneralNews(scrapy.Spider):
    name = 'energygeneralnews'

    def start_requests(self):
        # 替换为目标私有网站的登录接口地址
        login_url = "https://你的私有网站登录地址"
        # 替换为实际的账号密码,以及网站表单对应的字段名
        yield FormRequest(
            url=login_url,
            formdata={
                'username': '你的账号',
                'password': '你的密码'
            },
            callback=self.after_login
        )

    def after_login(self, response):
        # 验证登录是否成功,替换成网站登录后的特征文本(比如用户名、欢迎语)
        if "登录成功" in response.text:
            # 登录成功后,开始爬取目标分页页面
            for x in range(1, 18):
                target_url = f"https://你的私有网站列表页/Page-{x}.html"
                yield scrapy.Request(url=target_url, callback=self.parse)
        else:
            self.logger.error("登录失败,请检查账号密码或登录接口")

    def parse(self, response):
        for link in response.xpath('//*[@class="categoryArticle__content"]/a/@href'):
            yield scrapy.Request(
                url=link.get(),
                callback=self.parse_item
            )

    def parse_item(self, response):
        yield {
            'date': response.xpath('//*[@class="article_byline"]/text()[2]').re(r'\w+?\s\d\d,\s\d{4}'),
            'category': response.xpath('(//*[@itemprop="name"])[3]/text()').get(),
            'title': response.xpath('//*[@class="singleArticle__content"]/h1/text()').get(),
            'text':''.join([x.get().strip() for x in response.xpath('//*[@id="article-content"]//p//text()')])
        }

if __name__ == '__main__':
    process = CrawlerProcess()
    process.crawl(EnergyGeneralNews)
    process.start()

二、HTTP Basic认证(浏览器弹出账号密码输入框的场景)

这种是网站基于HTTP协议的基础认证,直接在请求里附带凭证即可,无需单独发送登录请求。

修改后的代码:

import scrapy
from scrapy.crawler import CrawlerProcess

class EnergyGeneralNews(scrapy.Spider):
    name = 'energygeneralnews'
    # 定义账号密码
    username = '你的账号'
    password = '你的密码'

    def start_requests(self):
        for x in range(1,18):
            target_url = f"https://你的私有网站列表页/Page-{x}.html"
            yield scrapy.Request(
                url=target_url,
                callback=self.parse,
                auth=(self.username, self.password)
            )

    def parse(self, response):
        for link in response.xpath('//*[@class="categoryArticle__content"]/a/@href'):
            yield scrapy.Request(
                url=link.get(),
                callback=self.parse_item,
                auth=(self.username, self.password)
            )

    def parse_item(self, response):
        yield {
            'date': response.xpath('//*[@class="article_byline"]/text()[2]').re(r'\w+?\s\d\d,\s\d{4}'),
            'category': response.xpath('(//*[@itemprop="name"])[3]/text()').get(),
            'title': response.xpath('//*[@class="singleArticle__content"]/h1/text()').get(),
            'text':''.join([x.get().strip() for x in response.xpath('//*[@id="article-content"]//p//text()')])
        }

if __name__ == '__main__':
    process = CrawlerProcess()
    process.crawl(EnergyGeneralNews)
    process.start()

也可以通过Scrapy配置文件统一设置:在settings.py里添加以下内容,所有请求会自动携带认证:

DOWNLOADER_MIDDLEWARES = {
    'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware': 400,
}
HTTP_AUTH_USERNAME = '你的账号'
HTTP_AUTH_PASSWORD = '你的密码'

内容的提问来源于stack exchange,提问作者yangyang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 13:06:24