You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy模拟ASP __doPostBack翻页失败,__EVENTTARGET未覆盖求助

Fixing Scrapy __doPostBack Pagination by Removing Unwanted Form Field

Hey there! Let’s work through this frustrating issue you’re hitting with Scrapy 1.5.0 and that stubborn ctl00$ContentBody$chkAll field. I’ve dealt with similar ASP.NET form quirks before, so here’s what you can do to fix it:

The Root Cause

When using FormRequest.from_response, Scrapy automatically pulls all fields from the page’s form—including that checkbox field you don’t want. Setting it to an empty string doesn’t work because ASP.NET still recognizes the field exists, and its presence is messing with the server’s handling of your __doPostBack request. The fix is to completely remove the field from your form data, not just blank it out.

Solution 1: Build the FormRequest Manually (Most Reliable)

Skip from_response entirely and construct your form data from scratch. This gives you full control over exactly which fields get sent:

def parse(self, response):
    # Extract critical ASP.NET hidden fields from the current page
    viewstate = response.css('input#__VIEWSTATE::attr(value)').get()
    viewstate_generator = response.css('input#__VIEWSTATEGENERATOR::attr(value)').get()
    event_validation = response.css('input#__EVENTVALIDATION::attr(value)').get()  # Include if present

    # Build your form data with ONLY the fields you need
    formdata = {
        '__VIEWSTATE': viewstate,
        '__VIEWSTATEGENERATOR': viewstate_generator,
        '__EVENTVALIDATION': event_validation,
        '__EVENTTARGET': 'ctl00$ContentBody$gvResults',  # Your pagination control's ID
        '__EVENTARGUMENT': 'Page$2',  # Target page number (adjust as needed)
        # Add any other required fields (like session-related ones) here—leave out chkAll entirely
    }

    yield scrapy.FormRequest(
        url=response.url,
        formdata=formdata,
        callback=self.parse_next_page
    )

This way, ctl00$ContentBody$chkAll never makes it into the request, so the server can’t use it to interfere with your pagination.

Solution 2: Modify the Request from from_response

If you prefer to keep using from_response, you can generate the request first, then delete the unwanted field before yielding it:

def parse(self, response):
    # Generate the initial form request with your __doPostBack parameters
    form_request = scrapy.FormRequest.from_response(
        response,
        formdata={
            '__EVENTTARGET': 'ctl00$ContentBody$gvResults',
            '__EVENTARGUMENT': 'Page$2'
        },
        dont_click=True,  # Important: don't auto-click any buttons
        callback=self.parse_next_page
    )

    # Remove the chkAll field if it exists in the form data
    unwanted_field = 'ctl00$ContentBody$chkAll'
    if unwanted_field in form_request.formdata:
        del form_request.formdata[unwanted_field]

    yield form_request

Key Notes:

  • Always extract fresh __VIEWSTATE, __VIEWSTATEGENERATOR, and __EVENTVALIDATION values from each response—these change with every page load, so reusing old values will break your requests.
  • Double-check the __EVENTTARGET value matches exactly what Fiddler shows in your manual pagination requests (case and spelling matter for ASP.NET controls).

内容的提问来源于stack exchange,提问作者reyman64

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:36:27