使用_dopostback方法抓取ASPX站点数据——PowerShell求助
Got it, let's walk through how to scrape all paginated data from that ASPX page using PowerShell. ASP.NET postback pages can be tricky because they rely on maintaining session state and specific hidden form variables, but we can handle that with a persistent web session and careful parsing of the page's HTML.
1. Set Up a Persistent Web Session
First, we need to keep track of cookies, viewstate, and other session data between requests. PowerShell's Invoke-WebRequest lets us do this with a session variable.
2. Fetch the Initial Page & Extract Critical Form Fields
ASPX postbacks require hidden fields like __VIEWSTATE and __EVENTVALIDATION to validate requests. We'll extract these from the initial page load, along with identifying the pagination control's target ID.
3. Loop Through All Pages
We'll either loop through all visible page numbers or keep clicking "Next" until there are no more pages. Each iteration will send a postback request with the correct page number as the event argument.
4. Parse & Collect Table Data
After each page load (initial or postback), we'll extract the table rows and add them to a collection for later export (like CSV).
# Get user credentials for the site $login = Get-Credential -Message "Enter your site credentials: " $baseUrl = "http://www.xxxxxxxx.com/index.aspx" # Initialize a persistent web session to maintain cookies and state $initialResponse = Invoke-WebRequest -Uri $baseUrl -Credential $login -SessionVariable webSession # Extract required hidden form fields (adjust regex if your page uses different IDs) $viewState = [regex]::Match($initialResponse.Content, 'id="__VIEWSTATE" value="([^"]+)"').Groups[1].Value $eventValidation = [regex]::Match($initialResponse.Content, 'id="__EVENTVALIDATION" value="([^"]+)"').Groups[1].Value # Get all page numbers from the pagination links (adjust XPath to match your page's pagination elements) $pageLinks = Select-Xml -Content $initialResponse.Content -XPath "//a[contains(@href, '__doPostBack')]" | ForEach-Object { $_.Node.href } $pageNumbers = $pageLinks | ForEach-Object { [regex]::Match($_, "__doPostBack\('([^']+)','(\d+)'\)").Groups[2].Value } | Sort-Object -Unique $maxPage = [int]$pageNumbers[-1] # Initialize array to store all scraped data $allScrapedData = @() # Loop through each page foreach ($page in 1..$maxPage) { Write-Host "Scraping page $page of $maxPage..." if ($page -eq 1) { # Use the initial response for the first page $currentResponse = $initialResponse } else { # Construct the postback form data (replace "GridView1" with your actual event target ID) $formData = @{ "__VIEWSTATE" = $viewState "__EVENTVALIDATION" = $eventValidation "__EVENTTARGET" = "GridView1" # Found in __doPostBack calls (e.g., __doPostBack('GridView1','2')) "__EVENTARGUMENT" = $page.ToString() # Add any other hidden form fields your page requires (check HTML for inputs like __VIEWSTATEGENERATOR) } # Send the postback request to load the next page $currentResponse = Invoke-WebRequest -Uri $baseUrl -Method Post -WebSession $webSession -Body $formData -Credential $login # Update viewstate and event validation for the next request $viewState = [regex]::Match($currentResponse.Content, 'id="__VIEWSTATE" value="([^"]+)"').Groups[1].Value $eventValidation = [regex]::Match($currentResponse.Content, 'id="__EVENTVALIDATION" value="([^"]+)"').Groups[1].Value } # Extract table rows (adjust XPath to match your table's ID and structure) $tableRows = Select-Xml -Content $currentResponse.Content -XPath "//table[@id='GridView1']//tr[not(th)]" # Skip header rows foreach ($row in $tableRows) { $cells = $row.Node.SelectNodes("td") # Create a custom object for each row (add/remove columns to match your table) $dataRow = [PSCustomObject]@{ Column1 = $cells[0].InnerText.Trim() Column2 = $cells[1].InnerText.Trim() Column3 = $cells[2].InnerText.Trim() } $allScrapedData += $dataRow } } # Export all data to a CSV file $allScrapedData | Export-Csv -Path "ASPX_Scraped_Data.csv" -NoTypeInformation -Encoding UTF8 Write-Host "Scraping complete! Data saved to ASPX_Scraped_Data.csv"
- Event Target ID: Replace
GridView1with the actual target from your page's__doPostBackcalls. Inspect a pagination link in your browser's dev tools to find this (e.g.,__doPostBack('ctl00$ContentPlaceHolder1$myGrid','3')usesctl00$ContentPlaceHolder1$myGridas the target). - Table XPath: Adjust the XPath for
//table[@id='GridView1']to match your table's actual ID. - Pagination Logic: If your page only shows "Previous/Next" buttons instead of all page numbers, modify the loop to check for the presence of a "Next" link and stop when it's no longer available.
- Additional Form Fields: Some pages require extra hidden fields like
__VIEWSTATEGENERATOR—check your page's HTML source and add them to theformDatahashtable.
内容的提问来源于stack exchange,提问作者NoahGrg

