基于日期的Python爬取求助:印尼央行汇率交易表抓取问题
Fixing Custom Date Range Scraping for BI Indonesia's Exchange Rate Table
Hey there! I see you're stuck trying to scrape the BI exchange rate transactions between 20 November 2015 and 20 November 2020. The core issue here is that this page uses AJAX POST requests (not URL parameters) to handle date filtering—so your initial GET request only pulls the default date data. Let's walk through how to fix this.
Step 1: Understand the Request Flow
First, open your browser's DevTools (F12), go to the Network tab, and filter for XHR requests. When you submit a custom date range on the page, you'll notice:
- The request is sent via POST, not GET
- It includes critical form parameters like view state values (required for ASP.NET pages to validate the request) and your custom start/end dates
Step 2: Updated Code to Handle POST Requests
Here's a complete working script that extracts the necessary validation parameters and submits your custom date range:
import requests from bs4 import BeautifulSoup import pandas as pd # Target URL url = "https://www.bi.go.id/id/moneter/informasi-kurs/transaksi-bi/Default.aspx" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/60.0.3112.101 Safari/537.36", "Referer": url } # Initialize a session to persist cookies and session data session = requests.Session() # First: GET the page to extract ASP.NET validation parameters initial_response = session.get(url, headers=headers) soup = BeautifulSoup(initial_response.content, "html.parser") # Grab required view state parameters (ASP.NET needs these to accept the POST) viewstate = soup.find("input", id="__VIEWSTATE")["value"] viewstate_generator = soup.find("input", id="__VIEWSTATEGENERATOR")["value"] event_validation = soup.find("input", id="__EVENTVALIDATION")["value"] # Your custom date range (match the page's required format: dd/mm/yyyy) start_date = "20/11/2015" end_date = "20/11/2020" # Build the form data to submit form_data = { "__VIEWSTATE": viewstate, "__VIEWSTATEGENERATOR": viewstate_generator, "__EVENTVALIDATION": event_validation, "ctl00$ContentPlaceHolder1$txtTanggalAwal": start_date, "ctl00$ContentPlaceHolder1$txtTanggalAkhir": end_date, "ctl00$ContentPlaceHolder1$btnTampilkan": "Tampilkan" # Match the submit button's name } # Second: Send the POST request with your date range post_response = session.post(url, headers=headers, data=form_data) post_soup = BeautifulSoup(post_response.content, "html.parser") # Extract the exchange rate table target_table = post_soup.find("table", id="ctl00_ContentPlaceHolder1_gvKurs") if target_table: # Convert table to DataFrame for easy handling df = pd.read_html(str(target_table))[0] print("Sample of the scraped data:") print(df.head()) # Save to CSV if needed df.to_csv("bi_exchange_rates_2015-2020.csv", index=False) else: print("Couldn't find the target table—double-check if the page structure has changed!")
Key Notes to Avoid Issues
- Use
requests.Session(): This keeps your session cookies consistent between the initial GET (to get validation params) and the POST (to submit dates) - Date Format: Make sure your dates follow
dd/mm/yyyy—that's what the BI page expects - Validation Params: ASP.NET pages require
__VIEWSTATE,__VIEWSTATEGENERATOR, and__EVENTVALIDATIONto confirm the request is legitimate. Always extract these fresh from the initial page load.
Troubleshooting Tips
- If you still get default date data: Double-check the form parameter names (they might change if the page updates). You can copy the full Form Data from DevTools to verify.
- Anti-scraping blocks: Add small delays between requests, use a more modern User-Agent, or try rotating proxies if needed.
内容的提问来源于stack exchange,提问作者adinda aulia
相关产品推荐
相关产品推荐

