替换URL页码并遍历多页面:线上超市订单历史爬取脚本实现求助
Hey Simon, I see you're working on scraping Jumbo's order history pages and having trouble generating URLs for each page from 1 to your PageNo (13). Let's get this sorted out for you—you already have the foundation with regex to extract the page number, now we just need to adapt that to replace the number dynamically as you loop through pages.
Method 1: Simple Regex Replacement (Quick & Straightforward)
Since you're already comfortable with regex, this approach builds on what you're familiar with. We'll compile a regex pattern to target just the page number value, then replace it with each number in your 1–13 range.
import re # Your base URL with the initial page number base_url = "https://www.jumbo.com/INTERSHOP/web/WFS/Jumbo-Grocery-Site/nl_NL/-/EUR/ViewOrderHistory-Paging?StatusType=Complete&PageNumber=1&PageableID=5DIKY7MVn2IAAAF5QrIGubww" max_page = 13 # Your PageNo variable # Regex pattern to match the page number between the fixed query segments page_pattern = re.compile(r"(?<=Complete&PageNumber=)\d+(?=&PageableID=)") # Loop through each page from 1 to max_page for page_num in range(1, max_page + 1): # Replace the page number in the URL current_page_url = page_pattern.sub(str(page_num), base_url) print(f"Processing page {page_num}: {current_page_url}") # Add your scraping logic here—e.g., send a request to current_page_url, parse the response
How this works:
- The regex uses positive lookbehind (
(?<=Complete&PageNumber=)) and positive lookahead ((?=&PageableID=)) to isolate only the numeric page value, so we don't accidentally modify other parts of the URL. re.sub()replaces that numeric value with the current page number in our loop.
Method 2: URL Query Parameter Manipulation (More Robust)
If you want a solution that's less dependent on the exact string structure of the URL (in case Jumbo ever reorders query parameters), using Python's built-in urllib.parse module is a better approach. This lets you explicitly modify the PageNumber parameter without string hacks.
from urllib.parse import urlparse, parse_qs, urlencode, urlunparse base_url = "https://www.jumbo.com/INTERSHOP/web/WFS/Jumbo-Grocery-Site/nl_NL/-/EUR/ViewOrderHistory-Paging?StatusType=Complete&PageNumber=1&PageableID=5DIKY7MVn2IAAAF5QrIGubww" max_page = 13 # Parse the URL into its components parsed_url = urlparse(base_url) # Convert HTML entity & to actual & so we can parse query parameters correctly query_params = parse_qs(parsed_url.query.replace("&", "&")) for page_num in range(1, max_page + 1): # Update the PageNumber parameter with the current page query_params["PageNumber"] = [str(page_num)] # Encode the query parameters back, then convert & back to & to match the original URL format updated_query = urlencode(query_params, doseq=True).replace("&", "&") # Reassemble the full URL current_page_url = urlunparse(parsed_url._replace(query=updated_query)) print(f"Processing page {page_num}: {current_page_url}") # Add your scraping logic here
Why this is better:
- It treats the URL's query parameters as a structured dictionary, so you don't have to worry about where
PageNumberappears in the parameter list. - It handles URL encoding properly, which prevents issues if parameters ever contain special characters.
Which method should you use?
- Use the regex method if you're confident the URL structure won't change anytime soon—it's quick and easy to implement with your existing code.
- Use the urllib method for a more future-proof solution that adapts to changes in the URL's parameter order or format.
Either approach will let you loop through all 13 pages of order history smoothly. Happy scraping!
内容的提问来源于stack exchange,提问作者Simon

