Python Requests响应缺失HTML:翻页功能异常排查求助
Let’s break down why your code isn’t picking up the "next" button and fix it step by step—this is a common scraping pitfall, so we’ll cover all the likely gaps:
1. You’re Using the Wrong HTML Parser
Your code uses BeautifulSoup(page2.content, 'html')—the 'html' parser isn’t a standard option for BeautifulSoup. It’s likely causing inconsistent parsing of the page structure, which could hide the pager element. Replace it with either:
'html.parser'(built into Python, no extra installation needed)'lxml'(faster, more robust—install withpip install lxmlfirst)
2. The Pager Element Might Not Be Inside #region-content
You’re limiting your search for the "next" button to the #region-content container, but some sites place pagination controls outside the main content area. Try searching the entire soup first, then fall back to the region if needed.
3. Recursion Isn’t Ideal for Pagination
Your recursive approach will hit Python’s recursion depth limit if there are many pages. A while loop is more reliable and avoids stack overflow issues.
4. Your Headers Are Still Incomplete
Even with a User-Agent, missing other critical request headers (like Accept, Accept-Language) can make the server return a simplified or different version of the page.模拟 a real browser’s request headers to ensure you get the full page content.
Fixed Code Implementation
Here’s the revised code addressing all these issues:
from bs4 import BeautifulSoup import requests from csv import reader def get_all_broker_links(): broker_links = [] # Mimic a real browser's request headers headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.cbp.gov/contact/find-broker-by-port' } with open('port.csv', 'r') as read_obj: csv_reader = reader(read_obj) for row in csv_reader: base_url = row[-1] page_num = 0 while True: try: # Ensure your base URL has a placeholder for the page number, e.g., "https://...?page={}" current_url = base_url.format(page_num) response = requests.get(current_url, headers=headers) response.raise_for_status() # Throw an error for HTTP status codes like 404/500 soup = BeautifulSoup(response.content, 'html.parser') # First check the entire page for the next button, then narrow to region-content next_page = soup.find('li', class_='pager-next') if not next_page: region_content = soup.find(id='region-content') next_page = region_content.find('li', class_='pager-next') if region_content else None # Extract broker links region_content = soup.find(id='region-content') if region_content: table_cells = region_content.find_all('td', class_='views-field views-field-title views-align-center') for cell in table_cells: link_tag = cell.find('a', href=True) if link_tag: full_link = 'https://www.cbp.gov' + link_tag['href'] broker_links.append(full_link) # Exit loop if no next page exists if not next_page: break page_num += 1 except requests.exceptions.RequestException as e: print(f"Error accessing {current_url}: {str(e)}") break return broker_links # Run the function all_links = get_all_broker_links() print(f"Collected {len(all_links)} broker links")
Final Checks to Verify
- Double-check that the URLs in
port.csvinclude a{}placeholder for the page number (e.g.,https://www.cbp.gov/contact/find-broker-by-port/4901?page={}). If not, adjust the base URL formatting. - Use your browser’s DevTools (Network tab) to inspect the actual request headers sent when navigating the site, then match them in your code for maximum accuracy.
- If the page still doesn’t show the pager, right-click the "next" button in your browser and select "Inspect" to confirm its exact class name and parent container—sometimes class names have subtle typos or variations.
内容的提问来源于stack exchange,提问作者Muhammad Ashfaq

