You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Requests响应缺失HTML:翻页功能异常排查求助

Troubleshooting "Next" Button Detection in CBP Broker Page Scraping

Let’s break down why your code isn’t picking up the "next" button and fix it step by step—this is a common scraping pitfall, so we’ll cover all the likely gaps:

1. You’re Using the Wrong HTML Parser

Your code uses BeautifulSoup(page2.content, 'html')—the 'html' parser isn’t a standard option for BeautifulSoup. It’s likely causing inconsistent parsing of the page structure, which could hide the pager element. Replace it with either:

  • 'html.parser' (built into Python, no extra installation needed)
  • 'lxml' (faster, more robust—install with pip install lxml first)

2. The Pager Element Might Not Be Inside #region-content

You’re limiting your search for the "next" button to the #region-content container, but some sites place pagination controls outside the main content area. Try searching the entire soup first, then fall back to the region if needed.

3. Recursion Isn’t Ideal for Pagination

Your recursive approach will hit Python’s recursion depth limit if there are many pages. A while loop is more reliable and avoids stack overflow issues.

4. Your Headers Are Still Incomplete

Even with a User-Agent, missing other critical request headers (like Accept, Accept-Language) can make the server return a simplified or different version of the page.模拟 a real browser’s request headers to ensure you get the full page content.

Fixed Code Implementation

Here’s the revised code addressing all these issues:

from bs4 import BeautifulSoup
import requests
from csv import reader

def get_all_broker_links():
    broker_links = []
    # Mimic a real browser's request headers
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
        'Accept-Language': 'en-US,en;q=0.5',
        'Referer': 'https://www.cbp.gov/contact/find-broker-by-port'
    }

    with open('port.csv', 'r') as read_obj:
        csv_reader = reader(read_obj)
        for row in csv_reader:
            base_url = row[-1]
            page_num = 0
            while True:
                try:
                    # Ensure your base URL has a placeholder for the page number, e.g., "https://...?page={}"
                    current_url = base_url.format(page_num)
                    response = requests.get(current_url, headers=headers)
                    response.raise_for_status()  # Throw an error for HTTP status codes like 404/500

                    soup = BeautifulSoup(response.content, 'html.parser')
                    
                    # First check the entire page for the next button, then narrow to region-content
                    next_page = soup.find('li', class_='pager-next')
                    if not next_page:
                        region_content = soup.find(id='region-content')
                        next_page = region_content.find('li', class_='pager-next') if region_content else None

                    # Extract broker links
                    region_content = soup.find(id='region-content')
                    if region_content:
                        table_cells = region_content.find_all('td', class_='views-field views-field-title views-align-center')
                        for cell in table_cells:
                            link_tag = cell.find('a', href=True)
                            if link_tag:
                                full_link = 'https://www.cbp.gov' + link_tag['href']
                                broker_links.append(full_link)

                    # Exit loop if no next page exists
                    if not next_page:
                        break

                    page_num += 1

                except requests.exceptions.RequestException as e:
                    print(f"Error accessing {current_url}: {str(e)}")
                    break

    return broker_links

# Run the function
all_links = get_all_broker_links()
print(f"Collected {len(all_links)} broker links")

Final Checks to Verify

  • Double-check that the URLs in port.csv include a {} placeholder for the page number (e.g., https://www.cbp.gov/contact/find-broker-by-port/4901?page={}). If not, adjust the base URL formatting.
  • Use your browser’s DevTools (Network tab) to inspect the actual request headers sent when navigating the site, then match them in your code for maximum accuracy.
  • If the page still doesn’t show the pager, right-click the "next" button in your browser and select "Inspect" to confirm its exact class name and parent container—sometimes class names have subtle typos or variations.

内容的提问来源于stack exchange,提问作者Muhammad Ashfaq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 20:48:01