You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup进行网页爬取无输出及分页导航问题求助

Hey there! Let's break down your problems and work through solutions step by step.

Fixing the Single-Page Scrape Issue

Your code runs without errors but returns nothing—this usually points to one of a few common issues. Let's debug and fix this first:

1. Verify if content is loaded dynamically

The requests library only fetches the initial raw HTML. If the table data is loaded with JavaScript (e.g., after the page loads), you won't see it in the response. To check:

  • Right-click the page → View Page Source. Search for the text of any row you see on the page. If it's missing, you'll need to use a tool like Selenium or Playwright to render the full page.

2. Correct your selector logic

Your code targets <tr> elements with class odd, but the site might use alternating classes (like odd and even) or a different class name entirely. Try this adjustment:

entries = soup.find_all('tr', {'class': ['odd', 'even']})

Add debug prints to confirm if rows are being found:

print(f"Number of rows found: {len(entries)}")

3. Add browser-like headers to avoid blocking

Many sites block requests that don't mimic a real browser. Add a User-Agent header:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}
response = requests.get(url, headers=headers)

4. Validate cell indices

Double-check the order of columns in the table. Print cell contents to confirm you're targeting the right indices:

for entry in entries:
    Cells = entry.find_all("td")
    if len(Cells) < 8:  # Skip header rows or incomplete entries
        continue
    print([cell.get_text(strip=True) for cell in Cells])  # Verify column order

Modified Single-Page Code

Here's a version with all these fixes and debug checks:

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = "https://www.ugc.ac.in/jobportal/search.aspx?tid=MTk5Mw=="
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}

File = []
response = requests.get(url, headers=headers)
print(f"Response status: {response.status_code}")  # Should be 200 if successful

soup = BeautifulSoup(response.text, "html.parser")
entries = soup.find_all('tr', {'class': ['odd', 'even']})
print(f"Rows found: {len(entries)}")

for entry in entries:
    Cells = entry.find_all("td")
    if len(Cells) < 8:
        continue
    columns = {
        'Gender': Cells[3].get_text(strip=True),
        'Category': Cells[4].get_text(strip=True),
        'Subject': Cells[5].get_text(strip=True),
        'NET Qualified': Cells[6].get_text(strip=True),
        'Month/Year': Cells[7].get_text(strip=True)
    }
    File.append(columns)

df = pd.DataFrame(File)
print(df)

Handling Pagination with Identical URLs

Since the URL doesn't change when navigating pages, the site uses ASP.NET View State (a common pattern for older portals). Here's how to navigate pages:

1. Inspect network traffic to find pagination parameters

  • Open Chrome DevTools (F12) → Network tab. Click the "Next" button on the page.
  • Look for a POST request to the same URL. In the "Form Data" section, you'll see critical fields like:
    • __VIEWSTATE and __EVENTVALIDATION: Hidden values that the site uses to track session state.
    • __EVENTTARGET: The ID of the "Next" button (e.g., ctl00$ContentPlaceHolder1$gvData$ctl23$ctl01).

2. Extract hidden fields and send POST requests

You need to extract these hidden values from each page, then include them in a POST request to load the next page. Use a requests.Session() to persist cookies between requests.

Example Pagination Code

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = "https://www.ugc.ac.in/jobportal/search.aspx?tid=MTk5Mw=="
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}

session = requests.Session()
all_data = []

# Load first page
response = session.get(url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# Scrape first page data (use the fixed single-page logic here)
entries = soup.find_all('tr', {'class': ['odd', 'even']})
for entry in entries:
    Cells = entry.find_all("td")
    if len(Cells) < 8:
        continue
    columns = {
        'Gender': Cells[3].get_text(strip=True),
        'Category': Cells[4].get_text(strip=True),
        'Subject': Cells[5].get_text(strip=True),
        'NET Qualified': Cells[6].get_text(strip=True),
        'Month/Year': Cells[7].get_text(strip=True)
    }
    all_data.append(columns)

# Navigate to subsequent pages (adjust range for total pages needed)
for _ in range(3):  # Example: scrape 3 more pages
    # Extract hidden state fields
    viewstate = soup.find('input', {'id': '__VIEWSTATE'}).get('value')
    event_validation = soup.find('input', {'id': '__EVENTVALIDATION'}).get('value')
    
    # Get next button ID (inspect the page to find this)
    next_button = soup.find('a', text='Next')
    if not next_button:
        break  # Exit loop if no more pages
    next_button_id = next_button.get('id')
    
    # Prepare form data for POST request
    form_data = {
        '__VIEWSTATE': viewstate,
        '__EVENTVALIDATION': event_validation,
        '__EVENTTARGET': next_button_id,
        '__EVENTARGUMENT': ''
    }
    
    # Send POST request to load next page
    response = session.post(url, data=form_data, headers=headers)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # Scrape current page data
    entries = soup.find_all('tr', {'class': ['odd', 'even']})
    for entry in entries:
        Cells = entry.find_all("td")
        if len(Cells) < 8:
            continue
        columns = {
            'Gender': Cells[3].get_text(strip=True),
            'Category': Cells[4].get_text(strip=True),
            'Subject': Cells[5].get_text(strip=True),
            'NET Qualified': Cells[6].get_text(strip=True),
            'Month/Year': Cells[7].get_text(strip=True)
        }
        all_data.append(columns)

# Convert to DataFrame
df = pd.DataFrame(all_data)
print(df)

Notes for Pagination

  • Replace the range(3) with the number of pages you need to scrape.
  • The next button ID might change per page, so always extract it from the current soup.
  • Add error handling for cases where the next button doesn't exist (end of pagination).

Hope this helps you resolve both issues!

内容的提问来源于stack exchange,提问作者Ishwarya Balasubramanian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 12:32:30