使用BeautifulSoup进行网页爬取无输出及分页导航问题求助
Hey there! Let's break down your problems and work through solutions step by step.
Fixing the Single-Page Scrape Issue
Your code runs without errors but returns nothing—this usually points to one of a few common issues. Let's debug and fix this first:
1. Verify if content is loaded dynamically
The requests library only fetches the initial raw HTML. If the table data is loaded with JavaScript (e.g., after the page loads), you won't see it in the response. To check:
- Right-click the page → View Page Source. Search for the text of any row you see on the page. If it's missing, you'll need to use a tool like Selenium or Playwright to render the full page.
2. Correct your selector logic
Your code targets <tr> elements with class odd, but the site might use alternating classes (like odd and even) or a different class name entirely. Try this adjustment:
entries = soup.find_all('tr', {'class': ['odd', 'even']})
Add debug prints to confirm if rows are being found:
print(f"Number of rows found: {len(entries)}")
3. Add browser-like headers to avoid blocking
Many sites block requests that don't mimic a real browser. Add a User-Agent header:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } response = requests.get(url, headers=headers)
4. Validate cell indices
Double-check the order of columns in the table. Print cell contents to confirm you're targeting the right indices:
for entry in entries: Cells = entry.find_all("td") if len(Cells) < 8: # Skip header rows or incomplete entries continue print([cell.get_text(strip=True) for cell in Cells]) # Verify column order
Modified Single-Page Code
Here's a version with all these fixes and debug checks:
import requests from bs4 import BeautifulSoup import pandas as pd url = "https://www.ugc.ac.in/jobportal/search.aspx?tid=MTk5Mw==" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } File = [] response = requests.get(url, headers=headers) print(f"Response status: {response.status_code}") # Should be 200 if successful soup = BeautifulSoup(response.text, "html.parser") entries = soup.find_all('tr', {'class': ['odd', 'even']}) print(f"Rows found: {len(entries)}") for entry in entries: Cells = entry.find_all("td") if len(Cells) < 8: continue columns = { 'Gender': Cells[3].get_text(strip=True), 'Category': Cells[4].get_text(strip=True), 'Subject': Cells[5].get_text(strip=True), 'NET Qualified': Cells[6].get_text(strip=True), 'Month/Year': Cells[7].get_text(strip=True) } File.append(columns) df = pd.DataFrame(File) print(df)
Handling Pagination with Identical URLs
Since the URL doesn't change when navigating pages, the site uses ASP.NET View State (a common pattern for older portals). Here's how to navigate pages:
1. Inspect network traffic to find pagination parameters
- Open Chrome DevTools (F12) → Network tab. Click the "Next" button on the page.
- Look for a POST request to the same URL. In the "Form Data" section, you'll see critical fields like:
__VIEWSTATEand__EVENTVALIDATION: Hidden values that the site uses to track session state.__EVENTTARGET: The ID of the "Next" button (e.g.,ctl00$ContentPlaceHolder1$gvData$ctl23$ctl01).
2. Extract hidden fields and send POST requests
You need to extract these hidden values from each page, then include them in a POST request to load the next page. Use a requests.Session() to persist cookies between requests.
Example Pagination Code
import requests from bs4 import BeautifulSoup import pandas as pd url = "https://www.ugc.ac.in/jobportal/search.aspx?tid=MTk5Mw==" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } session = requests.Session() all_data = [] # Load first page response = session.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # Scrape first page data (use the fixed single-page logic here) entries = soup.find_all('tr', {'class': ['odd', 'even']}) for entry in entries: Cells = entry.find_all("td") if len(Cells) < 8: continue columns = { 'Gender': Cells[3].get_text(strip=True), 'Category': Cells[4].get_text(strip=True), 'Subject': Cells[5].get_text(strip=True), 'NET Qualified': Cells[6].get_text(strip=True), 'Month/Year': Cells[7].get_text(strip=True) } all_data.append(columns) # Navigate to subsequent pages (adjust range for total pages needed) for _ in range(3): # Example: scrape 3 more pages # Extract hidden state fields viewstate = soup.find('input', {'id': '__VIEWSTATE'}).get('value') event_validation = soup.find('input', {'id': '__EVENTVALIDATION'}).get('value') # Get next button ID (inspect the page to find this) next_button = soup.find('a', text='Next') if not next_button: break # Exit loop if no more pages next_button_id = next_button.get('id') # Prepare form data for POST request form_data = { '__VIEWSTATE': viewstate, '__EVENTVALIDATION': event_validation, '__EVENTTARGET': next_button_id, '__EVENTARGUMENT': '' } # Send POST request to load next page response = session.post(url, data=form_data, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # Scrape current page data entries = soup.find_all('tr', {'class': ['odd', 'even']}) for entry in entries: Cells = entry.find_all("td") if len(Cells) < 8: continue columns = { 'Gender': Cells[3].get_text(strip=True), 'Category': Cells[4].get_text(strip=True), 'Subject': Cells[5].get_text(strip=True), 'NET Qualified': Cells[6].get_text(strip=True), 'Month/Year': Cells[7].get_text(strip=True) } all_data.append(columns) # Convert to DataFrame df = pd.DataFrame(all_data) print(df)
Notes for Pagination
- Replace the
range(3)with the number of pages you need to scrape. - The next button ID might change per page, so always extract it from the current soup.
- Add error handling for cases where the next button doesn't exist (end of pagination).
Hope this helps you resolve both issues!
内容的提问来源于stack exchange,提问作者Ishwarya Balasubramanian

