如何使用Python提取所有页面中「Next」按钮的链接并存入列表?
Hey there! Let's get your code working reliably to extract all those "Next" button links and save them to a list (and CSV). Your approach is on the right track, but we can make it more robust and easier to maintain. Here's how:
Key Issues in the Original Code
- Using regex to scrape the
hreffrom raw text is fragile—HTML structure changes will break this immediately. We should use BeautifulSoup to directly target the Next button element instead. - No handling for when there's no Next button left (this will throw errors when
re.searchreturnsNone). - URL conversion logic only accounts for one type of relative URL, which might miss cases where the link starts with a
/. - Variable reuse (like overwriting
sub_link) can lead to confusion and bugs.
Improved Working Code
import requests from bs4 import BeautifulSoup import pandas as pd def convert_relative_url(url): """Convert relative URLs to full absolute URLs for the target site""" if not url.startswith(('http://', 'https://')): # Handle both root-relative (/path) and page-relative (path) links base_url = 'https://money.rediff.com' if url.startswith('/'): return f"{base_url}{url}" return f"{base_url}/{url}" return url # Initialize list to store all unique Next links next_links = [] current_page_url = 'https://money.rediff.com/companies/groups/A' while current_page_url: try: # Fetch the page with error handling for HTTP issues response = requests.get(current_page_url) response.raise_for_status() # Trigger error for 4xx/5xx status codes soup = BeautifulSoup(response.content, 'html.parser') # Find the Next button by its text (more reliable than regex) next_button = soup.find('a', text=lambda text: text and 'Next' in text.strip()) if next_button: next_href = next_button.get('href') absolute_next_url = convert_relative_url(next_href) # Avoid duplicates and infinite loops if absolute_next_url not in next_links: next_links.append(absolute_next_url) current_page_url = absolute_next_url else: print("Detected a duplicate link—stopping to avoid infinite loop.") break else: # No more Next buttons to process print("No additional Next buttons found. Exiting loop.") break except requests.exceptions.RequestException as e: print(f"Failed to load page {current_page_url}: {str(e)}") break # Save results to CSV results_df = pd.DataFrame({'Link': next_links}) print(f"Successfully collected {len(next_links)} Next links.") print(f"DataFrame shape: {results_df.shape}") results_df.to_csv('results_1.csv', index=False)
What Changed & Why
- Targeted Element Search: Instead of regex, we use BeautifulSoup to find the
<a>tag containing "Next" text. This is way more resilient to minor HTML changes (like extra whitespace or nested elements). - Robust URL Handling: The
convert_relative_urlfunction handles all common relative URL formats, ensuring we always get a valid absolute link. - Error Prevention: We added HTTP error handling and checks for duplicate links to avoid crashes and infinite loops.
- Readability: Renamed variables to be more descriptive (e.g.,
next_linksinstead ofsub_link_edit) so the code is easier to follow later.
内容的提问来源于stack exchange,提问作者Math Inquiry
相关产品推荐
相关产品推荐

