You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取所有页面中「Next」按钮的链接并存入列表?

Hey there! Let's get your code working reliably to extract all those "Next" button links and save them to a list (and CSV). Your approach is on the right track, but we can make it more robust and easier to maintain. Here's how:

Key Issues in the Original Code
  • Using regex to scrape the href from raw text is fragile—HTML structure changes will break this immediately. We should use BeautifulSoup to directly target the Next button element instead.
  • No handling for when there's no Next button left (this will throw errors when re.search returns None).
  • URL conversion logic only accounts for one type of relative URL, which might miss cases where the link starts with a /.
  • Variable reuse (like overwriting sub_link) can lead to confusion and bugs.
Improved Working Code
import requests
from bs4 import BeautifulSoup
import pandas as pd

def convert_relative_url(url):
    """Convert relative URLs to full absolute URLs for the target site"""
    if not url.startswith(('http://', 'https://')):
        # Handle both root-relative (/path) and page-relative (path) links
        base_url = 'https://money.rediff.com'
        if url.startswith('/'):
            return f"{base_url}{url}"
        return f"{base_url}/{url}"
    return url

# Initialize list to store all unique Next links
next_links = []
current_page_url = 'https://money.rediff.com/companies/groups/A'

while current_page_url:
    try:
        # Fetch the page with error handling for HTTP issues
        response = requests.get(current_page_url)
        response.raise_for_status()  # Trigger error for 4xx/5xx status codes
        soup = BeautifulSoup(response.content, 'html.parser')
        
        # Find the Next button by its text (more reliable than regex)
        next_button = soup.find('a', text=lambda text: text and 'Next' in text.strip())
        
        if next_button:
            next_href = next_button.get('href')
            absolute_next_url = convert_relative_url(next_href)
            
            # Avoid duplicates and infinite loops
            if absolute_next_url not in next_links:
                next_links.append(absolute_next_url)
                current_page_url = absolute_next_url
            else:
                print("Detected a duplicate link—stopping to avoid infinite loop.")
                break
        else:
            # No more Next buttons to process
            print("No additional Next buttons found. Exiting loop.")
            break
            
    except requests.exceptions.RequestException as e:
        print(f"Failed to load page {current_page_url}: {str(e)}")
        break

# Save results to CSV
results_df = pd.DataFrame({'Link': next_links})
print(f"Successfully collected {len(next_links)} Next links.")
print(f"DataFrame shape: {results_df.shape}")
results_df.to_csv('results_1.csv', index=False)
What Changed & Why
  • Targeted Element Search: Instead of regex, we use BeautifulSoup to find the <a> tag containing "Next" text. This is way more resilient to minor HTML changes (like extra whitespace or nested elements).
  • Robust URL Handling: The convert_relative_url function handles all common relative URL formats, ensuring we always get a valid absolute link.
  • Error Prevention: We added HTTP error handling and checks for duplicate links to avoid crashes and infinite loops.
  • Readability: Renamed variables to be more descriptive (e.g., next_links instead of sub_link_edit) so the code is easier to follow later.

内容的提问来源于stack exchange,提问作者Math Inquiry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 20:57:26