Python网页爬取无结果问题求助:爬取Books to Scrape网站返回空DataFrame
Looking at your code, there are a few key issues causing the empty DataFrame. Let's break them down and fix them step by step:
1. Broken get_links Function
Your get_links function contains nested functions that are never called, and it doesn't return any value (so it defaults to None). When you pass None to extract_info, the loop over links doesn't run, leaving your all_books list empty.
2. Incorrect URL Construction
String concatenation for building book URLs leads to missing slashes (e.g., http://books.toscrape.com/cataloguea-light-in-the-attic_1000/index.html instead of the correct path with a slash after catalogue). Using urljoin from the urllib.parse module fixes this by properly handling relative paths.
3. Unreliable Price Extraction
While your current price selector works, targeting the specific price_color class makes the code more robust in case the page structure changes slightly.
Corrected Code
Here's the fixed version of your code with explanations:
import pandas as pd import requests from bs4 import BeautifulSoup as bs from urllib.parse import urljoin # Import for safe URL construction all_books = [] def get_page(url): page = requests.get(url) status = page.status_code soup = bs(page.text, "lxml") return [soup, status] def get_links(soup, base_url): links = [] listings = soup.find_all(class_="product_pod") for listing in listings: # Extract relative link from the product pod relative_link = listing.find("h5").a.get("href") # Build full absolute URL using urljoin to avoid path errors full_link = urljoin(base_url, relative_link) links.append(full_link) return links # Return the list of book links def extract_info(links): for link in links: res = requests.get(link).text book_soup = bs(res, "lxml") # Extract title from the h1 tag in product_main title = book_soup.find(class_="col-sm-6 product_main").h1.text.strip() # Extract price using the specific price_color class for reliability price = book_soup.find(class_="price_color").text.strip() book = {"title": title, "price": price} all_books.append(book) pg = 1 while True: url = f"http://books.toscrape.com/catalogue/page-{pg}.html" soup_status = get_page(url) if soup_status[1] == 200: print(f"scraping page {pg}") # Get valid book links from the current page book_links = get_links(soup_status[0], url) # Process each link to extract book info extract_info(book_links) pg += 1 else: print("The End") break df = pd.DataFrame(all_books) print(df)
Key Changes Made:
- Removed unused nested functions in
get_linksand ensured it returns the list of valid book links. - Added
urljointo safely construct absolute URLs from relative paths, eliminating broken links. - Updated price extraction to target the
price_colorclass directly for better reliability. - Ensured valid links are passed to
extract_info, so the loop runs and populatesall_books.
When you run this code, it should correctly scrape all books from the site and populate your DataFrame with titles and prices.
内容的提问来源于stack exchange,提问作者BJ Creative

