如何用Python与BeautifulSoup提取乌尔曼工业化学全书所有章节DOI
Alright, let's figure out how to extract chapter names and full DOI paths from Ullmann's Encyclopedia of Industrial Chemistry using Python and BeautifulSoup. I’ve dealt with similar scraping tasks before, so here’s a straightforward approach that handles both your sample case and more complex scenarios:
Step 1: Install Dependencies
First, make sure you have the required libraries installed. Run this in your terminal:
pip install beautifulsoup4 requests
(We’ll use requests if you need to fetch the actual webpage; if you’re working with local HTML files, you can skip it.)
Step 2: Basic Extraction (Your Sample HTML)
Let’s start with the exact HTML snippet you provided. We’ll parse it to get the chapter name and convert the DOI path to the full version:
from bs4 import BeautifulSoup # Your sample HTML sample_html = '''<h2 class="meta__title meta__title__margin"><span class="hlFld-Title"><a href="/doi/10.1002/14356007.c01_c01.pub2">Aerogels</a></span></h2>''' # Parse the HTML soup = BeautifulSoup(sample_html, 'html.parser') # Extract chapter name chapter_name = soup.find('span', class_='hlFld-Title').get_text(strip=True) # Convert raw DOI path to full version raw_doi_href = soup.find('span', class_='hlFld-Title').find('a')['href'] full_doi_path = raw_doi_href.replace('/doi/', '/doi/full/') # Output results print(f"Chapter Name: {chapter_name}") print(f"Full DOI Path: {full_doi_path}")
Running this will give you:
Chapter Name: Aerogels Full DOI Path: /doi/full/10.1002/14356007.c01_c01.pub2
Step 3: Handling Complex, Multi-Chapter HTML
If you’re working with a full page containing multiple chapters, we can loop through all matching elements to extract every entry:
from bs4 import BeautifulSoup # Example of a more complex HTML with multiple chapters complex_html = ''' <div class="content-wrapper"> <h2 class="meta__title meta__title__margin"><span class="hlFld-Title"><a href="/doi/10.1002/14356007.c01_c01.pub2">Aerogels</a></span></h2> <h2 class="meta__title meta__title__margin"><span class="hlFld-Title"><a href="/doi/10.1002/14356007.c02_c01.pub3">Biopolymers</a></span></h2> <h2 class="meta__title meta__title__margin"><span class="hlFld-Title"><a href="/doi/10.1002/14356007.c03_c02.pub1">Catalysis</a></span></h2> </div> ''' soup = BeautifulSoup(complex_html, 'html.parser') # Collect all chapters in a list of dictionaries chapter_data = [] for title_span in soup.find_all('span', class_='hlFld-Title'): # Get clean chapter name name = title_span.get_text(strip=True) # Get and format DOI path raw_doi = title_span.find('a')['href'] full_doi = raw_doi.replace('/doi/', '/doi/full/') # Add to list chapter_data.append({ 'chapter_name': name, 'full_doi_path': full_doi }) # Print all results for i, chapter in enumerate(chapter_data, 1): print(f"Chapter {i}:") print(f" Name: {chapter['chapter_name']}") print(f" Full DOI: {chapter['full_doi_path']}\n")
Step 4: Edge Case Handling
Sometimes, the DOI link might already include /full/ (though unlikely in your case). To avoid double-converting, add a simple check:
def format_full_doi(raw_href): if '/doi/full/' not in raw_href: return raw_href.replace('/doi/', '/doi/full/') return raw_href # Use this function instead of direct replace full_doi_path = format_full_doi(raw_doi_href)
Step 5: Fetching from a Live Webpage (Optional)
If you need to scrape the actual encyclopedia page, use requests to fetch the HTML first:
import requests from bs4 import BeautifulSoup url = "your_ullmann_page_url_here" response = requests.get(url) response.raise_for_status() # Raise error if request fails soup = BeautifulSoup(response.text, 'html.parser') # Now use the same extraction code as above
This approach is robust because it targets the specific hlFld-Title span class that wraps the chapter title and DOI link, which should be consistent across the encyclopedia’s pages.
内容的提问来源于stack exchange,提问作者pickenpack

