You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python与BeautifulSoup提取乌尔曼工业化学全书所有章节DOI

Alright, let's figure out how to extract chapter names and full DOI paths from Ullmann's Encyclopedia of Industrial Chemistry using Python and BeautifulSoup. I’ve dealt with similar scraping tasks before, so here’s a straightforward approach that handles both your sample case and more complex scenarios:

Step 1: Install Dependencies

First, make sure you have the required libraries installed. Run this in your terminal:

pip install beautifulsoup4 requests

(We’ll use requests if you need to fetch the actual webpage; if you’re working with local HTML files, you can skip it.)

Step 2: Basic Extraction (Your Sample HTML)

Let’s start with the exact HTML snippet you provided. We’ll parse it to get the chapter name and convert the DOI path to the full version:

from bs4 import BeautifulSoup

# Your sample HTML
sample_html = '''<h2 class="meta__title meta__title__margin"><span class="hlFld-Title"><a href="/doi/10.1002/14356007.c01_c01.pub2">Aerogels</a></span></h2>'''

# Parse the HTML
soup = BeautifulSoup(sample_html, 'html.parser')

# Extract chapter name
chapter_name = soup.find('span', class_='hlFld-Title').get_text(strip=True)

# Convert raw DOI path to full version
raw_doi_href = soup.find('span', class_='hlFld-Title').find('a')['href']
full_doi_path = raw_doi_href.replace('/doi/', '/doi/full/')

# Output results
print(f"Chapter Name: {chapter_name}")
print(f"Full DOI Path: {full_doi_path}")

Running this will give you:

Chapter Name: Aerogels
Full DOI Path: /doi/full/10.1002/14356007.c01_c01.pub2

Step 3: Handling Complex, Multi-Chapter HTML

If you’re working with a full page containing multiple chapters, we can loop through all matching elements to extract every entry:

from bs4 import BeautifulSoup

# Example of a more complex HTML with multiple chapters
complex_html = '''
<div class="content-wrapper">
    <h2 class="meta__title meta__title__margin"><span class="hlFld-Title"><a href="/doi/10.1002/14356007.c01_c01.pub2">Aerogels</a></span></h2>
    <h2 class="meta__title meta__title__margin"><span class="hlFld-Title"><a href="/doi/10.1002/14356007.c02_c01.pub3">Biopolymers</a></span></h2>
    <h2 class="meta__title meta__title__margin"><span class="hlFld-Title"><a href="/doi/10.1002/14356007.c03_c02.pub1">Catalysis</a></span></h2>
</div>
'''

soup = BeautifulSoup(complex_html, 'html.parser')

# Collect all chapters in a list of dictionaries
chapter_data = []
for title_span in soup.find_all('span', class_='hlFld-Title'):
    # Get clean chapter name
    name = title_span.get_text(strip=True)
    # Get and format DOI path
    raw_doi = title_span.find('a')['href']
    full_doi = raw_doi.replace('/doi/', '/doi/full/')
    # Add to list
    chapter_data.append({
        'chapter_name': name,
        'full_doi_path': full_doi
    })

# Print all results
for i, chapter in enumerate(chapter_data, 1):
    print(f"Chapter {i}:")
    print(f"  Name: {chapter['chapter_name']}")
    print(f"  Full DOI: {chapter['full_doi_path']}\n")

Step 4: Edge Case Handling

Sometimes, the DOI link might already include /full/ (though unlikely in your case). To avoid double-converting, add a simple check:

def format_full_doi(raw_href):
    if '/doi/full/' not in raw_href:
        return raw_href.replace('/doi/', '/doi/full/')
    return raw_href

# Use this function instead of direct replace
full_doi_path = format_full_doi(raw_doi_href)

Step 5: Fetching from a Live Webpage (Optional)

If you need to scrape the actual encyclopedia page, use requests to fetch the HTML first:

import requests
from bs4 import BeautifulSoup

url = "your_ullmann_page_url_here"
response = requests.get(url)
response.raise_for_status()  # Raise error if request fails

soup = BeautifulSoup(response.text, 'html.parser')

# Now use the same extraction code as above

This approach is robust because it targets the specific hlFld-Title span class that wraps the chapter title and DOI link, which should be consistent across the encyclopedia’s pages.

内容的提问来源于stack exchange,提问作者pickenpack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:44:58