You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Beautiful Soup爬取维基航海术语词汇表的列表处理问题

Handling Nested Lists in Term Descriptions When Scraping a Wikipedia Glossary

Hey there! Scraping Wikipedia glossaries can be a headache when descriptions aren't just plain text—especially when they include nested lists that break your extraction workflow. Let's fix this together for the nautical terms glossary you're working on.

The Core Issue

You're trying to pull each nautical term's title and its full description, but the descriptions often include unordered/ordered lists that aren't being captured correctly (either getting stripped entirely or losing their structure when you use get_text()).

Your Current Code (Completed for Context)

I'll fill in the missing parts of your snippet to show where the problem usually pops up:

import requests
from bs4 import BeautifulSoup

url = 'https://en.wikipedia.org/wiki/Glossary_of_nautical_terms'
page = requests.get(url)
get_title = []
get_desc = []
corrected_desc = []
output = ''

if page.status_code == 200:
    soup = BeautifulSoup(page.text, 'html.parser')
    # Targeting glossary terms (Wikipedia uses <dt class="glossary-term"> for these)
    terms = soup.find_all('dt', class_='glossary-term')
    for term in terms:
        title = term.get_text(strip=True)
        get_title.append(title)
        # The problem: this only grabs plain text, ignoring list structure
        desc = term.find_next_sibling('dd')
        get_desc.append(desc.get_text(strip=True))

The Fix: Preserve List Structure During Extraction

Instead of stripping all tags with get_text(), we'll traverse the description's elements and convert lists into a readable format (like Markdown bullets) so the structure stays intact. Here's a revised, working version:

import requests
from bs4 import BeautifulSoup

def format_description(desc_element):
    """Convert a description element (with lists) to structured text"""
    content = []
    # Loop through every child element in the description
    for child in desc_element.descendants:
        # Skip empty whitespace strings
        if child.name is None:
            text = child.strip()
            if text:
                content.append(text)
        # Handle unordered/ordered lists by converting to Markdown bullets
        elif child.name in ['ul', 'ol']:
            for li in child.find_all('li', recursive=False):
                content.append(f"- {li.get_text(strip=True)}")
        # Extract text from other tags (like links) without keeping the tags
        else:
            text = child.get_text(strip=True)
            if text:
                content.append(text)
    # Join all parts into a clean, structured description
    return '\n'.join(content)

url = 'https://en.wikipedia.org/wiki/Glossary_of_nautical_terms'
page = requests.get(url)
term_data = []

if page.status_code == 200:
    soup = BeautifulSoup(page.text, 'html.parser')
    glossary_terms = soup.find_all('dt', class_='glossary-term')
    
    for term in glossary_terms:
        # Extract the term title
        term_title = term.get_text(strip=True)
        # Get the corresponding description block
        desc_block = term.find_next_sibling('dd')
        if desc_block:
            formatted_desc = format_description(desc_block)
            term_data.append({
                'title': term_title,
                'description': formatted_desc
            })

# Test with the first 3 terms to verify output
for entry in term_data[:3]:
    print(f"**{entry['title']}**\n{entry['description']}\n---")

Key Improvements:

  • The format_description function preserves list structure by converting <li> items to Markdown bullet points.
  • It skips empty whitespace to avoid messy, spaced-out output.
  • It handles nested elements (like links within descriptions) by extracting their text while maintaining the overall flow.

Quick Tips:

  • For faster parsing, install lxml (pip install lxml) and replace 'html.parser' with 'lxml' in the BeautifulSoup call.
  • Add a custom User-Agent header to your request to comply with Wikipedia's guidelines:
    headers = {'User-Agent': 'YourScraper/1.0 (your.contact@email.com)'}
    page = requests.get(url, headers=headers)
    
  • Some terms might have multiple description blocks—if you encounter that, use term.find_next_siblings('dd') to capture all of them.

内容的提问来源于stack exchange,提问作者Bromira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:51:12