基于Beautiful Soup爬取维基航海术语词汇表的列表处理问题
Hey there! Scraping Wikipedia glossaries can be a headache when descriptions aren't just plain text—especially when they include nested lists that break your extraction workflow. Let's fix this together for the nautical terms glossary you're working on.
The Core Issue
You're trying to pull each nautical term's title and its full description, but the descriptions often include unordered/ordered lists that aren't being captured correctly (either getting stripped entirely or losing their structure when you use get_text()).
Your Current Code (Completed for Context)
I'll fill in the missing parts of your snippet to show where the problem usually pops up:
import requests from bs4 import BeautifulSoup url = 'https://en.wikipedia.org/wiki/Glossary_of_nautical_terms' page = requests.get(url) get_title = [] get_desc = [] corrected_desc = [] output = '' if page.status_code == 200: soup = BeautifulSoup(page.text, 'html.parser') # Targeting glossary terms (Wikipedia uses <dt class="glossary-term"> for these) terms = soup.find_all('dt', class_='glossary-term') for term in terms: title = term.get_text(strip=True) get_title.append(title) # The problem: this only grabs plain text, ignoring list structure desc = term.find_next_sibling('dd') get_desc.append(desc.get_text(strip=True))
The Fix: Preserve List Structure During Extraction
Instead of stripping all tags with get_text(), we'll traverse the description's elements and convert lists into a readable format (like Markdown bullets) so the structure stays intact. Here's a revised, working version:
import requests from bs4 import BeautifulSoup def format_description(desc_element): """Convert a description element (with lists) to structured text""" content = [] # Loop through every child element in the description for child in desc_element.descendants: # Skip empty whitespace strings if child.name is None: text = child.strip() if text: content.append(text) # Handle unordered/ordered lists by converting to Markdown bullets elif child.name in ['ul', 'ol']: for li in child.find_all('li', recursive=False): content.append(f"- {li.get_text(strip=True)}") # Extract text from other tags (like links) without keeping the tags else: text = child.get_text(strip=True) if text: content.append(text) # Join all parts into a clean, structured description return '\n'.join(content) url = 'https://en.wikipedia.org/wiki/Glossary_of_nautical_terms' page = requests.get(url) term_data = [] if page.status_code == 200: soup = BeautifulSoup(page.text, 'html.parser') glossary_terms = soup.find_all('dt', class_='glossary-term') for term in glossary_terms: # Extract the term title term_title = term.get_text(strip=True) # Get the corresponding description block desc_block = term.find_next_sibling('dd') if desc_block: formatted_desc = format_description(desc_block) term_data.append({ 'title': term_title, 'description': formatted_desc }) # Test with the first 3 terms to verify output for entry in term_data[:3]: print(f"**{entry['title']}**\n{entry['description']}\n---")
Key Improvements:
- The
format_descriptionfunction preserves list structure by converting<li>items to Markdown bullet points. - It skips empty whitespace to avoid messy, spaced-out output.
- It handles nested elements (like links within descriptions) by extracting their text while maintaining the overall flow.
Quick Tips:
- For faster parsing, install
lxml(pip install lxml) and replace'html.parser'with'lxml'in the BeautifulSoup call. - Add a custom
User-Agentheader to your request to comply with Wikipedia's guidelines:headers = {'User-Agent': 'YourScraper/1.0 (your.contact@email.com)'} page = requests.get(url, headers=headers) - Some terms might have multiple description blocks—if you encounter that, use
term.find_next_siblings('dd')to capture all of them.
内容的提问来源于stack exchange,提问作者Bromira

