如何用Python、BeautifulSoup和re将网页提取文本转为字典?
Hey there! Let's get your script cleaned up, make it Pythonic, and get that perfect dictionary output you need. Here's a refactored version with explanations:
Refactored Pythonic Script
from urllib.request import Request, urlopen from bs4 import BeautifulSoup import re url = "https://www.itis.gov/servlet/SingleRpt/SingleRpt?search_topic=TSN&search_value=183833#null" page = urlopen(Request(url, headers={'User-Agent': 'Mozilla/5.0'})) soup = BeautifulSoup(page, 'html.parser') taxonomy_dict = {} target_categories = {'Kingdom', 'Phylum', 'Division', 'Class', 'Order', 'Family', 'Genus', 'Species'} for tr in soup.find_all('tr'): text = tr.get_text().strip() if not text: continue # Match lines starting with our target categories, split into key and value segment match = re.match(r'^(Kingdom|Phylum|Division|Class|Order|Family|Genus|Species)(.*)$', text) if match: key = match.group(1) value_segment = match.group(2).strip() # Clean values based on category type if key == 'Species': # Keep only the first two words to preserve binomial nomenclature cleaned_value = ' '.join(value_segment.split()[:2]) else: # For other categories, take just the first word (ignore author/year details) cleaned_value = value_segment.split()[0] taxonomy_dict[key] = cleaned_value print(taxonomy_dict)
Key Improvements & Breakdown
- PEP8 Compliance: Split imports onto separate lines, use descriptive snake_case variable names, and keep the logic concise but readable
- Precise Pattern Matching: The regex cleanly separates each line into the category key and the rest of the content, avoiding messy string splits
- Targeted Value Cleaning:
- For most categories, we strip out extra author/year info by taking only the first word of the value segment
- For
Species, we specifically retain the first two words to maintain the required binomial format (e.g., "Lycaon pictus")
- Clean Flow: The loop skips empty lines and only processes relevant rows, making the script efficient and easy to debug
Expected Output
Running this will produce exactly the dictionary you're looking for:
{ 'Kingdom': 'Animalia', 'Phylum': 'Chordata', 'Class': 'Mammalia', 'Order': 'Carnivora', 'Family': 'Canidae', 'Genus': 'Lycaon', 'Species': 'Lycaon pictus' }
内容的提问来源于stack exchange,提问作者Ninja Warrior 11
相关产品推荐
相关产品推荐

