You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python、BeautifulSoup和re将网页提取文本转为字典?

Hey there! Let's get your script cleaned up, make it Pythonic, and get that perfect dictionary output you need. Here's a refactored version with explanations:

Refactored Pythonic Script

from urllib.request import Request, urlopen
from bs4 import BeautifulSoup
import re

url = "https://www.itis.gov/servlet/SingleRpt/SingleRpt?search_topic=TSN&search_value=183833#null"
page = urlopen(Request(url, headers={'User-Agent': 'Mozilla/5.0'}))
soup = BeautifulSoup(page, 'html.parser')

taxonomy_dict = {}
target_categories = {'Kingdom', 'Phylum', 'Division', 'Class', 'Order', 'Family', 'Genus', 'Species'}

for tr in soup.find_all('tr'):
    text = tr.get_text().strip()
    if not text:
        continue
    
    # Match lines starting with our target categories, split into key and value segment
    match = re.match(r'^(Kingdom|Phylum|Division|Class|Order|Family|Genus|Species)(.*)$', text)
    if match:
        key = match.group(1)
        value_segment = match.group(2).strip()
        
        # Clean values based on category type
        if key == 'Species':
            # Keep only the first two words to preserve binomial nomenclature
            cleaned_value = ' '.join(value_segment.split()[:2])
        else:
            # For other categories, take just the first word (ignore author/year details)
            cleaned_value = value_segment.split()[0]
        
        taxonomy_dict[key] = cleaned_value

print(taxonomy_dict)

Key Improvements & Breakdown

  • PEP8 Compliance: Split imports onto separate lines, use descriptive snake_case variable names, and keep the logic concise but readable
  • Precise Pattern Matching: The regex cleanly separates each line into the category key and the rest of the content, avoiding messy string splits
  • Targeted Value Cleaning:
    • For most categories, we strip out extra author/year info by taking only the first word of the value segment
    • For Species, we specifically retain the first two words to maintain the required binomial format (e.g., "Lycaon pictus")
  • Clean Flow: The loop skips empty lines and only processes relevant rows, making the script efficient and easy to debug

Expected Output

Running this will produce exactly the dictionary you're looking for:

{
    'Kingdom': 'Animalia',
    'Phylum': 'Chordata',
    'Class': 'Mammalia',
    'Order': 'Carnivora',
    'Family': 'Canidae',
    'Genus': 'Lycaon',
    'Species': 'Lycaon pictus'
}

内容的提问来源于stack exchange,提问作者Ninja Warrior 11

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:43:05