You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从重复标签中提取图片描述文本?

Extracting Image Description with BeautifulSoup for Your Digital Humanities Project

Hey there! Since you're new to programming and dipping your toes into Digital Humanities, let's walk through exactly how to pull that description text using BeautifulSoup—super straightforward once you get the hang of it.

Step 1: Install BeautifulSoup

First, you'll need to install the BeautifulSoup library if you haven't already. Open your terminal and run:

pip install beautifulsoup4

We'll use Python's built-in html.parser for parsing, so no extra installs needed for that.

Step 2: Modify Your Existing Script

Here's how to update your urllib code to extract the description specifically. I'll add comments to explain each part:

import urllib.request
from bs4 import BeautifulSoup  # Import BeautifulSoup

url = "http://pitts.emory.edu/dia/image_details.cfm?ID=17250"
f = urllib.request.urlopen(url)
page_html = f.read().decode('utf-8')

# Parse the raw HTML into a navigable BeautifulSoup object
soup = BeautifulSoup(page_html, 'html.parser')

# Locate the <b> tag that contains the "Description:" label
description_label = soup.find('b', string='Description:')

# Grab the text immediately after this <b> tag (your target description!)
# .next_sibling gets the text node right after the <b> element, strip() cleans up extra spaces/line breaks
image_description = description_label.next_sibling.strip()

# Print or save your extracted description
print(image_description)

How This Works

Let's break down the key parts:

  • BeautifulSoup(page_html, 'html.parser'): Converts messy raw HTML into a structured "tree" of elements we can easily search and navigate.
  • soup.find('b', string='Description:'): Finds the exact <b> tag with the text "Description:"—this acts as our anchor point, since we know the actual description text comes right after it.
  • description_label.next_sibling: Targets the text node that follows the <b> tag, and strip() removes any unwanted leading/trailing whitespace or line breaks.

Alternative (More Robust) Approach

If the page ever has minor formatting changes (like extra spaces in the "Description:" label), you can use a lambda to match partial text instead:

# Find any <b> tag where the text contains "Description:" (ignores extra spaces)
description_label = soup.find('b', string=lambda text: text and 'Description:' in text)
image_description = description_label.next_sibling.strip()

This makes your script a bit more resilient to small tweaks in the page's HTML structure.

Since you're working on a DH project, you can easily adapt this script to loop through multiple image URLs and collect all descriptions—just wrap this logic in a function and iterate over your list of image IDs!

内容的提问来源于stack exchange,提问作者D Roberts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:30:54