You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup4提取网页div中带链接的文本内容?

How to Transform HTML Content with BeautifulSoup as a Python Newbie

Hey there! Let's tackle this problem step by step—since you're new to Python and BeautifulSoup, I'll keep things clear and straightforward. Your goal is to take a <div> element with mixed text and links, and rewrite it so each link shows as Link Text ("Actual URL") while keeping the rest of the text intact.

The Approach

Instead of trying to slice or use before/after methods (which can get messy with mixed content), we'll traverse every child node inside the <div>. For each node:

  • If it's plain text, we just keep it as-is.
  • If it's an <a> tag, we format it to include both the link text and its URL in parentheses.

Example Code

Here's a complete, runnable solution:

from bs4 import BeautifulSoup, NavigableString, Tag

# Your original HTML content
html = '<div>Some TEXT with <a href="some Link">some LINK</a> and some continuing TEXT with following <a href="some Link">some LINK</a> inside.</div>'

# Parse the HTML with BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')
div_element = soup.find('div')

# Initialize an empty list to build our result
result_parts = []

# Iterate over every child node in the div
for node in div_element.contents:
    # Check if the node is plain text
    if isinstance(node, NavigableString):
        result_parts.append(node.strip())  # Strip extra whitespace (optional but clean)
    # Check if the node is an <a> tag
    elif isinstance(node, Tag) and node.name == 'a':
        link_text = node.get_text(strip=True)
        link_url = node.get('href')
        # Format the link as "Link Text ("URL")"
        formatted_link = f'{link_text} ("{link_url}")'
        result_parts.append(formatted_link)

# Join all parts into a single string
final_result = ' '.join(result_parts)
print(final_result)

What This Does

  • div_element.contents gives us all direct children of the <div>—this includes both text snippets and <a> tags, in the order they appear.
  • We check each node's type:
    • NavigableString is BeautifulSoup's way of representing plain text nodes. We add this text directly to our result.
    • Tag represents HTML elements (like <a>). For links, we extract the text inside the tag with .get_text() and the URL with .get('href'), then format them into the string you want.
  • Finally, we join all parts together into a single, clean string.

Output

Running this code will give you exactly what you're looking for:

Some TEXT with some LINK ("some Link") and some continuing TEXT with following some LINK ("some Link") inside.

This method works even if your <div> has more complex mixed content (like multiple links, extra text, or other tags)—it's flexible and easy to adjust if you need to handle other element types later.

内容的提问来源于stack exchange,提问作者dustixx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:38:12