You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用lxml提取同一div内多HTML标签文本的方法

Extracting Text from Multiple Child Tags in a Single Div with lxml (Python)

Hey there! No worries at all—getting started with lxml can feel a bit tricky at first, but let's break this down step by step with clear, runnable examples.

First, make sure you have lxml installed:

pip install lxml

We'll use a sample HTML snippet to demonstrate different extraction approaches. Let's assume your HTML looks something like this:

<div class="content-container">
    <h2>Sample Heading</h2>
    <a href="/link1">First Link Text</a>
    <div class="sub-div">Text inside a sub div</div>
    <p>Paragraph text here</p>
    <a href="/link2">Second Link Text</a>
</div>

1. Extract All Text from All Child Tags (XPath)

This method grabs text from every descendant tag inside your target div, then cleans up extra whitespace:

from lxml import html

# Sample HTML content
sample_html = """
<div class="content-container">
    <h2>Sample Heading</h2>
    <a href="/link1">First Link Text</a>
    <div class="sub-div">Text inside a sub div</div>
    <p>Paragraph text here</p>
    <a href="/link2">Second Link Text</a>
</div>
"""

# Parse the HTML string into a tree structure
tree = html.fromstring(sample_html)

# Use XPath to get all text nodes under the target div
all_text_nodes = tree.xpath('//div[@class="content-container"]//text()')

# Clean up whitespace and filter out empty strings
cleaned_texts = [text.strip() for text in all_text_nodes if text.strip()]

# Print the result
print("All extracted text:")
for text in cleaned_texts:
    print(f"- {text}")

2. Extract Text from Specific Child Tags (XPath)

If you only want text from certain tags (like <a> and <div>), you can narrow down the XPath query:

# Target only <a> and <div> tags inside the container
specific_text_nodes = tree.xpath('//div[@class="content-container"]//*[self::a or self::div]/text()')
cleaned_specific_texts = [text.strip() for text in specific_text_nodes if text.strip()]

print("\nText from <a> and <div> tags only:")
for text in cleaned_specific_texts:
    print(f"- {text}")

3. Using CSS Selectors (Alternative to XPath)

lxml also supports CSS selectors via the cssselect module. Here's how to do the same tasks with CSS:

# Extract text from all child tags using CSS selectors
all_child_elements = tree.cssselect('div.content-container *')
css_all_texts = []
for elem in all_child_elements:
    # Check if the element has non-empty text
    elem_text = elem.text.strip() if elem.text else ""
    if elem_text:
        css_all_texts.append(elem_text)

print("\nAll text via CSS selectors:")
for text in css_all_texts:
    print(f"- {text}")

# Target specific tags with CSS selectors
specific_elements = tree.cssselect('div.content-container a, div.content-container div')
css_specific_texts = [elem.text.strip() for elem in specific_elements if elem.text.strip()]

print("\nText from <a> and <div> tags via CSS selectors:")
for text in css_specific_texts:
    print(f"- {text}")

Key Notes:

Hope these examples help you get started! If you run into edge cases with your actual HTML structure, just adjust the selectors to match what you're working with.

内容的提问来源于stack exchange,提问作者Carlos N.P.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:52:54