You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于XPath与Python lxml提取HTML中h4及关联文本、链接的技术咨询

Your current script pulls all elements separately, but to pair each h4 with its corresponding links and text content, we need to process each topLinks container individually. Here's the corrected approach using XPath to group the content properly:

Corrected Python Script

from lxml.html import fromstring

url='''
<div class="topLinks">
 <div class="hd left">
 </div><div class="hd-middle middle">
 <h4>TTTTTTTTTTTTT</h4></div><div class="hd right"></div><div class="boxMiddle"><ul><li><a href="FullStory.aspx?gid=4&id=6516" title="1399/03/18" target="_blank">PPPPPPPPPPPPPPP<img class="new" src="images/new.png"></a></li><li><a href="http://register1.sanjesh.org/fanni99up" title="1399/03/11" target="_blank">CCCCCCCCCCCCC</a></li><li><a href="http://www6.sanjesh.org/download/fani99/FaniNote99.pdf" title="1399/03/11" target="_blank"> ZZZZZZZZ </a></li><li><a href="FullStory.aspx?gid=4&id=6509" title="1399/03/11" target="_blank">FFFFFF</a></li><li><a href="FullStory.aspx?gid=4&id=6498" title="1399/02/21" target="_blank">XXXXXXXXXXXXXX </a></li></ul></div><div class="boxBottom"></div></div>
 <div class="topLinks"><div class="hd left_alter"></div><div class="hd-middle middle_alter">
 <h4>CCCCCCCCCCCC</h4></div><div class="hd right_alter"></div><div class="boxMiddle_alter"><ul><li><a href="http://register1.sanjesh.org/rgempiactax99/" title="1399/03/18" target="_blank">GGGGGGGGGGGGGGGG <img class="new" src="images/new.png"></a></li><li><a href="FullStory.aspx?gid=11&id=6515" title="1399/03/18" target="_blank">FFFFFFFFF<img class="new" src="images/new.png"></a></li><li><a href="http://register2.sanjesh.org/RGKhanevadehConsult/" title="1399/03/12" target="_blank">HHHHHHHHH</a></li><li><a href="FullStory.aspx?gid=11&id=6512" title="1399/03/12" target="_blank">FFFFFFFF</a></li><li><a href="FullStory.aspx?gid=11&id=6505" title="1399/02/24" target="_blank">NNNNNNNNNNNNNNNNNNNNNNNNNN</a></li><li><a href="http://dl.sanjesh.org/NOETDownload/DownloadHandler.ashx?id=1271" title="1398/12/12" target="_blank">OOOOOOOOOOOO</a></li><li><a href="FullStory.aspx?gid=11&id=6480" title="1399/01/26" target="_blank">JJJJJJJ</a></li></ul></div><div class="boxBottom_alter"></div></div>
'''

tree = fromstring(url)

# Iterate over each topLinks container to group content
for section in tree.xpath("//div[@class='topLinks']"):
    # Extract the h4 title for this section
    h4_title = section.xpath(".//h4/text()")[0].strip()
    print(f"*Section Title:* {h4_title}\n")
    
    # Get all link elements within the current section
    link_elements = section.xpath(".//a")
    for index, link in enumerate(link_elements, 1):
        # Clean up link text and extract URL
        link_text = link.xpath("text()")[0].strip()
        link_url = link.xpath("@href")[0]
        print(f"{index}. Text: `{link_text}` | URL: `{link_url}`")
    
    # Add a separator between sections for readability
    print("\n" + "-"*50 + "\n")

Key Improvements Explained

  • Grouped Processing: By looping over each topLinks div and using .// in XPath (the dot restricts searches to the current section), we ensure links are paired with their correct h4 title.
  • Whitespace Cleaning: Using .strip() removes extra spaces from text elements for cleaner output.
  • Readable Formatting: The output clearly maps each section title to its associated links and text.

Sample Output

*Section Title:* TTTTTTTTTTTTT

1. Text: `PPPPPPPPPPPPPPP` | URL: `FullStory.aspx?gid=4&id=6516`
2. Text: `CCCCCCCCCCCCC` | URL: `http://register1.sanjesh.org/fanni99up`
3. Text: `ZZZZZZZZ` | URL: `http://www6.sanjesh.org/download/fani99/FaniNote99.pdf`
4. Text: `FFFFFF` | URL: `FullStory.aspx?gid=4&id=6509`
5. Text: `XXXXXXXXXXXXXX` | URL: `FullStory.aspx?gid=4&id=6498`

--------------------------------------------------

*Section Title:* CCCCCCCCCCCC

1. Text: `GGGGGGGGGGGGGGGG` | URL: `http://register1.sanjesh.org/rgempiactax99/`
2. Text: `FFFFFFFFF` | URL: `FullStory.aspx?gid=11&id=6515`
3. Text: `HHHHHHHHH` | URL: `http://register2.sanjesh.org/RGKhanevadehConsult/`
4. Text: `FFFFFFFF` | URL: `FullStory.aspx?gid=11&id=6512`
5. Text: `NNNNNNNNNNNNNNNNNNNNNNNNNN` | URL: `FullStory.aspx?gid=11&id=6505`
6. Text: `OOOOOOOOOOOO` | URL: `http://dl.sanjesh.org/NOETDownload/DownloadHandler.ashx?id=1271`
7. Text: `JJJJJJJ` | URL: `FullStory.aspx?gid=11&id=6480`

--------------------------------------------------

内容的提问来源于stack exchange,提问作者jack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 22:42:32