基于XPath与Python lxml提取HTML中h4及关联文本、链接的技术咨询
Grouped Extraction of h4 Titles, Link Texts, and URLs
Your current script pulls all elements separately, but to pair each h4 with its corresponding links and text content, we need to process each topLinks container individually. Here's the corrected approach using XPath to group the content properly:
Corrected Python Script
from lxml.html import fromstring url=''' <div class="topLinks"> <div class="hd left"> </div><div class="hd-middle middle"> <h4>TTTTTTTTTTTTT</h4></div><div class="hd right"></div><div class="boxMiddle"><ul><li><a href="FullStory.aspx?gid=4&id=6516" title="1399/03/18" target="_blank">PPPPPPPPPPPPPPP<img class="new" src="images/new.png"></a></li><li><a href="http://register1.sanjesh.org/fanni99up" title="1399/03/11" target="_blank">CCCCCCCCCCCCC</a></li><li><a href="http://www6.sanjesh.org/download/fani99/FaniNote99.pdf" title="1399/03/11" target="_blank"> ZZZZZZZZ </a></li><li><a href="FullStory.aspx?gid=4&id=6509" title="1399/03/11" target="_blank">FFFFFF</a></li><li><a href="FullStory.aspx?gid=4&id=6498" title="1399/02/21" target="_blank">XXXXXXXXXXXXXX </a></li></ul></div><div class="boxBottom"></div></div> <div class="topLinks"><div class="hd left_alter"></div><div class="hd-middle middle_alter"> <h4>CCCCCCCCCCCC</h4></div><div class="hd right_alter"></div><div class="boxMiddle_alter"><ul><li><a href="http://register1.sanjesh.org/rgempiactax99/" title="1399/03/18" target="_blank">GGGGGGGGGGGGGGGG <img class="new" src="images/new.png"></a></li><li><a href="FullStory.aspx?gid=11&id=6515" title="1399/03/18" target="_blank">FFFFFFFFF<img class="new" src="images/new.png"></a></li><li><a href="http://register2.sanjesh.org/RGKhanevadehConsult/" title="1399/03/12" target="_blank">HHHHHHHHH</a></li><li><a href="FullStory.aspx?gid=11&id=6512" title="1399/03/12" target="_blank">FFFFFFFF</a></li><li><a href="FullStory.aspx?gid=11&id=6505" title="1399/02/24" target="_blank">NNNNNNNNNNNNNNNNNNNNNNNNNN</a></li><li><a href="http://dl.sanjesh.org/NOETDownload/DownloadHandler.ashx?id=1271" title="1398/12/12" target="_blank">OOOOOOOOOOOO</a></li><li><a href="FullStory.aspx?gid=11&id=6480" title="1399/01/26" target="_blank">JJJJJJJ</a></li></ul></div><div class="boxBottom_alter"></div></div> ''' tree = fromstring(url) # Iterate over each topLinks container to group content for section in tree.xpath("//div[@class='topLinks']"): # Extract the h4 title for this section h4_title = section.xpath(".//h4/text()")[0].strip() print(f"*Section Title:* {h4_title}\n") # Get all link elements within the current section link_elements = section.xpath(".//a") for index, link in enumerate(link_elements, 1): # Clean up link text and extract URL link_text = link.xpath("text()")[0].strip() link_url = link.xpath("@href")[0] print(f"{index}. Text: `{link_text}` | URL: `{link_url}`") # Add a separator between sections for readability print("\n" + "-"*50 + "\n")
Key Improvements Explained
- Grouped Processing: By looping over each
topLinksdiv and using.//in XPath (the dot restricts searches to the current section), we ensure links are paired with their correcth4title. - Whitespace Cleaning: Using
.strip()removes extra spaces from text elements for cleaner output. - Readable Formatting: The output clearly maps each section title to its associated links and text.
Sample Output
*Section Title:* TTTTTTTTTTTTT 1. Text: `PPPPPPPPPPPPPPP` | URL: `FullStory.aspx?gid=4&id=6516` 2. Text: `CCCCCCCCCCCCC` | URL: `http://register1.sanjesh.org/fanni99up` 3. Text: `ZZZZZZZZ` | URL: `http://www6.sanjesh.org/download/fani99/FaniNote99.pdf` 4. Text: `FFFFFF` | URL: `FullStory.aspx?gid=4&id=6509` 5. Text: `XXXXXXXXXXXXXX` | URL: `FullStory.aspx?gid=4&id=6498` -------------------------------------------------- *Section Title:* CCCCCCCCCCCC 1. Text: `GGGGGGGGGGGGGGGG` | URL: `http://register1.sanjesh.org/rgempiactax99/` 2. Text: `FFFFFFFFF` | URL: `FullStory.aspx?gid=11&id=6515` 3. Text: `HHHHHHHHH` | URL: `http://register2.sanjesh.org/RGKhanevadehConsult/` 4. Text: `FFFFFFFF` | URL: `FullStory.aspx?gid=11&id=6512` 5. Text: `NNNNNNNNNNNNNNNNNNNNNNNNNN` | URL: `FullStory.aspx?gid=11&id=6505` 6. Text: `OOOOOOOOOOOO` | URL: `http://dl.sanjesh.org/NOETDownload/DownloadHandler.ashx?id=1271` 7. Text: `JJJJJJJ` | URL: `FullStory.aspx?gid=11&id=6480` --------------------------------------------------
内容的提问来源于stack exchange,提问作者jack
相关产品推荐
相关产品推荐

