如何获取通过XPath定位的单个<tr>元素的原始HTML
Got it, let's break this down for you. You've already pulled all the <tr> elements into the tr_value list—great start! To grab the full HTML of a single <tr> element, the exact approach depends on which library you're using (lxml and Scrapy are the most common tools for this kind of work in Python). Here's how to do it in both cases:
Using lxml
Once you've picked your target <tr> element (for example, the first one in the list), you can use lxml's etree.tostring() function to convert it to an HTML string. Here's the code:
from lxml import etree # Grab the first <tr> element from your list single_tr = tr_value[0] # Convert the element to its full HTML string tr_html = etree.tostring(single_tr, encoding='unicode', method='html') # Print or use the HTML as needed print(tr_html)
encoding='unicode'ensures you get a regular string instead of raw bytes.method='html'makes sure the output is proper HTML (not XML, which might have self-closing tags that don't match standard HTML).
Using Scrapy
Scrapy's Selector objects have a couple of handy ways to get an element's outer HTML. Here are two reliable methods:
Method 1: Use the outer-html() XPath function
This is the most straightforward way in Scrapy:
# Get the first <tr> element's full HTML single_tr_html = tr_value[0].xpath('outer-html()').get()
Method 2: Access the underlying lxml element
If you prefer using lxml's tostring() (like the example above), you can access the raw lxml element via the root attribute:
from lxml import etree # Grab the first <tr> and convert to HTML single_tr_html = etree.tostring(tr_value[0].root, encoding='unicode', method='html')
And just a quick note: if you want a different <tr> element (not the first one), just adjust the index—for the second element, use tr_value[1], third uses tr_value[2], and so on (since Python lists are zero-indexed).
内容的提问来源于stack exchange,提问作者Nirdesh Kumar

