Scrapy/Python如何处理缺失闭合TR/TD标签的网页表格?
Hey there! I’ve dealt with way too many wonky HTML tables missing closing <tr>/<td> tags, so I totally get the frustration of moving from a quick JS hack to a more robust Scrapy setup. Let’s break down your options, starting with the easiest (and most maintainable) approaches first:
1. Let Scrapy’s Selectors Do the Heavy Lifting (Most Recommended)
Scrapy uses lxml under the hood, which has a super forgiving HTML parser that automatically fixes unclosed tags and builds a valid DOM tree—even for garbage HTML. You might not need any manual splitting at all!
Here’s how to test this quickly:
- Fire up the Scrapy shell with your target URL:
scrapy shell https://your-target-url.com - Try selecting rows first:
rows = response.xpath('//table//tr') # Or using CSS selectors: rows = response.css('table tr') - For each row, extract the cells:
for row in rows: cells = row.xpath('.//td/text()').getall() # Clean up whitespace if needed: cells = [cell.strip() for cell in cells if cell.strip()] print(cells)
If this works, you’re golden! The parser will handle all the unclosed tags silently, just like a browser does.
Pro Tip: Use html5lib for Extra Tolerance
If lxml’s default parser still struggles with extremely broken HTML, install html5lib (a parser that mimics browser behavior perfectly) and specify it when creating your selector:
from scrapy.selector import Selector # Parse with html5lib instead of lxml sel = Selector(text=response.text, type='html', parser='html5lib') rows = sel.xpath('//table//tr')
This is overkill for most cases, but it’s a lifesaver for truly mangled HTML.
2. When You Do Need to Split (Last Resort)
If for some reason the selector approach fails, avoid splitting the entire response.data directly—this is error-prone (e.g., <tr> might appear in scripts, comments, or cell text). Instead:
- First extract only the table’s HTML content to limit your scope:
table_html = response.xpath('//table').get() if not table_html: # Handle missing table case pass - Split using a regex to account for
<tr>tags with attributes (like<tr class="odd">):import re # Split on any <tr> tag (with or without attributes) row_chunks = re.split(r'<tr[^>]*>', table_html)[1:] # Skip the empty first chunk - Then process each row chunk to extract cells:
for chunk in row_chunks: # Split cells, clean up tags and whitespace cells = [re.sub(r'<.*?>', '', cell).strip() for cell in re.split(r'<td[^>]*>', chunk)[1:] if cell.strip()] print(cells)
Note on response.data.split()
You mentioned confusion about this—response.data returns raw bytes, so splitting it directly would require decoding first (e.g., response.data.decode('utf-8').split(...)). It’s always better to use response.text instead, since it handles encoding automatically based on the response headers.
Final Takeaway
Stick with Scrapy’s selectors first—they’re designed to handle messy HTML, and your code will be way easier to maintain than manual string splitting. Only fall back to regex/split when absolutely necessary.
内容的提问来源于stack exchange,提问作者Peter Nguyen

