使用lxml抓取局域网网站计数器:特定数字提取问题求助
如何用Python lxml提取HTML表格中特定分类的计数器数据?
我正在写脚本自动化抓取局域网网站上的计数器数据,现在卡住了。目标是提取表格里1-1、2-4等固定分类右侧的第一个数字,最终输出格式是task - counter 1-1 = 6490 2-4 = 442这样的。
网站上的HTML表格结构如下:
<TR><td><p align="left" style="margin-left: 30;"><b>title</b></p></td><td><p> </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;">table one</p></td><td><p> Task average </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;"></p></td><td><p> number number </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;">1-1 C</p></td><td><p> 6490 1 </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;">2-4 C</p></td><td><p> 442 2 </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;">5-10 C</p></td><td><p> 44 6 </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;">11-20 C</p></td><td><p> 3 15 </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;">21-30 C</p></td><td><p> 2 25 </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;">31-50 C</p></td><td><p> 1 40 </p></td> </TR> <TR><td><p align="left" style="margin-left: 40;">sum</p></td><td><p> 6982 1 </p></td> </TR>
我现在用的代码是这样的:
import requests from lxml import html pageContent=requests.get( 'http://x.html') tree = html.fromstring(pageContent.content) scraped = tree.xpath('//p/text()') print(scraped)
但执行后输出了一大堆无关的内容,试了其他方法也没成功,求帮忙解决!
解决方案
你的问题其实是出在太宽泛的XPath上——//p/text()会把页面里所有<p>标签的文本都抓下来,自然会混进一堆无关内容。咱们换个思路,精准定位到包含目标分类的行,再提取对应数字就行。
具体实现思路
- 锁定目标行:找到
<tr>中第一个<td>里的<p>文本包含1-1、2-4这类分类的行; - 提取计数器数字:对每个目标行,取第二个
<td>里的<p>文本,清理后分割出第一个有效数字; - 格式化输出:按照你要的格式拼接结果。
完整可运行代码
import requests from lxml import html # 定义你需要抓取的目标分类,后续要加新分类直接往列表里加就行 target_categories = ["1-1", "2-4"] page_content = requests.get('http://x.html') tree = html.fromstring(page_content.content) result_parts = [] for category in target_categories: # 用XPath精准定位到包含目标分类的行,再取对应列的文本 # 用contains是因为<p>里除了分类还有"C"和空格,模糊匹配更稳妥 row_text = tree.xpath(f'//tr[td/p[contains(text(), "{category}")]]/td[2]/p/text()') if row_text: # 清理文本里的空格和换行,分割后取第一个元素就是目标数字 clean_text = row_text[0].strip() counter_num = clean_text.split()[0] result_parts.append(f"{category} = {counter_num}") # 拼接成你想要的输出格式 final_output = f"task - counter {' '.join(result_parts)}" print(final_output)
代码细节说明
- XPath定位逻辑:
//tr[td/p[contains(text(), "{category}")]]会过滤出所有符合条件的行,避免抓取表头、总和这些无关行; - 文本处理:
strip()去掉文本前后的空格,split()按任意空格分割后取第一个元素,完美避开后面的average数值; - 扩展性:如果之后需要抓取
5-10这类其他分类,只需要在target_categories列表里添加对应的字符串即可。
如果你的局域网网站需要登录验证或者有反爬机制,可能还需要在requests.get里添加headers或者cookies参数,但根据你提供的信息,上面的代码应该能直接解决当前的问题。
内容的提问来源于stack exchange,提问作者Gadzin
相关产品推荐
相关产品推荐

