如何使用Python解析网页表格?已获取table_rows但无法提取数据
提取网页表格数据的解决方案
嘿,看起来你已经搞定了最棘手的定位环节——找到目标表格并拿到所有行数据了!接下来提取具体数值其实很简单,我给你整理了完整的代码和步骤:
完整代码示例
options = webdriver.ChromeOptions() options.add_argument('headless') driver = webdriver.Chrome(r'C:\Users\wendle\BrowserDrivers\chromedriver.exe', options=options) driver.get(wsr) html = driver.page_source soup = BeautifulSoup(html, 'lxml') table = soup.find_all('table') table_rows = table[2].find_all('tr') # 初始化列表存储提取后的结构化数据 tool_records = [] # 遍历数据行:跳过前两行(第一行是DIFFUSION标题栏,第二行是表头) for row in table_rows[2:]: # 获取当前行的所有单元格 cells = row.find_all('td') # 逐个提取字段文本,strip()清理多余空格和换行 tool_id = cells[0].get_text(strip=True) tool_type = cells[1].get_text(strip=True) status = cells[2].get_text(strip=True) date_time = cells[3].get_text(strip=True) minutes = cells[4].get_text(strip=True) employee = cells[5].get_text(strip=True) comments = cells[6].get_text(strip=True) # 将数据存入字典,结构清晰方便后续处理 record = { 'ToolId': tool_id, 'Type': tool_type, 'Status': status, 'Date/Time': date_time, 'Min': minutes, 'Employee': employee, 'Comments': comments } tool_records.append(record) # 打印验证提取结果 for entry in tool_records: print(entry)
关键步骤说明
- 跳过非数据行:
table_rows[2:]是因为你的表格前两行是标题栏和表头,从第三行开始才是实际的工具数据,直接跳过这两行就能精准抓取目标内容。 - 提取单元格文本:用
find_all('td')拿到当前行的所有单元格,再通过索引cells[0]到cells[6]对应每个字段,get_text(strip=True)会自动去掉文本前后的空格、换行符,让数据更整洁。 - 结构化存储:把每一行数据存入字典再添加到列表,后续不管是导出到Excel、CSV还是做数据分析,都能轻松处理。
如果之后需要调整字段或者处理更复杂的表格结构,随时调整就行!
内容的提问来源于stack exchange,提问作者james wendle
相关产品推荐
相关产品推荐

