You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python按指定格式抓取HTML表格并导出为CSV文件

Fixing Your HTML Table to CSV Extraction Code

Let's break down the issues in your code and fix them step by step:

Key Problems in Your Current Code

  • XPath context issues: Using // in XPath inside loops will search the entire document instead of the current row (<tr>) you're iterating over. This means you're pulling all rows' data every time instead of just the current one.
  • Incorrect loop structure: You're looping over all <td> elements after grabbing all tables, which leads to redundant and incorrect data collection.
  • Wrong data aggregation: You're appending each field's list to results and then writing each list as a separate row, instead of combining them into a single row per table row.
  • Missing header row: You never write the required CSV header (Link,name,brand,description).

Corrected Code

# First, write the CSV header
writer.writerow(["Link", "name", "brand", "description"])

# Iterate over each table row in the tbody (skip the header row)
for tr in doc.xpath("//table/tbody/tr"):
    # Get the link from the first td's anchor (using relative XPath with .// to stay in current tr)
    link = tr.xpath(".//td[1]/a/@href")[0]
    # Get text content for name, brand, description (again, relative to current tr)
    name = tr.xpath(".//td[1]//text()")[0].strip()
    brand = tr.xpath(".//td[2]//text()")[0].strip()
    description = tr.xpath(".//td[3]//text()")[0].strip()
    
    # Combine into a single row and write to CSV
    writer.writerow([link, name, brand, description])

Explanation of Changes

  1. Write the header first: We start by writing the required CSV header row to match your desired output.
  2. Target specific rows: We directly loop over <tr> elements inside the table's <tbody> to process each data row individually.
  3. Relative XPath: Using .// instead of // ensures we only select elements within the current <tr> we're processing, so we get the correct data for each row.
  4. Extract single values: Since each <td> has exactly one text value, we access the first element of the XPath result list ([0]) and use .strip() to clean up any extra whitespace.
  5. Single row per table row: We combine all four fields into a single list and write it as one CSV row, which matches your desired output format.

This code will generate the CSV exactly as you specified.

内容的提问来源于stack exchange,提问作者Riyas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:07:52