You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为爬取到的表格数据新增对应内部ID列?

实现方法

核心逻辑很简单:在遍历ID的循环层内拿到当前迭代的ID值,等当前ID对应的页面表格解析完成后,直接给解析出的所有行统一追加ID字段即可,不需要做复杂的参数传递。
下面是两种最常用实现的代码示例,你可以根据自己的技术栈选:

基于Pandas处理表格(写起来最省事)

如果你平时用Pandas做表格数据处理,不需要逐行循环加ID,整表赋值新列就能自动给所有行带上当前ID:

import pandas as pd
import requests

# 你的内部ID遍历列表
internal_ids = [2301, 2302, 2303, 2304]
crawl_results = []

for aid in internal_ids:
    # 替换成你自己的URL拼接规则
    target_url = f"https://your-target-site.com/detail/{aid}"
    resp = requests.get(target_url, timeout=10)
    # 解析当前页的目标表格,按实际情况选对应的表格索引
    page_table = pd.read_html(resp.text)[0]
    # 直接新增ID列,当前页所有行自动绑定当前循环的ID
    page_table["source_internal_id"] = aid
    crawl_results.append(page_table)

# 合并所有页的数据导出即可
final_data = pd.concat(crawl_results, ignore_index=True)
final_data.to_csv("crawl_output.csv", index=False, encoding="utf-8-sig")

原生解析不依赖Pandas的场景

如果你是用BeautifulSoup这类库逐行解析表格,每解析完一行数据,直接把当前ID塞进行数据里就行:

import requests
from bs4 import BeautifulSoup

internal_ids = [2301, 2302, 2303, 2304]
all_rows = []

for aid in internal_ids:
    target_url = f"https://your-target-site.com/detail/{aid}"
    resp = requests.get(target_url, timeout=10)
    soup = BeautifulSoup(resp.text, "html.parser")
    table = soup.find("table", class_="data-table") # 替换成你实际的表格定位规则
    # 提取表头
    header = [th.get_text(strip=True) for th in table.find_all("th")]
    # 遍历数据行
    for tr in table.find_all("tr")[1:]: # 跳过表头行
        cell_vals = [td.get_text(strip=True) for td in tr.find_all("td")]
        row = dict(zip(header, cell_vals))
        # 给当前行追加对应ID
        row["source_internal_id"] = aid
        all_rows.append(row)

踩坑提醒:ID赋值的逻辑一定要写在「遍历ID的for循环内部」、「遍历表格行的for循环外部」,不要写在最外层,不然会出现所有行都绑定最后一个ID、数据和ID错位的问题。

内容的提问来源于stack exchange,提问作者Chris G.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 17:18:42