You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现产品数据文件对比并生成新增产品文件的问题求助

Python实现产品数据文件对比并生成新增产品文件的问题求助

我刚接触Python,正在开发一个小工具,功能是拿一份从URL下载的产品信息主文件,和本地的旧数据文件做对比——核心是通过SKU列找出主文件里的新增产品,然后生成两个文件:一个只包含新增SKU的列表,另一个保留主文件所有列但只放新增产品的行。

文件名说明

  • downloaded_data.csv:从URL下载的主文件,包含完整的产品信息
  • current-data.csv:用来对比的旧数据文件,以此为基准找新增产品
  • new-products.csv:最终生成的文件,要保留主文件的所有列,但仅包含旧文件里没有的产品行
  • comparison-file.csv:辅助文件,只输出所有新增的SKU列表

示例:downloaded_data.csv的列结构(新文件需完全保留这些列)

TitleImage URLSKU
Product Name Oneimageurl.jpgSKU123
Product Name Twoimageurl.jpgSKU456

示例:current-data.csv的列结构(通过SKU/ID列做匹配)

IDImage URLProduct Name
SKU123imageurl.jpgProduct Name One
SKU789imageurl.jpgProduct Name Three

示例:预期生成的new-products.csv

TitleImage URLSKU
Product Name Threeimageurl.jpgSKU789

我写了下面的Python代码,但运行后发现new-products.csv直接是主文件的副本,完全没有过滤出新增产品。需要说明的是,旧数据文件的列和主文件不一样,列顺序也不同,我只打算用SKU作为唯一匹配标识。有没有大佬能帮我看看问题出在哪,或者给点改进方向呀?谢谢啦!

import pandas as pd
import requests

pd.set_option("display.max_rows", None)

# Dowload CSV file
url = "URL GOES HERE"
response = requests.get(url)

# Check if the request was successful (status code 200)
if response.status_code == 200:
    # Save the content of the response to a local CSV file
    with open("downloaded_data.csv", "wb") as f:
        f.write(response.content)
    print("CSV file downloaded successfully")
else:
    print("Failed to download csv file. Status code: ", response.status_code)

# Read the CSV file into a Pandas DataFrame
master_file_compare = pd.read_csv("downloaded_data.csv", usecols=[29], names=['SKU'])
account_data = pd.read_csv("account-data.csv", usecols=[5], names=['SKU'])

# Merging both dataframe using left join
comparison_result = pd.merge(master_file_compare,account_data, on='SKU', how='left', indicator=True)

# Filtering only the rows that are available in left (master_file_compare)
comparison_result = comparison_result.loc[comparison_result['_merge'] == 'left_only']

comparison_result.to_csv('comparison-file.csv', encoding='utf8')

# Compare Comparison file to master and generate data file
with open('downloaded_data.csv', 'r', encoding='utf8') as in_file, open('new-products.csv', 'w', encoding='utf8') as out_file:
    for line in in_file:
        if line.split(',')[0].strip() not in comparison_result:
            out_file.write(line)

# print(account_data.head)
print(comparison_result)

备注:内容来源于stack exchange,提问作者Calogero Savarino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 15:34:50