You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多CSV处理遇Pandas设置警告与解析错误,求解决方案

问题解决方案

1. 解决Setting with Copy警告

这个警告的核心原因是df_filtered是原DataFrame的切片视图,而非独立的副本,直接给它新增列会触发pandas的视图/副本歧义检查。

修复方法:过滤数据时直接生成独立副本即可:

# 替换原过滤代码
df_filtered = df[df['date'] >= last_month].copy()

或者用loc语法明确生成副本,可读性更强:

df_filtered = df.loc[df['date'] >= last_month, :].copy()

后续再给df_filtered['author']赋值时,就不会触发警告了。

2. 解决ParserError解析错误

这个错误是因为部分CSV文件格式不规范——比如某行的字段数量和表头不匹配,或者字段内容里包含了未转义的分隔符(比如逗号)。

可以通过以下几种方式针对性修复:

  • 指定正确分隔符:如果你的CSV不是用逗号分隔(比如制表符),显式指定sep参数:
    df = pd.read_csv(file_path, sep='\t')
    
  • 切换Python解析引擎:Python引擎比默认的C引擎对不规范格式的容忍度更高:
    df = pd.read_csv(file_path, engine='python')
    
  • 跳过错误行:如果只是少数行格式错误不影响整体数据,直接跳过这些行:
    df = pd.read_csv(file_path, on_bad_lines='skip')
    
  • 处理带引号的字段:如果字段内容包含分隔符(比如"Smith, John"),指定quotechar确保正确解析:
    df = pd.read_csv(file_path, quotechar='"')
    

完整修复后的关键代码片段

for filename in os.listdir(input_path):
    file_path = os.path.join(input_path, filename)
    if os.path.isfile(file_path):
        # 修复ParserError的read_csv参数组合
        df = pd.read_csv(file_path, engine='python', on_bad_lines='skip')

        df['date'] = pd.to_datetime(df['date'], errors='coerce').dt.tz_localize(None)
        last_month = last_month.replace(tzinfo=None)

        # 修复Setting with Copy警告,生成独立副本
        df_filtered = df[df['date'] >= last_month].copy()

        articles = df_filtered['url'].tolist()
        authors = []
        for url in articles:
            print(f"Getting data from {url}")
            headers = {'User-Agent': "Mozilla/5.0 (Macintosh; Intel Mac OS X 14.5; rv:128.0) Gecko/20100101 Firefox/128.0"}
            response = requests.get(url, headers=headers)

            if response.status_code == 200:
                content = response.text
            else:
                print(f"Failed to fetch article from {url}")
                authors.append("none")
                continue

            soup = BeautifulSoup(content, 'html.parser')
            author = get_author(soup)
            print(f"Author: {author}")
            authors.append(author)

        df_filtered['author'] = authors

        output_file_path = os.path.join(output_path, filename)
        df_filtered.to_csv(output_file_path, index=False)

内容的提问来源于stack exchange,提问作者Andy C.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 19:51:02