You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame上的Lambda Function无法运行问题求助

问题:Lambda函数结合自定义urltotext提取PDF文本失败

问题分析

你当前代码中result.assign(Content = lambda x: urltotext(x['Source']))无法正常工作,核心原因如下:

  • x['Source']传入的是整个Pandas Series(整列数据),但urltotext函数仅能处理单个URL字符串,无法批量处理数据。
  • 函数中直接写入本地文件DailyCA.pdf,处理多个链接时会导致文件被重复覆盖,多线程场景下还会引发IO冲突。
  • 异常处理过于笼统,无法定位具体错误(如网络请求失败、PDF解析错误等)。

修复方案

1. 替换批量处理方式

将assign的lambda写法改为apply,对Source列的每个元素单独调用urltotext:

result['Content'] = result['Source'].apply(urltotext)

2. 优化urltotext函数(避免本地文件IO)

使用内存字节流替代本地文件,既提升效率又避免文件冲突,同时优化异常信息输出:

from io import BytesIO

def urltotext(link):
    try:
        # 复用已创建的缓存会话,保持请求一致性
        resp = s.get(link)
        resp.raise_for_status()  # 触发HTTP错误抛出
        # 直接用内存字节流读取PDF,无需写入本地文件
        doc = fitz.open("pdf", BytesIO(resp.content)) 
        text = [page.get_text('text') for page in doc]
        doc.close()
        return 'shodhpage'.join(text)
    except Exception as e:
        # 打印错误信息,方便调试具体链接的问题
        print(f"处理链接{link}时出错: {str(e)}")
        return "none"

3. 其他优化建议

  • 移除未使用的库:PyPDF2、multiprocessing(当前代码未用到)。
  • 复用requests_cache.CachedSession对象发起PDF请求,减少重复请求并保持会话一致性。

修改后的完整代码

import requests
import pandas as pd
from datetime import datetime
from datetime import date
import json
import fitz
import requests_cache 
from io import BytesIO


def urltotext(link):
    try:
        # 复用缓存会话
        resp = s.get(link)
        resp.raise_for_status()
        doc = fitz.open("pdf", BytesIO(resp.content)) 
        text = [page.get_text('text') for page in doc]
        doc.close()
        return 'shodhpage'.join(text)
    except Exception as e:
        print(f"处理链接{link}时出错: {str(e)}")
        return "none"


def all():
    global s  # 让urltotext可以复用会话对象
    print("Started Pulling")
    currentd = date.today()
    s = requests_cache.CachedSession('demo_cache', backend='sqlite')
    headers =   {'Host':'www.nseindia.com','User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:82.0) Gecko/20100101 Firefox/82.0','Accept':'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language':'en-US,en;q=0.5', 'Accept-Encoding':'gzip, deflate, br','DNT':'1', 'Connection':'keep-alive', 'Upgrade-Insecure-Requests':'1','Pragma':'no-cache','Cache-Control':'no-cache',  }
    url = 'https://www.nseindia.com/'
    step = s.get(url,headers=headers)
    today = datetime.now().strftime('%d-%m-%Y')
    api_url = f'https://www.nseindia.com/api/corporate-announcements?index=equities&from_date=01-01-2022&to_date=18-08-2022'
    resp = s.get(api_url,headers=headers).json()
    print("API Read")
    result = pd.DataFrame(resp)
    result.drop(['difference', 'dt','exchdisstime','csvName','old_new','orgid','seq_id','bflag','symbol','sort_date'], axis = 1, inplace = True)
    result.rename(columns = {'an_dt':'DateandTime', 'attchmntFile':'Source','attchmntText':'Topic','desc':'Type','smIndustry':'Sector','sm_name':'Company Name','sm_isin':'ISIN'}, inplace = True)
    result[['Date','Time']] = result.DateandTime.str.split(expand=True)
    result = result[result['Type'].str.contains("Loss of Share Certificates|Copy of Newspaper Publication") == False]
    result['Type'] = result['Type'].astype(str)
    result['Type'].replace("Certificate under SEBI (Depositories and Participants) Regulations, 2018",'Junk' , inplace = True)
    result = result[result['Type'].str.contains("Junk") == False]
    result = result[result["Type"].str.contains("Trading Window") == False]
    result = result[result["Type"].str.contains("Loss of share certificate") == False]
    result = result[result["Type"].str.contains("Loss of share certificates") == False]
    result = result[result["Type"].str.contains("Disclosure under SEBI Takeover Regulations") == False]
    result = result[result["Type"].str.contains("Newspaper Advertisements") == False]
    result = result[result["Type"].str.contains("-") == False]
    result.drop_duplicates(subset='Source', keep = 'first', inplace = True)
    result['Temporary']=pd.to_datetime(result['Date']+' '+result['Time'])
    result['Date']=result['Temporary'].dt.strftime('%b %d, %Y')
    result['Time']=result['Temporary'].dt.strftime('%R %p')
    result['DateTime'] = result['Temporary'].dt.strftime('%m/%d/%Y %I:%M %p')
    result.drop(['DateandTime', 'Temporary'], axis = 1, inplace = True)
    # 替换为apply处理每个URL
    result['Content'] = result['Source'].apply(urltotext)
    result.to_csv("2018-Test.csv")

all()

关键说明

  • apply方法确保每个URL被单独传入urltotext处理,匹配函数的单参数设计。
  • 内存字节流BytesIO避免了本地文件的写入和覆盖问题,大幅提升处理速度。
  • 复用requests_cache会话,减少重复请求,同时保持请求头一致性,降低被目标网站拦截的概率。
  • 新增错误打印,方便快速排查具体链接的处理失败原因。

内容的提问来源于stack exchange,提问作者Jay shankarpure

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 02:24:39