You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Data Factory自定义活动中用Dedupe遇TypedDict导入错误求助

ImportError: cannot import name 'TypedDict' when using dedupe in Azure Data Factory Custom Activity

Problem Description

我在Azure Data Factory中创建了一个自定义活动,已经配置好Batch账户和池,简单的Python代码能成功运行,但运行以下去重代码时,在import dedupe语句处抛出ImportError: cannot import name 'TypedDict'错误。

My Code

import os
import csv
import re
import logging
import optparse
import pandas as pd
import numpy as np
import dedupe
import pickle
if __name__ == '__main__':
    optp = optparse.OptionParser()
    optp.add_option('-v', '--verbose', dest='verbose', action='count', help='Increase verbosity (specify multiple times for more)' )
    (opts, args) = optp.parse_args()
    log_level = logging.WARNING
    if opts.verbose:
        if opts.verbose == 1:
            log_level = logging.INFO
        elif opts.verbose >= 2:
            log_level = logging.DEBUG
    logging.getLogger().setLevel(log_level)
    input_file = 'testfile.csv'
    output_file = 'out.csv'
    settings_file = 'settings'
    training_file = 'trainingfile.json'
    print('importing data ...')
    data = pd.read_csv(input_file)
    data.set_index('Id')
    data = data.where(pd.notnull(data), None)
    data_d = data.to_dict('index')
    if os.path.exists(settings_file):
        print('reading from', settings_file)
        with open(settings_file, 'rb') as f:
            deduper = dedupe.StaticDedupe(f)
    else:
        fields = [
            {'field' : 'fname', 'type': 'String'},
            {'field' : 'lname', 'type': 'ShortString','has missing' : True},
            {'field' : 'rnamr', 'type': 'String', 'has missing' : True},
            {'field' : 'add', 'type': 'String', 'has missing' : True},
            {'field' : 'username', 'type': 'ShortString', 'has missing' : True},
            {'field' : 'namecommon', 'type': 'String', 'has missing' : True},
            {'field' : 'wenurl', 'type': 'String', 'has missing' : True},
            {'field' : 'country', 'type': 'ShortString', 'has missing' : True},
        ]
        print('Creating deduper')
        deduper = dedupe.Dedupe(fields)

Full Error Traceback

Traceback (most recent call last):
  File "main.py", line 8, in <module>
    import dedupe
  File "C:\Anaconda\lib\site-packages\dedupe\__init__.py", line 6, in <module>
    from dedupe._init import * # noqa
  File "C:\Anaconda\lib\site-packages\dedupe\_init.py", line 1, in <module>
    from dedupe.api import StaticDedupe, Dedupe # noqa: F401
  File "C:\Anaconda\lib\site-packages\dedupe\api.py", line 21, in <module>
    import dedupe.core as core
  File "C:\Anaconda\lib\site-packages\dedupe\core.py", line 22, in <module>
    from dedupe._typing import (RecordPairs, RecordID, Blocks, Data, Literal)
  File "C:\Anaconda\lib\site-packages\dedupe\_typing.py", line 15, in <module>
    from typing_extensions import TypedDict, Protocol, Literal
ImportError: cannot import name 'TypedDict'

Solution

这个问题我之前也碰到过,核心原因是你的Python环境里的typing_extensions库版本不兼容或者安装不完整。dedupe库需要从typing_extensions中导入TypedDict类,但你的环境里的typing_extensions版本过低(低于3.7.4.3,这是TypedDict被加入到typing_extensions的最低版本),或者安装过程中出现了问题。

你可以按照以下步骤解决:

  1. 检查当前typing_extensions版本
    在你的Azure Batch池的Python环境中运行以下命令,确认版本:

    pip show typing_extensions
    

    如果版本低于3.7.4.3,就需要升级。

  2. 升级typing_extensions到兼容版本
    使用pip升级:

    pip install --upgrade typing_extensions
    

    如果你用的是Anaconda环境,可以用conda命令:

    conda install -c conda-forge typing_extensions --upgrade
    
  3. 确保Batch池所有节点环境一致
    因为Azure Batch池可能有多个节点,你需要保证所有节点都完成了版本升级。最简单的方式是在自定义活动的启动脚本中加入上述升级命令,这样每次运行代码前都会自动检查并升级依赖。

  4. (可选)升级dedupe库到最新稳定版
    如果升级typing_extensions后还是报错,可能是dedupe版本和依赖库不兼容,尝试升级dedupe:

    pip install --upgrade dedupe
    

完成以上步骤后,重新运行你的去重代码,应该就能解决这个导入错误了。


内容的提问来源于stack exchange,提问作者Sindu_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 09:12:39