You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python无法从utils导入指定函数如何解决?(关联nltk)

解决Jupyter Notebook中ImportError: 无法从utils导入process_tweet的问题

问题复现

在Google Colab和Jupyter Notebook中测试代码时,Jupyter Notebook抛出以下错误:

---------------------------------------------------------------------------
ImportError                               Traceback (most recent call last)
Cell In[5], line 1
----> 1 from utils import process_tweet, build_freqs

ImportError: cannot import name 'process_tweet' from 'utils' (C:\Users\Gabriel\anaconda3\Lib\site-packages\utils\__init__.py)

用户运行的代码如下:

import nltk                                  
from nltk.corpus import twitter_samples      
import matplotlib.pyplot as plt              
import numpy as np                           

nltk.download('twitter_samples')
nltk.download('stopwords')

!pip install utils 
from utils import process_tweet, build_freqs

解决方案

  • 卸载第三方utils包:你通过pip install utils安装的是通用工具库,其中并不包含process_tweet和build_freqs这两个自定义函数,先卸载这个包:
    pip uninstall -y utils
    
  • 创建自定义utils.py文件:这两个函数是NLP教学场景(比如Coursera自然语言处理专项课程)中常用的自定义工具函数,你需要在当前Jupyter Notebook的工作目录下新建一个utils.py文件,将以下代码复制进去:
    import re
    import string
    import numpy as np
    
    from nltk.corpus import stopwords
    from nltk.stem import PorterStemmer
    from nltk.tokenize import TweetTokenizer
    
    def process_tweet(tweet):
        """处理推文文本:分词、去停用词、词干提取等"""
        stemmer = PorterStemmer()
        stopwords_english = stopwords.words('english')
        # 移除@提及、链接、#符号
        tweet = re.sub(r'^RT[\s]+', '', tweet)
        tweet = re.sub(r'https?://[^\s\n\r]+', '', tweet)
        tweet = re.sub(r'#', '', tweet)
        # 分词
        tokenizer = TweetTokenizer(preserve_case=False, strip_handles=True, reduce_len=True)
        tweet_tokens = tokenizer.tokenize(tweet)
    
        tweets_clean = []
        for word in tweet_tokens:
            if (word not in stopwords_english and  
                word not in string.punctuation): 
                stem_word = stemmer.stem(word)  # 词干提取
                tweets_clean.append(stem_word)
    
        return tweets_clean
    
    def build_freqs(tweets, ys):
        """构建词频字典:键为(词, 标签),值为出现次数"""
        yslist = np.squeeze(ys).tolist()
        freqs = {}
        for y, tweet in zip(yslist, tweets):
            for word in process_tweet(tweet):
                pair = (word, y)
                if pair in freqs:
                    freqs[pair] += 1
                else:
                    freqs[pair] = 1
        return freqs
    
  • 验证导入:确保utils.py和你的Jupyter Notebook文件在同一个文件夹下,重新运行导入语句即可正常使用这两个函数。

内容的提问来源于stack exchange,提问作者philosophy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:02:52