You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用Pattern库ngrams方法触发UnboundLocalError问题求助

Fixing UnboundLocalError in Pattern's ngrams with PySpark

Let's break down why you're hitting this error and how to fix it quickly.

Root Cause

Looking at your code, you're passing a list of strings to Pattern's ngrams() function (since each value in your RDD is a list like [u'Campbell was born in Myrtle Bank.']). But if you check the ngrams() source code, it only handles three input types: basestring, Sentence, and Text.

When you pass a list, none of the isinstance() checks trigger, so the variable s never gets assigned. Then when the code reaches for s in s:, it throws an UnboundLocalError because s doesn't exist yet.

Solutions

Instead of passing a list of strings directly, join them into a single string first. This fits what ngrams() expects, and you won't need to modify the Pattern library.

Modify your business code like this:

corpus = text\
 .mapValues(lambda v: ngrams(' '.join(v), self.max_ngram))\  # Join list into one string
 .flatMap(lambda (target, tokens): (((target, t), 1) for t in tokens))\
 .reduceByKey(add)\
 .map(lambda ((target, token), count): (token, (target, count)))

This works because ' '.join(v) turns your list of sentences into a single string, which ngrams() can tokenize and process correctly—triggering the isinstance(string, basestring) branch and properly initializing s.

2. Patch Pattern's ngrams() Function

If you need to keep passing lists directly (e.g., to preserve sentence boundaries without joining), modify the ngrams() function to handle list inputs. Update the function in pattern/text/__init__.py like this:

def ngrams(string, n=3, punctuation=PUNCTUATION, continuous=False):
    """ Returns a list of n-grams (tuples of n successive words) from the given string. Alternatively, you can supply a Text or Sentence object. With continuous=False, n-grams will not run over sentence markers (i.e., .!?). Punctuation marks are stripped from words. """
    def strip_punctuation(s, punctuation=set(punctuation)):
        return [w for w in s if (isinstance(w, Word) and w.string or w) not in punctuation]
    if n <= 0:
        return []
    s = []  # Initialize s with default empty list
    if isinstance(string, basestring):
        s = [strip_punctuation(s.split(" ")) for s in tokenize(string)]
    elif isinstance(string, Sentence):
        s = [strip_punctuation(string)]
    elif isinstance(string, Text):
        s = [strip_punctuation(s) for s in string]
    elif isinstance(string, list):
        # Handle list of strings: process each element as a separate text segment
        for item in string:
            if isinstance(item, basestring):
                s.extend([strip_punctuation(split_item.split(" ")) for split_item in tokenize(item)])
    # Add an else clause if you want to handle other types gracefully
    if continuous:
        s = [sum(s, [])]
    g = []
    # Rename loop variable to avoid overwriting the s list
    for segment in s:
        g.extend([tuple(segment[i:i+n]) for i in range(len(segment)-n+1)])
    return g

Key changes here:

  • Initialize s as an empty list upfront to avoid undefined errors.
  • Add an elif isinstance(string, list) branch to process each string in your list.
  • Rename the loop variable from s to segment to prevent accidentally overwriting the s list (a subtle bug in the original code!).

3. Validate Your RDD Data

Double-check that your text processing pipeline is outputting the data you expect. Your initial text.map(...).groupByKey().map(...) produces tuples where the value is a list of strings—this is correct, but you just need to adapt it for ngrams() as shown above.

Why Your Previous Library Modification Failed

Chances are you either didn't add the list handling branch, or you forgot to initialize s before the if checks. The original code only sets s inside specific isinstance() blocks, so any unsupported input leaves s undefined.

内容的提问来源于stack exchange,提问作者userofstackoverflow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:15:44