调用Pattern库ngrams方法触发UnboundLocalError问题求助
Let's break down why you're hitting this error and how to fix it quickly.
Root Cause
Looking at your code, you're passing a list of strings to Pattern's ngrams() function (since each value in your RDD is a list like [u'Campbell was born in Myrtle Bank.']). But if you check the ngrams() source code, it only handles three input types: basestring, Sentence, and Text.
When you pass a list, none of the isinstance() checks trigger, so the variable s never gets assigned. Then when the code reaches for s in s:, it throws an UnboundLocalError because s doesn't exist yet.
Solutions
1. Fix Your Input to ngrams() (Recommended)
Instead of passing a list of strings directly, join them into a single string first. This fits what ngrams() expects, and you won't need to modify the Pattern library.
Modify your business code like this:
corpus = text\ .mapValues(lambda v: ngrams(' '.join(v), self.max_ngram))\ # Join list into one string .flatMap(lambda (target, tokens): (((target, t), 1) for t in tokens))\ .reduceByKey(add)\ .map(lambda ((target, token), count): (token, (target, count)))
This works because ' '.join(v) turns your list of sentences into a single string, which ngrams() can tokenize and process correctly—triggering the isinstance(string, basestring) branch and properly initializing s.
2. Patch Pattern's ngrams() Function
If you need to keep passing lists directly (e.g., to preserve sentence boundaries without joining), modify the ngrams() function to handle list inputs. Update the function in pattern/text/__init__.py like this:
def ngrams(string, n=3, punctuation=PUNCTUATION, continuous=False): """ Returns a list of n-grams (tuples of n successive words) from the given string. Alternatively, you can supply a Text or Sentence object. With continuous=False, n-grams will not run over sentence markers (i.e., .!?). Punctuation marks are stripped from words. """ def strip_punctuation(s, punctuation=set(punctuation)): return [w for w in s if (isinstance(w, Word) and w.string or w) not in punctuation] if n <= 0: return [] s = [] # Initialize s with default empty list if isinstance(string, basestring): s = [strip_punctuation(s.split(" ")) for s in tokenize(string)] elif isinstance(string, Sentence): s = [strip_punctuation(string)] elif isinstance(string, Text): s = [strip_punctuation(s) for s in string] elif isinstance(string, list): # Handle list of strings: process each element as a separate text segment for item in string: if isinstance(item, basestring): s.extend([strip_punctuation(split_item.split(" ")) for split_item in tokenize(item)]) # Add an else clause if you want to handle other types gracefully if continuous: s = [sum(s, [])] g = [] # Rename loop variable to avoid overwriting the s list for segment in s: g.extend([tuple(segment[i:i+n]) for i in range(len(segment)-n+1)]) return g
Key changes here:
- Initialize
sas an empty list upfront to avoid undefined errors. - Add an
elif isinstance(string, list)branch to process each string in your list. - Rename the loop variable from
stosegmentto prevent accidentally overwriting theslist (a subtle bug in the original code!).
3. Validate Your RDD Data
Double-check that your text processing pipeline is outputting the data you expect. Your initial text.map(...).groupByKey().map(...) produces tuples where the value is a list of strings—this is correct, but you just need to adapt it for ngrams() as shown above.
Why Your Previous Library Modification Failed
Chances are you either didn't add the list handling branch, or you forgot to initialize s before the if checks. The original code only sets s inside specific isinstance() blocks, so any unsupported input leaves s undefined.
内容的提问来源于stack exchange,提问作者userofstackoverflow

