如何在Python中使用自定义Snowball .sbl算法及立陶宛语词干提取器?
Got it, let's walk through how to get that Lithuanian Snowball stemmer working in Python step by step. I've tinkered with custom Snowball stemmers before, so here's what you need to do:
First, make sure you have the Snowball compiler installed—it's the tool that converts the .sbl rule file into usable C code.
- On Linux: Grab it via your package manager with
sudo apt install snowball - On macOS: Use Homebrew:
brew install snowball - On Windows: You'll need to build it from the Snowball source (no pre-built packages, but the build process is straightforward with MinGW or MSVC)
Once installed, run this command in your terminal, replacing lithuanian.sbl with your actual .sbl filename:
snowball -o lithuanian.c lithuanian.sbl
This spits out a lithuanian.c file with the stemmer logic translated to C.
Option A: Quick and dirty with ctypes
This method lets you use the compiled C code directly without modifying existing Python packages.
First, compile the C file into a shared library:
- Linux/macOS:
gcc -shared -fPIC -o liblithuanian.so lithuanian.c - Windows (with MinGW):
gcc -shared -o lithuanian.dll lithuanian.c
Then, write a small Python wrapper using ctypes to call the stemmer functions:
import ctypes # Load the compiled shared library lib = ctypes.CDLL('./liblithuanian.so') # Use './lithuanian.dll' on Windows # Define the opaque stemmer struct (we don't need to know its internal details) class Stemmer(ctypes.Structure): pass # Set up function signatures for the Snowball API lib.stemmer_new.restype = ctypes.POINTER(Stemmer) # Initialize the stemmer—pass the language name as bytes (must match your .sbl's definition) stemmer = lib.stemmer_new(b'lithuanian') lib.stemmer_stem.restype = ctypes.POINTER(ctypes.c_char) lib.stemmer_stem.argtypes = [ctypes.POINTER(Stemmer), ctypes.POINTER(ctypes.c_char), ctypes.c_int] def stem_word(word): """Stem a single Lithuanian word""" word_bytes = word.encode('utf-8') # Call the stemmer function result_ptr = lib.stemmer_stem(stemmer, word_bytes, len(word_bytes)) # Copy the result from the static buffer and decode to string return ctypes.string_at(result_ptr).decode('utf-8') # Test it out! print(stem_word("vilnius")) print(stem_word("lietuviškas")) # Clean up resources when done lib.stemmer_delete.argtypes = [ctypes.POINTER(Stemmer)] lib.stemmer_delete(stemmer)
Option B: Build a custom PyStemmer version
If you prefer using the familiar PyStemmer interface, you can add your Lithuanian stemmer to the package:
- Grab the PyStemmer source code from its official repository.
- Copy your
lithuanian.cfile into thesrc/libstemmerdirectory of the source code. - Open
src/libstemmer/modules.txtand add a new line with justlithuanian(this tells the build system to include your stemmer). - Recompile and install PyStemmer:
python setup.py build_ext --inplace pip install .
Once installed, you can use it just like any other built-in stemmer:
import Stemmer stemmer = Stemmer.Stemmer('lithuanian') print(stemmer.stemWord("lietuviškas")) # Or stem multiple words at once print(stemmer.stemWords(["vilnius", "kaunas", "klaipėda"]))
Make sure your .sbl file follows Snowball's syntax correctly—if you run into compilation errors, double-check the rule definitions. The Snowball documentation has examples if you need to tweak anything.
内容的提问来源于stack exchange,提问作者Lukas

