Python语言编码概念错误:芬兰语借词识别函数开发问询
get_loan_words(filename) Got it, let's break down how to build this function step by step, based on the rules you laid out for Finnish loanwords!
核心规则回顾
First, let's recap the key marker: Finnish loanwords contain at least one of these non-standard orthographic letters: å, c, q, x, f, g, w, z (and we need to account for both uppercase and lowercase versions).
Step 1: Read and process the file
We'll need to read the text file, extract individual words, and filter out any punctuation or non-word characters that might stick to the edges of words. Regular expressions work great for this.
Step 2: Identify loanwords
For each extracted word, check if it includes any of the target letters. Using a set for the target letters makes this check fast and efficient.
Full Code Implementation
import re def get_loan_words(filename): # Define the set of loanword marker characters (lowercase for case-insensitive check) loan_chars = {'å', 'c', 'q', 'x', 'f', 'g', 'w', 'z'} loan_words = set() # Use a set to avoid duplicate words try: # Open and read the file with UTF-8 encoding (standard for Finnish text) with open(filename, 'r', encoding='utf-8') as file: content = file.read() # Extract all valid Finnish words (includes standard ä/ö and their uppercase variants) words = re.findall(r'\b[a-zA-ZåäöÅÄÖ]+\b', content) for word in words: lower_word = word.lower() # Check if any character in the word is a loanword marker if any(char in loan_chars for char in lower_word): loan_words.add(word) # Return sorted list of unique loanwords for readability return sorted(loan_words) except FileNotFoundError: print(f"Oops! The file {filename} couldn't be found.") return [] except Exception as e: print(f"Something went wrong while reading the file: {e}") return []
Code Explanations
- Regex for word extraction:
r'\b[a-zA-ZåäöÅÄÖ]+\b'ensures we only grab actual Finnish words, excluding punctuation, numbers, and other non-letter characters. We include the standard Finnish special charactersäandöso we don't miss valid words. - Case insensitivity: Converting each word to lowercase before checking means we catch both
Cafeandcafeas loanwords. - Duplicate prevention: Using a
setstores each loanword only once, even if it appears multiple times in the text. We convert it to a sorted list at the end for a clean output. - Error handling: Basic exception handling makes the function more robust, letting the user know if the file is missing or if another issue pops up.
Test Example
If your text file suomi_teksti.txt has this content:
Minä käysin Berliinissä ja ostin waffeleita. Cafeissa oli mahtava kahvi, ja siellä oli myös q-tasoinen kone.
Calling get_loan_words("suomi_teksti.txt") will return:
['Berliinissä', 'Cafeissa', 'kone', 'q-tasoinen', 'waffeleita']
All these words contain at least one of the loanword marker characters, so they're correctly identified.
内容的提问来源于stack exchange,提问作者John

