You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python语言编码概念错误:芬兰语借词识别函数开发问询

实现芬兰语借词识别函数 get_loan_words(filename)

Got it, let's break down how to build this function step by step, based on the rules you laid out for Finnish loanwords!

核心规则回顾

First, let's recap the key marker: Finnish loanwords contain at least one of these non-standard orthographic letters: å, c, q, x, f, g, w, z (and we need to account for both uppercase and lowercase versions).

Step 1: Read and process the file

We'll need to read the text file, extract individual words, and filter out any punctuation or non-word characters that might stick to the edges of words. Regular expressions work great for this.

Step 2: Identify loanwords

For each extracted word, check if it includes any of the target letters. Using a set for the target letters makes this check fast and efficient.

Full Code Implementation

import re

def get_loan_words(filename):
    # Define the set of loanword marker characters (lowercase for case-insensitive check)
    loan_chars = {'å', 'c', 'q', 'x', 'f', 'g', 'w', 'z'}
    loan_words = set()  # Use a set to avoid duplicate words
    
    try:
        # Open and read the file with UTF-8 encoding (standard for Finnish text)
        with open(filename, 'r', encoding='utf-8') as file:
            content = file.read()
            # Extract all valid Finnish words (includes standard ä/ö and their uppercase variants)
            words = re.findall(r'\b[a-zA-ZåäöÅÄÖ]+\b', content)
            
            for word in words:
                lower_word = word.lower()
                # Check if any character in the word is a loanword marker
                if any(char in loan_chars for char in lower_word):
                    loan_words.add(word)
        
        # Return sorted list of unique loanwords for readability
        return sorted(loan_words)
    except FileNotFoundError:
        print(f"Oops! The file {filename} couldn't be found.")
        return []
    except Exception as e:
        print(f"Something went wrong while reading the file: {e}")
        return []

Code Explanations

  • Regex for word extraction: r'\b[a-zA-ZåäöÅÄÖ]+\b' ensures we only grab actual Finnish words, excluding punctuation, numbers, and other non-letter characters. We include the standard Finnish special characters ä and ö so we don't miss valid words.
  • Case insensitivity: Converting each word to lowercase before checking means we catch both Cafe and cafe as loanwords.
  • Duplicate prevention: Using a set stores each loanword only once, even if it appears multiple times in the text. We convert it to a sorted list at the end for a clean output.
  • Error handling: Basic exception handling makes the function more robust, letting the user know if the file is missing or if another issue pops up.

Test Example

If your text file suomi_teksti.txt has this content:

Minä käysin Berliinissä ja ostin waffeleita. Cafeissa oli mahtava kahvi, ja siellä oli myös q-tasoinen kone.

Calling get_loan_words("suomi_teksti.txt") will return:

['Berliinissä', 'Cafeissa', 'kone', 'q-tasoinen', 'waffeleita']

All these words contain at least one of the loanword marker characters, so they're correctly identified.

内容的提问来源于stack exchange,提问作者John

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:28:59