You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python塞尔维亚语日期本地化及西里尔文转拉丁文方案问询

Fixing Serbian Cyrillic Display & Cyrillic-to-Latin Conversion in Python (CMD Compatible)

Problem Breakdown

Your code is hitting two main issues:

  • Cyrillic garbled output: Windows CMD's default encoding doesn't align with the locale settings, turning Serbian Cyrillic characters into question marks.
  • Missing universal Cyrillic-to-Latin conversion: You need a reliable way to translate Cyrillic weekdays to their Latinized counterparts to match your text file's content.

Step 1: Fix Cyrillic Display in Python & CMD

First, we'll ensure Python outputs UTF-8 correctly and tweak CMD to support Cyrillic rendering.

Adjust Python's Output Encoding

Update your code to force UTF-8 output and add cross-system locale fallbacks:

import sys
import calendar
import locale

# Force standard output to use UTF-8
sys.stdout.reconfigure(encoding='utf-8')

# Set Serbian locale (handles Linux/macOS and Windows differences)
try:
    locale.setlocale(locale.LC_ALL, 'sr_SP.UTF-8')  # Linux/macOS
except:
    locale.setlocale(locale.LC_ALL, 'sr-SP')  # Windows fallback

def main():
    cyrillic_days = list(calendar.day_name)
    print("Serbian Cyrillic weekdays:")
    print(cyrillic_days)
    
    day_input = input("Enter a weekday (Cyrillic): ")
    if day_input in cyrillic_days:
        print(f"Match found: {day_input}")
    else:
        print("Error: Input not in weekday list")

if __name__ == "__main__":
    main()

Configure CMD for Cyrillic

Before running the script in CMD, execute this command to switch to UTF-8 encoding:

chcp 65001

This lets CMD render UTF-8 Cyrillic characters correctly.


Step 2: Add Official Serbian Cyrillic-to-Latin Conversion

We'll use a mapping dictionary following Serbia's official transliteration rules to convert Cyrillic to Latinized text:

import sys
import calendar
import locale

# Official Serbian Cyrillic to Latin transliteration mapping
cyrillic_to_latin = {
    'а': 'a', 'б': 'b', 'в': 'v', 'г': 'g', 'д': 'd', 'ђ': 'đ',
    'е': 'e', 'ж': 'ž', 'з': 'z', 'и': 'i', 'ј': 'j', 'к': 'k',
    'л': 'l', 'љ': 'lj', 'м': 'm', 'н': 'n', 'њ': 'nj', 'о': 'o',
    'п': 'p', 'р': 'r', 'с': 's', 'т': 't', 'ћ': 'ć', 'у': 'u',
    'ф': 'f', 'х': 'h', 'ц': 'c', 'ч': 'č', 'џ': 'dž', 'ш': 'š',
    'А': 'A', 'Б': 'B', 'В': 'V', 'Г': 'G', 'Д': 'D', 'Ђ': 'Đ',
    'Е': 'E', 'Ж': 'Ž', 'З': 'Z', 'И': 'I', 'Ј': 'J', 'К': 'K',
    'Л': 'L', 'Љ': 'Lj', 'М': 'M', 'Н': 'N', 'Њ': 'Nj', 'О': 'O',
    'П': 'P', 'Р': 'R', 'С': 'S', 'Т': 'T', 'Ћ': 'Ć', 'У': 'U',
    'Ф': 'F', 'Х': 'H', 'Ц': 'C', 'Ч': 'Č', 'Џ': 'Dž', 'Ш': 'Š'
}

def convert_cyrillic_to_latin(text):
    return ''.join([cyrillic_to_latin.get(char, char) for char in text])

# Force UTF-8 output
sys.stdout.reconfigure(encoding='utf-8')

# Set locale with fallback
try:
    locale.setlocale(locale.LC_ALL, 'sr_SP.UTF-8')
except:
    locale.setlocale(locale.LC_ALL, 'sr-SP')

def main():
    cyrillic_days = list(calendar.day_name)
    latin_days = [convert_cyrillic_to_latin(day) for day in cyrillic_days]
    
    print("Cyrillic weekdays:", cyrillic_days)
    print("Latinized weekdays:", latin_days)
    
    # Read Latinized weekdays from text file (one per line)
    with open('latin_weekdays.txt', 'r', encoding='utf-8') as f:
        file_latin_days = [line.strip() for line in f]
    
    day_input = input("Enter a Cyrillic weekday: ")
    if day_input in cyrillic_days:
        converted_day = convert_cyrillic_to_latin(day_input)
        print(f"Latinized version: {converted_day}")
        if converted_day in file_latin_days:
            print("Matches the text file content!")
        else:
            print("Does not match the text file content!")
    else:
        print("Error: Input not in Cyrillic weekday list")

if __name__ == "__main__":
    main()

Important Notes

  • Locale compatibility: Different OSes use different locale names (e.g., sr_SP.UTF-8 for Linux/macOS, sr-SP for Windows) — the try-except block handles this automatically.
  • CMD requirement: Always run chcp 65001 before executing the script in CMD; without this, even correct UTF-8 output will appear garbled.
  • Transliteration accuracy: The mapping follows Serbia's official rules, including special characters like ђ → đ and љ → lj.

内容的提问来源于stack exchange,提问作者Neo Cortex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:47:17