A character-recognition system for Hangeul

Författare

Johan Sageryd

Summary, in English

This work presents a rule-based character-recognition system for the Korean script, Hangeul. An input raster image representing one Korean character (Hangeul syllable) is thinned down to a skeleton, and the individual lines extracted. The lines, along with information on how they are interconnected, are translated into a set of hierarchical graphs, which can be easily traversed and compared with a set of reference structures represented in the same way. Hangeul consists of consonant and vowel graphemes, which are combined into blocks representing syllables. Each reference structure describes one possible variant of such a grapheme. The reference structures that best match the structures found in the input are combined to form a full Hangeul syllable. Testing all of the 11 172 possible characters, each rendered as a 200-pixel-squared raster image using the gothic font AppleGothic Regular, had a recognition accuracy of 80.6 percent. No separation logic exists to be able to handle characters whose graphemes are overlapping or conjoined; with such characters removed from the set, thereby reducing the total number of characters to 9 352, an accuracy of 96.3 percent was reached. Hand-written characters were also recognised, to a certain degree. The work shows that it is possible to create a workable character-recognition system with reasonably simple means.

Avdelning/ar

Language Technology Program

Publiceringsår

2009

Språk

Engelska

Fulltext

Dokumenttyp

Examensarbete för magisterexamen (Ett år)

Ämne

Technology and Engineering
Languages and Literatures

Nyckelord

graph matching
graph thinning
Hangeul
Korean
character recognition

Handledare

Johan Frid

A character-recognition system for Hangeul

Summary, in English

Kontaktinformation

Information om www.lu.se

Följ oss på sociala medier

Samarbeten och nätverk